Google's Gemini 3.8 Flash arrives as AI token pricing competition sharpens across the sector
The AI inference market is repricing noticeably, and the cadence of model releases has accelerated to match. Google has released Gemini 3.8 Flash, its third Flash-series model in six weeks, and the company is pitching…
Key takeaways
- Google has released Gemini 3.8 Flash, its third Flash-series model in six weeks, with introductory API rates of $0.75 per million input tokens and $3.75 per million output tokens through year-end.
- The release ships two variants: a standard general-purpose model for agentic tasks and coding, and Gemini 3.8 Flash Cyber, tuned for vulnerability detection and mitigation.
- When the promotional window closes, listed rates rise to $1.50 per million input tokens and $7.50 per million output tokens.
- No frontier-level Gemini Pro model has shipped since early 2026, raising developer questions about whether the promised Gemini 3.5 Pro will arrive at all.
- Google's introductory pricing reflects sector-wide pressure as other AI labs have cut rates amid more cautious business spending on AI.
The AI inference market is repricing noticeably, and the cadence of model releases has accelerated to match. Google has released Gemini 3.8 Flash, its third Flash-series model in six weeks, and the company is pitching introductory API rates of $0.75 per million input tokens and $3.75 per million output tokens through the end of the year.
Two variants ship simultaneously. The standard Gemini 3.8 Flash is positioned as a general-purpose workhorse suited to agentic tasks and software development, and Google describes the combined package as its best reasoning and coding model to date. The second, Gemini 3.8 Flash Cyber, runs on the same underlying model but has been tuned specifically for vulnerability detection and mitigation.
The pricing signal
Google frames the introductory rate as a limited window. When it closes, listed rates move to $1.50 per million input tokens and $7.50 per million output tokens. At the current release pace, new models are likely to arrive before the promotional period ends.
The per-token competition matters at the sector level. Other AI labs have recently cut their rates as businesses have grown more cautious about AI spending. Google's introductory pricing reflects the same pressure: price is currently doing work that model differentiation alone cannot.
What the release cadence signals
No frontier-level Gemini Pro model has shipped since early 2026, and the frequency of Flash releases has raised questions in the developer community about whether Gemini 3.5 Pro, which was promised, will arrive at all. Google has not committed to a timeline. The pitch for this release, API access at introductory rates, mirrors what the company offered developers for Gemini 3.7 Flash just a couple of weeks ago.
The Cyber variant is the detail that separates this release from a straightforward iteration. Tuning the Flash base specifically for vulnerability detection and mitigation gives security teams a more focused tool. That is a different commercial bet than releasing another general reasoning upgrade.
The read-through for the broader sector: when the largest model providers hold promotional pricing to retain developer adoption, the implied cost of switching between platforms falls. Google's third Flash release in six weeks is, on balance, as much a pricing story as a capability one.
Related reading
Source · 來源