Gemini 3.8 Flash pricing begins at the same global token rates as its predecessor, but Google says the new model can consume more tokens on difficult jobs. That leaves developers with two distinct cost questions: how much the model uses per task today, and Google’s scheduled rate increase on Jan. 1, 2027.
Google announced Gemini 3.8 Flash and the restricted Gemini 3.8 Flash Cyber on Sept. 2. The company says the general-purpose model improves on Gemini 3.7 Flash for software engineering, autonomous-agent work and multi-step reasoning. Those are Google’s performance claims, based in part on the company’s reported benchmark results, not a guarantee that every production workload will perform better.
For now, Google lists global introductory rates of $0.75 per 1 million input tokens and $3.75 per 1 million output tokens through Dec. 31, 2026. Google Cloud’s pricing page says those global rates will rise to $1.50 for input and $7.50 for output per million tokens beginning Jan. 1, 2027. Regional pricing can differ.
Why can Gemini 3.8 Flash cost more at the same token price?
Google’s answer is that 3.8 Flash may do more work on complex requests. The company says it can take additional reasoning steps and repeatedly call tools, particularly when users select higher effort levels. Google also warns that the model may use more tokens to pursue better performance.
That means equal per-token rates do not produce equal per-task spending. A task that uses more input or output tokens costs more at the same listed rate. The extra usage is workload-dependent, so there is no one percentage increase that applies to every application.
Artificial Analysis estimated that 3.8 Flash cost about 40% more per task than 3.7 Flash in its own measurements, according to reporting by The Verge and The Register. That estimate was tied to higher output-token use and more agent turns in the evaluator’s tests. It is not Google billing data, and it should not be read as a forecast for every customer’s bill.
What can developers do to limit Gemini 3.8 Flash costs?
Google recommends lower effort levels for applications where compute efficiency matters more than extra reasoning. It also says Gemini 3.7 Flash remains supported for efficiency-first workloads. Developers using paid Gemini API access can monitor usage in Google AI Studio, according to Google’s billing documentation.
The other price change is much less conditional. Google’s published global input and output rates for both 3.8 Flash and 3.7 Flash are set to double after the introductory period ends. Teams evaluating the newer model should therefore compare token consumption on their own jobs before year-end, while also budgeting for the listed 2027 rates.
Google also introduced Gemini 3.8 Flash Cyber, which it says is aimed at vulnerability detection and automated patching. Access is limited to trusted defenders through Google’s Fairwind Program, rather than offered as a general public model.
Google’s launch details are available in its announcement, and its Cloud pricing page lists the introductory and scheduled standard rates.
This story draws on original reporting from The Verge.