Google has released Gemini 3.6 Flash, replacing the 3.5 Flash model it promoted at I/O in May, and added two more variants: Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber. The practical pitch to developers is lower cost and fewer wasted tokens, which is the part of the AI boom that finance departments have started reading.
The missing piece is Gemini 3.5 Pro. Google said at I/O that its next flagship model was due in June. That did not happen. Google now says the model is being tested with unnamed partners and will ship “as soon as it’s ready,” according to the company.
What changed in Gemini 3.6 Flash
Google says Gemini 3.6 Flash responds to feedback on 3.5 Flash, with better coding, stronger multimodal performance, and standard support for computer use through the Gemini API. Those are company claims, and benchmarks remain benchmarks, but Google did publish some numbers.
On DeepSWE, a coding evaluation, Google says 3.6 Flash scored 49 percent, up from 37 percent for 3.5 Flash. On OSWorld, which tests computer-use tasks, the new model reached 83 percent, compared with 78.4 percent for its predecessor.
The larger change may be token use. Google says Gemini 3.6 Flash uses about 17 percent fewer tokens than 3.5 Flash. In agent-style workflows, where a model may call tools, inspect results, and try again, fewer steps and fewer tokens can cut the bill quickly.
Google is also lowering the output price. Gemini 3.6 Flash costs $1.50 per 1 million input tokens and $7.50 per 1 million output tokens through the API. Gemini 3.5 Flash was priced at $1.50 for input and $9 for output.
Gemini 3.6 Flash begins rolling out in the API now and will replace 3.5 Flash in the Gemini app, according to Google.
Flash Lite gets cheaper speed, mostly
Google is keeping the 3.5 line alive with Gemini 3.5 Flash Lite. The company calls it its most efficient modern AI model and says it can generate 350 tokens per second. Google positions it for large-scale agent systems where latency and cost matter more than chasing frontier-model bragging rights.
The price is $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. That is still more expensive than the earlier 3.1 Flash Lite, which cost $0.25 for input and $1.50 for output. Google says 3.5 Flash Lite is available to developers and in the Gemini app, and will appear frequently in Google Search, where its speed may fit AI Overviews.
A cyber model with restricted access
Gemini 3.5 Flash Cyber is Google’s first large language model tuned specifically for cybersecurity work. Google says it performs nearly as well at finding and fixing security problems as Anthropic’s larger and more expensive Claude Mythos model, while keeping the efficiency profile of Flash.
Google also acknowledges the obvious problem: a model that can find vulnerabilities can help defenders or attackers. The company says Gemini 3.5 Flash Cyber will launch soon only as a limited pilot inside Google DeepMind’s CodeMender agent. Access will be restricted to trusted partners and governments, according to Google.
Google also says it has started pre-training Gemini 4, which it describes as more ambitious than its earlier AI efforts. The company gave no release schedule and did not say whether more 3.x models will arrive before Gemini 4.
This story draws on original reporting from Ars Technica.