Gemini 3.6 Flash Ships — and the Savings Are in the Output Column

Gemini 3.6 Flash release, July 21, 2026

In This Article

  1. What Google shipped on July 21
  2. The price change, stated precisely
  3. What did not change
  4. The benchmark numbers
  5. Flash Cyber is a pilot, not a product
  6. Why it matters
  7. Common questions

Key Takeaways

Google released three Gemini models on July 21, 2026, and the version numbers do not read left to right. In a post bylined by Tulsee Doshi, Senior Director of Product Management, the company announced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber together. TechCrunch and 9to5Google reported the same three models on the same day. Despite carrying the higher number, 3.6 Flash is the newer generation; the 3.5-labeled models beside it are not its successors. The flagship 3.5 Pro was absent from the whole announcement.

What Google shipped on July 21

Two of the three are generally available. The Gemini API changelog entry for July 21, 2026 lists gemini-3.6-flash and gemini-3.5-flash-lite as GA — stable, production-ready, not preview. It describes 3.6 Flash as offering improved token efficiency and code and agentic planning at a lower price point than 3.5 Flash, and Flash-Lite as a low-latency, cost-effective subagent option for high-volume automation. The same entry notes that the temperature, top_p and top_k sampling parameters are deprecated for these latest models — a small change with real consequences if your code sets them.

Google lists availability across the Gemini API, Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, the Gemini Enterprise app, and the Gemini app, with Flash-Lite additionally going to Google Search. Third-party surfaces moved the same day: GitHub's changelog put Gemini 3.6 Flash into Copilot for Pro, Pro+, Max, Business and Enterprise, across VS Code, Visual Studio, Copilot CLI, the Copilot cloud agent, the Copilot app, JetBrains, Xcode and Eclipse. GitHub says the rollout is gradual and that Business and Enterprise administrators must enable the "Gemini 3.6 Flash Preview" policy in Copilot settings first. In GitHub's own words, its early testing found the model "demonstrated higher task-completion rates and better token efficiency than Gemini 3.5 Flash across coding and agentic workflows."

The price change, stated precisely

This release has been summarized as "cheaper," and that is half right. Per Google's pricing page, standard tier, per million tokens:

Standard-tier pricing, per 1M tokens

ModelInputOutput
Gemini 3.6 Flash$1.50$7.50
Gemini 3.5 Flash (predecessor)$1.50$9.00
Gemini 3.5 Flash-Lite$0.30$2.50

The input price is identical. Only output fell, from $9.00 to $7.50 — a 16.7% cut. Context caching for 3.6 Flash is listed at $0.15 per million tokens plus $1.00 per million tokens per hour of storage, and Google's page shows no context-length price tiering on the standard tier for these models, with Batch and Flex modes at 50% off. Flash Cyber does not appear in the pricing table at all.

The second half of the savings is behavioral rather than posted. Google says 3.6 Flash consumes about 17% fewer output tokens than 3.5 Flash; 9to5Google attributes that figure to the Artificial Analysis Index rather than to Google's own measurement. Google's blog separately cites reductions of up to 65% on the DeepSWE benchmark specifically — a best case on one evaluation, not a general rate.

16.7%
The output-price cut, $9.00 to $7.50 per million tokens. Input held flat at $1.50.
Separately, Google reports about 17% fewer output tokens consumed, a figure 9to5Google attributes to the Artificial Analysis Index.

What did not change

The context window. Google's model pages list 1,048,576 input tokens and a 65,536-token output limit for gemini-3.6-flash and the identical figures for gemini-3.5-flash. The DeepMind model card phrases the same thing as "up to 1M" context and a 64K token output. Inputs remain text, image, video, audio and PDF; output is text. If you have been reading this release as a context upgrade, it is not one — and our guide to what a million-token context window really buys you still applies unchanged.

What did move is recency. The DeepMind model card gives a March 2026 knowledge cutoff, with the caveat that some domains have limited information back to January 2025. 9to5Google independently reports the cutoff advancing from January 2025 to March 2026.

The benchmark numbers

Google's blog gives four head-to-head comparisons against 3.5 Flash: DeepSWE 49% vs. 37%, MLE-Bench 63.9% vs. 49.7%, OSWorld-Verified 83.0% vs. 78.4%, and GDPval-AA v2 1421 vs. 1349. The DeepMind model card adds absolute scores that were not in that comparison set: SWE-Bench Pro 58.7%, CharXiv 85.2% without tools and 89.4% with, GDM-MRCR v2 long context at 91.8% for 128k and 54.0% at 1M, and Terminal-bench 2.1 at 78.0%.

Flash-Lite's numbers are its own and should not be read as 3.6 Flash results: per Google's blog, SWE-Bench Pro 54.2% against 49.6% for the prior Flash-Lite generation, Terminal-Bench 2.1 at 54% against 31%, and GDM-MRCR v2 at 72.2% against 60.1%.

One number we are deliberately not printing is throughput. Published tokens-per-second figures for this model vary widely across trackers, and Artificial Analysis re-measures over time, so any single number would be stale on arrival. Google's blog gives an output-speed figure for Flash-Lite only, not for 3.6 Flash.

Flash Cyber is a pilot, not a product

Gemini 3.5 Flash Cyber is fine-tuned to find and fix cybersecurity vulnerabilities. Google states access is limited to governments and trusted partners through the CodeMender limited-access pilot program, and cites "competitive frontier performance" on the CyberGym benchmark without publishing a number. It does not appear in the public API changelog or the pricing table, which is consistent with restricted access.

Read the access terms before you plan around it

For teams working in regulated or federal environments, a vulnerability-finding model is the most interesting item in this release — and the one you cannot currently buy. It is a pilot with a named access gate, no published benchmark score, and no listed price. Treat it as a signal about direction, not as a tool you can put on a roadmap.

Why it matters

The following is our analysis, not reported fact. The economics of this release favor output-heavy workloads specifically. If your traffic is long prompts producing short answers — document extraction, classification, retrieval over big contexts — the posted price change does nothing for you, because input is where your tokens are and input did not move. If your traffic is agentic, where the model reasons at length and emits long tool-calling traces, you get the 16.7% posted cut and, if the efficiency figure holds on your workload, fewer tokens emitted on top of it. Those compound. That is a meaningfully different bill for a coding agent than for a RAG pipeline, and it is worth running against your own token mix in our LLM price calculator before assuming a flat saving.

The second thing worth naming is the shape of Google's release cadence. A Flash-tier model shipped GA while the Pro flagship stayed in partner testing. Google's blog says 3.5 Pro is currently testing with partners and will be broadly available as soon as it is ready, and that Gemini 4 pre-training has begun, described as the company's most ambitious pre-training run yet. TechCrunch reports Logan Kilpatrick saying the team is testing 3.5 Pro with partners and hopes it will "land soon," and attributes to Bloomberg that Google faced internal delays meeting performance goals — we did not read the Bloomberg original, so take that as a reported chain. We covered the delay itself in Gemini 3.5 Pro's missed deadlines.

For practitioners the practical read is unchanged by any of it: the mid-tier is where the real work runs, and it is improving faster than the flagship ships. Test 3.6 Flash against whatever you currently default to, on your own tasks, and check whether the deprecated sampling parameters break anything before you promote it. Our guide on choosing a model tier covers the method; the complete Gemini guide covers the family.

Run it against your own token mix

The savings here depend entirely on whether your workload is input-heavy or output-heavy. Our free calculator compares Gemini 3.6 Flash against Flash-Lite, GPT-5.6 and Claude on your actual volumes.

Open the price calculator

Sources: Google — Introducing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber; Gemini API changelog; Gemini API pricing; gemini-3.6-flash model docs; gemini-3.5-flash model docs; DeepMind model card; TechCrunch; 9to5Google; GitHub Changelog. Analysis and framing by Precision AI Academy.

Common questions

Is Gemini 3.6 Flash actually cheaper than 3.5 Flash? On output, yes — $7.50 against $9.00 per million tokens. On input, no: both are $1.50. Whether your bill falls depends on your input-to-output ratio.

Why is 3.6 newer than the 3.5 models released beside it? The version numbers track model generations, not release order. All three shipped July 21, 2026, and 3.6 Flash is the newer generation. Flash-Lite and Flash Cyber carry 3.5 labels.

Can I use Gemini 3.5 Flash Cyber? Only if you are in the CodeMender limited-access pilot, which Google says is open to governments and trusted partners. It is not in the public API or pricing table.

Do I need to change my API code? Possibly. The changelog lists temperature, top_p and top_k as deprecated for the latest models. Check whether your client sets them.

Where is Gemini 3.5 Pro? Still not generally available as of this release. Google says it is testing with partners and will ship when ready, and that Gemini 4 pre-training has begun.

About Precision AI Academy

Precision AI Academy publishes practical AI news, plain-language analysis, and 137 free courses for builders and working professionals. It is a sister site of Precision Federal, a federal software and AI firm. We verify the numbers, cite the primary sources, and skip the hype.