The short answer
Gemini 3.7 Flash arrived on 13 August 2026, three weeks after 3.6 Flash, with a 16-point DeepSWE improvement and an introductory price that halves the per-token cost until 31 December 2026. At $0.75 per million input tokens and $3.75 per million output tokens, it is now cheaper than 3.6 Flash ($1.50 / $7.50) while performing measurably better on coding and automation benchmarks. It ranks first on FrontierCode 1.1 Main — ahead of Claude Sonnet 5 and GPT-5.6 Terra — and scores 34% on GDP.pdf enterprise document comprehension, compared to 28% for Sonnet 5 and 24.7% for GPT-5.6 Terra.
The right response is the same as with any model release: route a slice of real agent traffic to 3.7 Flash, measure cost per completed task and revision rate on your own workloads, and promote it only where quality holds.
What Google shipped
Gemini 3.7 Flash is positioned as Google's "most intelligent workhorse model" for coding and agents, arriving 23 days after 3.6 Flash. The model keeps the same 1-million-token context window and up to 64,000 output tokens as its predecessor but delivers better first-pass accuracy, improved instruction adherence, and more deliberate multi-step planning. It uses the same transformer-based mixture-of-experts architecture as 3.6 Flash.
Availability is broad from day one. Developer access is through the Gemini API and Google AI Studio; Android Studio includes 3.7 Flash for in-IDE coding assistance; Google Antigravity supports it for agent-first workflows; and the Gemini Enterprise Agent Platform gives enterprises a governed deployment path. For consumer users, it is rolling out through Gemini Spark to AI Pro and Ultra subscribers in 160+ countries. Teams already building AI agent pipelines on Google's stack can access it today with no migration required.
Google notes the model “better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity,” and “puts in more effort into multi-step planning and tool calls.” For agentic workloads, those improvements translate to fewer failed tool calls and fewer revision cycles per completed task — which matters more to cost than the per-token rate alone.
Benchmark results and competition
The benchmark story is notable because 3.7 Flash does not just improve on 3.6 Flash — it beats the current frontier models from Anthropic and OpenAI across multiple coding-relevant tasks.
DeepSWE v1.1 (long-horizon software engineering): 65.3% for 3.7 Flash vs 49.0% for 3.6 Flash — a 16.3-point gain. This benchmark measures how well an agent completes multi-step engineering tasks from scratch, making it one of the most relevant indicators for teams building automated coding pipelines.
FrontierCode 1.1 Main (production code quality across 100 programming tasks in multiple languages): 43.6% vs 34.4%. Google claims first place overall, ahead of Claude Sonnet 5 and GPT-5.6 Terra. This benchmark emphasises compile-correctness, code style, and task completion rather than just token prediction, so it tends to track real-world coding performance more closely than general reasoning evals.
GDP.pdf (enterprise business document comprehension): 34.0% vs 22.0% for 3.6 Flash, compared to 28% for Claude Sonnet 5 and 24.7% for GPT-5.6 Terra. For teams building pipelines that process contracts, compliance documents, or financial reports alongside code, this matters.
AutomationBench (enterprise workflow automation): 30.4% vs 17.0% — a near-doubling that reflects better multi-tool orchestration and multi-step planning.
Arena.ai WebDev (web development tasks): Elo 1588 vs 1538 for 3.6 Flash.
The competitive gap is narrow in absolute percentage terms, but the direction is consistent: 3.7 Flash is currently the strongest coding-focused model at the efficiency tier, from any provider. For teams evaluating model choices for GenAI integration projects, it is now the benchmark to test against.
Pricing: half the cost through 2026
The pricing change is the part with the most immediate operational impact. Gemini 3.7 Flash launches at $0.75 per million input tokens and $3.75 per million output tokens — exactly half the $1.50 input and $7.50 output of Gemini 3.6 Flash. This introductory rate is confirmed through 31 December 2026, after which pricing rises to the 3.6 Flash-equivalent standard rates ($1.50 input / $7.50 output) from 1 January 2027.
For teams with high-volume agent fleets, the math is straightforward: if your workloads hold quality on 3.7 Flash, running the same jobs now costs half as much. A team spending $10,000 per month on 3.6 Flash agent tokens could spend $5,000 on 3.7 Flash through year-end, or keep spending the same amount and run twice the workload. Those savings compound across every agent run that completes successfully in fewer tokens — and the benchmark improvements suggest 3.7 Flash does indeed complete long-horizon tasks in fewer steps.
The standard-rate reset in January 2027 is worth noting in planning. Teams that optimise agent stacks around the introductory price should model the January cost impact now, or plan a re-evaluation before year-end when the next efficiency tier will likely be available anyway, given Google's three-week release cadence this summer.
What it means for US & EU teams
For US teams building product features that use AI agents — automated code review, CI pipeline triage, multi-step data extraction — Gemini 3.7 Flash is the first meaningful price cut alongside a quality improvement at the efficiency tier. The right action is not to flip a switch but to route a representative fraction of real traffic, measure cost-per-successful-task and failure rate, and promote it only where the numbers hold. Agent workloads are specific to your codebase, tools, and prompt design; benchmarks give you a prior, not a guarantee.
EU teams have an additional step. The model is available globally through the Gemini API and, for enterprises, through Vertex AI and the Gemini Enterprise Agent Platform. But availability is not the same as compliance. Workloads that process personal data — common in FinTech, HealthTech, and legal-tech pipelines — need the specific processing region, data handling terms, and logging configurations confirmed against GDPR and the EU AI Act before agents touch production data. Google offers EU-region endpoints through Vertex AI; confirm your deployment is on one of them if data residency matters.
The broader pattern worth tracking: Google has shipped three efficiency-tier Gemini updates in roughly two months (3.5 Flash, 3.6 Flash, 3.7 Flash), while Gemini 3.5 Pro — the flagship tier — remains delayed. Gemini 3.5 Pro was last updated in February 2026 and has not shipped as of this writing, according to Bloomberg. The competitive race at the moment is in efficiency and price, not frontier capability. For engineering leaders evaluating AI model strategy, that means the economics of high-volume agentic work are improving faster than raw model intelligence, and the winning stack is one that can swap the routing layer cheaply when the next update drops.
How to evaluate it this quarter
The introductory pricing window through year-end makes this a practical moment to run a proper evaluation, not just a back-of-envelope projection. Here is the process that avoids the common failure mode of switching all traffic on benchmark evidence alone.
- Define cost-per-successful-task on your workload. Measure tokens, dollars, and task-completion rate per real agent job, not per-token list price or benchmark score.
- Route a slice, not all traffic. Send 10–20% of real agent runs to 3.7 Flash and compare against your current model side by side for one to two weeks.
- Track revision and failure rates. A model that completes coding tasks in fewer self-corrections saves whole generation cycles — watch that metric alongside token cost.
- Confirm region and compliance for EU workloads. Verify data-residency endpoint, logging scope, and AI Act obligations before agents process regulated data.
- Cap spend per agent. A cheaper model invites higher run volume. Set per-agent and per-day spend caps before promoting to full traffic, so efficiency gains are not eaten by volume growth.
- Model the January repricing. The standard rates from 1 January 2027 equal today’s 3.6 Flash prices. Forecast the impact now so there are no budget surprises at year-end.
Frequently asked questions
What is Gemini 3.7 Flash?
Gemini 3.7 Flash is Google's latest efficiency-tier model, released on 13 August 2026. It is tuned for software engineering, AI agent workflows, and multi-step planning, and is available through the Gemini API, Google AI Studio, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and Gemini Spark in 160+ countries.
How is Gemini 3.7 Flash different from Gemini 3.6 Flash?
Gemini 3.7 Flash scores 65.3% on DeepSWE, up from 49.0% for 3.6 Flash, and 43.6% on FrontierCode 1.1 production code quality, up from 34.4%. Enterprise workflow automation (AutomationBench) more than doubles to 30.4% from 17.0%. Pricing is also halved during the introductory period: $0.75 input / $3.75 output per million tokens, versus $1.50 / $7.50 for 3.6 Flash.
How much does Gemini 3.7 Flash cost?
$0.75 per million input tokens and $3.75 per million output tokens through 31 December 2026. From 1 January 2027, pricing rises to $1.50 input and $7.50 output per million tokens, matching the current 3.6 Flash standard rates.
Is Gemini 3.7 Flash available in the EU?
Yes, through the Gemini API, Gemini Enterprise Agent Platform, and Vertex AI. EU teams handling personal or regulated data should confirm data residency endpoint, logging configuration, and compliance with GDPR and the EU AI Act before using the model in production agentic pipelines.
Should we switch coding agents from 3.6 Flash to 3.7 Flash now?
Test before switching. Route 10–20% of real agent traffic to 3.7 Flash, measure cost per completed task, revision rate, and task-success rate, and promote it only where quality holds on your actual codebase and tooling. Benchmark scores are directional; the savings are real only on your specific workloads.
Sources
Google — Introducing Gemini 3.7 Flash: our most intelligent workhorse model (primary source), 13 August 2026
Bloomberg — Google Unveils Gemini 3.7 Flash Model as Gemini 3.5 Pro Delay Persists, 13 August 2026
Axios — Google’s Gemini 3.7 Flash arrives before Gemini 3.5 Pro, 13 August 2026
SiliconAngle — Google launches Gemini 3.7 Flash for coding, AI agent projects, 13 August 2026