Gemini 3.7 Flash: Google's Best AI Model Is Now a Flash (2026)

Gemini 3.7 Flash is Google's best available model as of August 2026. Intro pricing, independent benchmarks, and whether you should switch.

Quick answer. Gemini 3.7 Flash, released August 13, 2026, is Google's best available model right now, scoring 56 on the independent Artificial Analysis Intelligence Index. Intro API pricing is $0.75 input / $3.75 output per 1M tokens through December 31, 2026, rising to $1.50/$7.50 in January. With Gemini 3.5 Pro delayed indefinitely, this Flash model is the one to run.

Google shipped Gemini 3.7 Flash on August 13, 2026, calling it "our most intelligent workhorse model yet for coding and agents." That marketing line undersells the strange position Google is in: with Gemini 3.5 Pro delayed indefinitely, 3.7 Flash is not just Google's best cheap model. As of August 2026, it is Google's best model, full stop.

This guide covers what shipped, the two-layer pricing story (including a discrepancy between Google's official rate and what OpenRouter actually charges), the benchmarks with clear provenance labels, and whether you should switch. Every number here comes from the Google launch post and model card, the OpenRouter models API (machine-read August 20, 2026), Artificial Analysis, or attributed practitioner reports.

What is Gemini 3.7 Flash?

Gemini 3.7 Flash is the third Flash release Google has shipped this summer, following 3.5 Flash (May 19) and 3.6 Flash (July 21). Core specs from the model card and OpenRouter listing:

  • Released: August 13, 2026. Model ID gemini-3.7-flash (on OpenRouter: google/gemini-3.7-flash).
  • Context window: 1M tokens input (1,048,576), 64K output.
  • Knowledge cutoff: March 2026 (some domains January 2025).
  • Modalities: text, images, audio, and video in; text only out.
  • Positioning: coding and agentic work. Google claims significantly better real-world software-engineering and agent benchmarks, with fewer failed agent loops, plus higher-fidelity web and app code from design mocks, including auditing a codebase against a mock for 1:1 design parity. Those claims are Google's own; the benchmark section below separates vendor numbers from independent ones.

Launch-day availability was broad: Google Antigravity, Google AI Studio, Android Studio, the Gemini Enterprise Agent Platform and Enterprise app, and OpenRouter same-day. The same launch expanded Gemini Spark, Google's persistent personal agent, from US-only Ultra subscribers to Google AI Pro and Ultra subscribers in 160+ countries.

Why is Google's best model a Flash model?

This is the part most launch coverage skims, and it is the real story of summer 2026.

At I/O on May 19, Google shipped Gemini 3.5 Flash and promised Gemini 3.5 Pro was "already being used internally" with a rollout "next month." June came and went. On June 25, Business Insider reported the release had slipped to July. On July 16, CNBC reported Alphabet shares fell on the delay, and Bloomberg (via 9to5Google) reported the reason: Google was "taking time to try to improve [3.5 Pro's] capabilities, particularly in coding," after late-June training runs with updated data "yielded disappointing results." Google's last official word, on July 21, was that it is "currently testing 3.5 Pro, an upgraded Flash model, and other models with partners."

The upgraded Flash models shipped: 3.6 Flash on July 21, then 3.7 Flash on August 13. The Pro did not. As of August 20, 2026, there is no gemini-3.5-pro model ID anywhere, no model card, and no launch date. Google's Pro flagship remains Gemini 3.1 Pro (preview), launched back in February, and on Google's own published numbers 3.7 Flash now outscores the Pro line on most benchmarks. The practitioner consensus, as one Hacker News commenter (re-thc) put it: "Flash is better than Pro for now."

So the honest framing for anyone picking a Gemini model in August 2026: the Flash tier is the flagship. If you have been holding out for 3.5 Pro, there is nothing to hold out for yet, and no date to plan around.

How much does Gemini 3.7 Flash cost?

Pricing is a two-layer story, and quoting it carelessly is how posts get this wrong. The canonical price is Google's official API rate. But there is an introductory-pricing clock on it, and OpenRouter is currently serving the model at half the official intro rate.

TierInput / 1MOutput / 1MNotes
Official intro (until Dec 31, 2026)$0.75$3.75The canonical citable price
Batch / Flex (intro)$0.375$1.87550% off intro rate
Priority tier$1.35$6.75 
Standard (from Jan 1, 2027)$1.50$7.502x the intro rate
OpenRouter (as of Aug 20, 2026)$0.375$1.875Half Google's official intro rate; see below

The OpenRouter discrepancy, flagged carefully: reading OpenRouter's models API on August 20, 2026, google/gemini-3.7-flash is listed at $0.375/$1.875, with a :batch variant at $0.1875/$0.9375. That matches Google's batch tier, not its standard intro rate. Whether this is a routing arrangement or a listing quirk, treat $0.75/$3.75 as the price to budget against and the OpenRouter rate as an as-of-August-2026 observation that may change without notice.

There is also a free tier in AI Studio, but Google no longer publishes flat rate-limit tables; the actual limits are only visible in your AI Studio dashboard.

The intro-pricing clock matters. 3.6 Flash launched with the same structure: intro until December 31, 2026, then $1.50/$7.50. The context is that 3.5 Flash launched in May at $1.50/$9.00, roughly a 3x hike over its predecessor, and was widely seen as mispriced. The community reads the intro pricing on 3.6 and 3.7 as a face-saving reset. One Hacker News commenter (GodelNumbering) was blunt: "They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5' ... I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)." Another (KptMarchewa) called the January reversion "added specifically so you migrate out." If you build on the intro rate, budget for the 2x jump in January 2027.

How good is Gemini 3.7 Flash? Benchmarks, with provenance

Three grades of evidence here, and they should not be mixed.

Independent: Artificial Analysis

As of August 2026, Artificial Analysis scores Gemini 3.7 Flash at 56 on its Intelligence Index, up from 52 for 3.6 Flash. That places it above Claude Sonnet 5 (55) and below GPT-5.6 Terra (57) and Muse Spark 1.2 (57). AA's differentiated finding on Gemini is throughput: it sits in the fastest tier for output speed and end-to-end response time, at the cost of high verbosity. Output tokens per AA task rose from 26k (3.6 Flash) to 37k, so part of the intelligence gain is bought with more thinking tokens. Even so, the price cut means it is cheaper per task: AA's cost to run its full suite was $485 for 3.7 Flash versus $1,068 for Grok 4.6.

Vendor self-reported: Google's model card

Google's model card compares 3.7 Flash against 3.6 Flash, Claude Sonnet 5, and GPT-5.6 Terra. These numbers are Google-selected and Google-run; read them as a vendor's best case, not neutral ground truth.

Benchmark [Google self-reported]3.7 Flash3.6 FlashSonnet 5GPT-5.6 Terra
FrontierCode 1.143.6%34.4%42.7%41.3%
DeepSWE v1.165.3%48.6%53.8%69.6%
Code Arena (web-dev Elo)1588153815411523
Terminal-Bench 2.185.8%78.0%80.4%87.4%
Terminal-Bench 3.014.9%5.4%14.6%20.8%
AutomationBench30.4%17.0%10.7%23.6%
OSWorld-2.0 (computer use)47.9%33.8%-50.2%
HLE-Verified53.6%51.2%31.0%51.1%
LVBench (long video)85.4%84.2%68.5%78.9%
GDM-MRCR v2 8-needle @128k97.0%91.8%81.5%93.5%

The pattern in Google's own table is coherent: 3.7 Flash leads on web-dev code generation, browser/app automation, long-video and long-context retrieval, and PDF comprehension, while GPT-5.6 Terra keeps the lead on the hardest agentic-coding suites (DeepSWE, Terminal-Bench 3.0, OSWorld). On knowledge-work Elo (GDPVal-AA v2), 3.7 Flash trails Sonnet 5, Terra, and especially Muse Spark 1.2 (1628 vs 1525). Google did not pick a table it sweeps, which lends it some credibility, but it remains the vendor's own harness.

Practitioner reports (anecdotal)

From the 967-point Hacker News launch thread: nickandbro called it "genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore." xnx: Google is "not currently in the lead for maximum model capability, but it is still very competitive (or even best) in the multidimensional capability, cost, and speed frontier." On the skeptical side, markasoftware noted GPT-5.6 Sol high is "almost the same speed if you take into account drastically lower token use ... Gemini is 7x faster but 5x more tokens," and fmind-dev's use-case framing rings true for agent builders: Flash is "one of the best 'good-enough' models" but "often not strong enough for heavy refactoring and long running development loops."

How does it compare to 3.5 Flash, 3.6 Flash, and the budget field?

Within Google's own family, the answer is simple: 3.7 Flash dominates. Prices below are OpenRouter listings as of August 20, 2026, with Google's official rate noted where it differs.

ModelContext$ in / $ out per 1MRead
Gemini 3.7 Flash1M0.375 / 1.875 (official intro 0.75 / 3.75)AA Index 56; the pick
Gemini 3.6 Flash1M0.75 / 3.75AA Index 52; superseded after 3 weeks
Gemini 3.5 Flash1M1.50 / 9.00Worst value in its own family now
Gemini 3.5 Flash-Lite1M0.30 / 2.50350 tok/s tier for bulk work
Gemini 3.1 Pro (preview)1M2.00 / 12.00The "flagship" that Flash now outscores
DeepSeek V4 Flash1M0.089 / 0.177The extreme-budget option
Qwen 3.8 27B1M0.45 / 3.20Open-weights alternative

vs 3.5 Flash: there is no case for staying. 3.5 Flash costs $1.50/$9.00, more than double 3.7 Flash's intro rate, for a model two generations behind. If you are still on it, switching is a config change that cuts your bill and improves quality.

vs 3.6 Flash: same intro price, four AA Index points lower, and Google's card shows big gaps on FrontierCode (34.4% vs 43.6%) and DeepSWE (48.6% vs 65.3%) [Google self-reported]. 3.6's one advantage was terseness; 3.7 spends noticeably more output tokens per task. For latency-sensitive, high-volume pipelines where verbosity is cost, test both. Everyone else should move.

vs DeepSeek V4 Flash: for text-only workloads, DeepSeek is 13-26x cheaper with comparable intelligence per practitioner reads of the AA data (euazOn on HN, anecdotal), though DeepSeek has announced a price hike. Gemini's counterargument is multimodality (image, audio, video in) and serving scale; as PunchTornado put it on HN: "did you try to ingest 1M documents per hour with any provider except GCP ... The only thing that works at scale is gemini flash." We compared the budget tier in depth in our DeepSeek V4 Flash vs Qwen 3.7 Flash piece.

vs Qwen 3.8 27B: at $0.45/$3.20 on OpenRouter it undercuts Gemini's official intro rate, and being open-weights, you can self-host it. If data residency or self-hosting matters, start with our Qwen 3.8 27B complete guide. If you want a managed multimodal API with the fastest response times in AA's measurements, Gemini 3.7 Flash is the stronger default.

How do you access Gemini 3.7 Flash?

Four routes, all live since launch day:

  1. Google AI Studio - free tier for prototyping (limits shown in your dashboard, not published as flat tables anymore).
  2. Gemini API - model ID gemini-3.7-flash:
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"contents":[{"parts":[{"text":"Audit this component against the attached mock."}]}]}'
  1. Enterprise - Gemini Enterprise Agent Platform and the Enterprise app carry it; Google's launch materials list agent-platform availability, with tunable variants reported for the Vertex side.
  2. OpenRouter - google/gemini-3.7-flash, currently at the discounted rate noted above.

In the consumer apps, the launch also expanded Gemini Spark, the persistent agent introduced at I/O, to Google AI Pro and Ultra subscribers in 160+ countries.

One migration note for terminal users: the standalone Gemini CLI stopped working for Google-auth tiers on June 18, 2026, replaced by Antigravity CLI. The open-source google-gemini/gemini-cli repo is still releasing (v0.56.0 stable, August 19), but the hosted-auth path is gone.

Should you switch to Gemini 3.7 Flash?

Switch now if you are on any older Gemini Flash: it is cheaper than 3.5 Flash, smarter than 3.6, and the change is one string. Switch too if you are paying Gemini 3.1 Pro rates ($2/$12) for work a Flash-class model handles; on Google's own numbers the Flash now outscores the Pro line on most published benchmarks.

Test first if you run long agentic coding loops. The hardest-suite numbers (DeepSWE, Terminal-Bench 3.0, OSWorld) still favor GPT-5.6 Terra even in Google's own table, and practitioners consistently describe Flash as the fast good-enough tier rather than the heavy-refactoring tier. A pattern several developers report: use a frontier model to design and review, and Gemini Flash to implement fast in the middle.

Budget honestly: the $0.75/$3.75 rate expires December 31, 2026. Model your 2027 costs at $1.50/$7.50, and treat the OpenRouter discount as a bonus while it lasts.

FAQ

Is Gemini 3.7 Flash better than Gemini 3.1 Pro?

On most of Google's published benchmarks, yes, and at a fraction of the price ($0.75/$3.75 intro vs $2/$12). The practitioner consensus as of August 2026 is that Google's Flash tier currently beats its own Pro tier. Gemini 3.1 Pro remains the official Pro flagship only because Gemini 3.5 Pro has not shipped.

When will Gemini 3.5 Pro be released?

There is no date. It was promised for June 2026 at I/O, slipped to July, and was then delayed indefinitely; Bloomberg reported the holdup is coding performance after disappointing late-June training runs. Google's last official statement (July 21, 2026) says it is testing with partners and will release "as soon as it's ready." As of August 20, 2026, no gemini-3.5-pro model ID exists.

How much does Gemini 3.7 Flash cost?

Official API intro pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens through December 31, 2026, then $1.50/$7.50 from January 1, 2027. Batch jobs are 50% off. OpenRouter currently lists it at $0.375/$1.875 as of August 2026.

What is the context window of Gemini 3.7 Flash?

1M tokens of input (1,048,576) and up to 64K output tokens. It accepts text, images, audio, and video as input, and outputs text only. No current Gemini model offers a 2M context; that figure was an expectation attached to the unreleased 3.5 Pro.

Is Gemini 3.7 Flash the best AI model from Google in 2026?

As of August 2026, yes. It scores 56 on the Artificial Analysis Intelligence Index, above Claude Sonnet 5 (55) and above every other Gemini model, including the 3.1 Pro flagship. Across the wider field it trails GPT-5.6 Terra and Muse Spark 1.2 (both 57), so it is Google's best rather than the industry's best.

Is there a free way to use Gemini 3.7 Flash?

Yes. Google AI Studio includes a free tier for Gemini 3.7 Flash. Google no longer publishes flat rate-limit tables, so check the AI Studio dashboard for your project's actual limits.

For the full Gemini family picture, including the 3.5 Flash launch and the Pro delay timeline, see our Gemini 3.5 complete guide.