Gemini 3.7 Flash: Release Date, Specs & Pricing (2026)

Gemini 3.7 Flash is Google's best available model as of August 2026. Intro pricing, independent benchmarks, and whether you should switch.

Quick answer. Google released Gemini 3.7 Flash on August 13, 2026. It is still a stable, fully supported model with no announced retirement date, a 1,048,576-token context window, and pricing of $0.75 input / $3.75 output per 1M tokens through December 31, 2026. Gemini 3.8 Flash superseded it on September 2, 2026 at identical pricing.

Update · September 3, 2026

Gemini 3.8 Flash is now out. Google shipped gemini-3.8-flash on September 2, 2026 at exactly the same price as 3.7 Flash — $0.75 / $3.75 per 1M tokens, same 1M context, same 64K output ceiling, same modalities. If you want Google's current Flash model rather than the 3.7 reference, start here:

Gemini 3.8 Flash: the complete guide  ·  Gemini 3.8 Flash vs 3.7 Flash: the head-to-head

This page stays the reference for Gemini 3.7 Flash itself — its release date, verified specs, pricing and where it now sits in the lineup.

Gemini 3.7 Flash was Google's headline Flash release of August 2026, and for three weeks it was the most capable model Google served. That window closed on September 2 when Gemini 3.8 Flash went generally available. This page is the 3.7 Flash record: what it is, when exactly it shipped, what it costs today, whether it is going away, and the one pricing behaviour that catches people out.

Every number below was checked against a primary source on September 3, 2026 — Google's Gemini API changelog, pricing page, model list and deprecations page, the Google DeepMind Flash model page, Artificial Analysis, and the live OpenRouter listing. Anything that could not be verified today was cut rather than hedged.

When was Gemini 3.7 Flash released?

August 13, 2026. Google's own API changelog dates the general-availability launch of gemini-3.7-flash to that day, describing "substantial improvements across software engineering, web development, and agentic workflows" and confirming introductory pricing through December 31, 2026. The OpenRouter listing for google/gemini-3.7-flash carries a creation timestamp of 2026-08-13 17:03 UTC, which matches same-day third-party availability.

It went straight to GA — there was no -preview stage for this one, unlike the Gemini 3.1 Pro line which is still preview-only.

DateEventSource
Jul 21, 2026Gemini 3.6 Flash and 3.5 Flash-Lite appearOpenRouter creation timestamps
Aug 13, 2026Gemini 3.7 Flash GA, intro pricing announcedGoogle API changelog
Sep 1, 2026Agentic video understanding added to 3.7 Flash (and 3.6 Flash, 3.5 Flash-Lite) — up to 88% fewer tokens on long-form videoGoogle API changelog
Sep 2, 2026Gemini 3.8 Flash GA — supersedes 3.7 as Google's flagship FlashGoogle API changelog
Dec 31, 2026Introductory pricing ends for 3.7 FlashGoogle pricing page
Jan 1, 2027Standard pricing takes effect: $1.50 / $7.50 per 1MGoogle pricing page

Note the September 1 entry. Google added a genuinely useful capability to 3.7 Flash the day before launching its replacement — agentic video understanding, which Google says cuts token usage by up to 88% on long-form content. That is not the behaviour of a model being quietly wound down.

Is Gemini 3.7 Flash still available, or has it been deprecated?

Still available, still stable, and there is no published retirement date.

Two checks confirm this as of September 3, 2026. First, gemini-3.7-flash is listed on Google's Gemini API model page with a status of Stable — not preview, not legacy, not deprecated. Second, Google's deprecations page, which does carry explicit shutdown dates for a long list of models (gemini-3.1-flash-lite retires May 7, 2027; gemini-2.5-flash-image on October 2, 2026; gemini-omni-flash-preview on September 30, 2026), lists no shutdown date for gemini-3.7-flash. The same is true of 3.6 Flash and 3.5 Flash.

Google's own precedent here is worth knowing when you plan a migration window: the shutdown notices it does publish typically give around 6 to 12 months of runway. Since no notice exists for 3.7 Flash yet, the practical read is that you have at least that long from whenever a notice appears — but pin your model string to gemini-3.7-flash explicitly rather than relying on any alias, and watch the deprecations page rather than assuming.

Superseded is not the same as deprecated. 3.7 Flash is superseded by 3.8 Flash. It is not deprecated by anyone.

What are Gemini 3.7 Flash's specifications?

These come from Google's model documentation page for gemini-3.7-flash, cross-checked against the OpenRouter API listing.

AttributeValue
Model ID (Gemini API / Vertex)gemini-3.7-flash
Model ID (OpenRouter)google/gemini-3.7-flash
StatusStable / GA
Input token limit1,048,576 (1M)
Output token limit65,536 (64K)
Input modalitiesText, image, video, audio, PDF
Output modalitiesText only
Thinking modesLow, medium, high
SupportedContext caching, code execution, computer use (preview), file search, function calling, Google Search grounding, Google Maps grounding, structured outputs
Not supportedImage generation, audio generation, Live API
Serving tiersStandard, Batch API, Flex, Priority

One correction to a claim that circulated widely at launch: the output ceiling is 65,536 tokens, not 64,000. And there is no 2M-token Gemini context window in production — every current Gemini 3.x model, 3.7 Flash and 3.8 Flash included, tops out at 1,048,576 input tokens.

How much does Gemini 3.7 Flash cost in 2026?

Straight from Google's pricing page. All figures are per 1 million tokens, and the January 2027 column is not a rumour — it is published on the same page as the current rates.

TierInput (to Dec 31, 2026)Output (to Dec 31, 2026)Input (from Jan 1, 2027)Output (from Jan 1, 2027)
Standard$0.75$3.75$1.50$7.50
Batch API$0.375$1.875$0.75$3.75
Priority$1.35$6.75$2.70$13.50
Cached input (read)$0.075$0.15
Cache storage$0.50 per 1M tokens / hour$1.00 per 1M tokens / hour
Free tier (AI Studio)FreeFreeRate-limited; limits shown in your AI Studio dashboard

Three things fall out of that table that are worth acting on:

  • Batch is half price. Anything that does not need a synchronous response — nightly enrichment, backfills, bulk classification, document extraction — belongs on the Batch API at $0.375 / $1.875. That is a 50% cut for a scheduling change.
  • Cached reads are 10x cheaper than fresh input. $0.075 versus $0.75. If you send the same system prompt, tool schema or document corpus on every call, context caching is the single largest lever available. Just account for the $0.50/1M/hour storage charge — caching only pays if the cache is actually hit often enough to beat that clock.
  • Everything doubles on January 1, 2027. Standard, Batch, Priority and caching all go to exactly 2x. If you are modelling 2027 spend, model it at $1.50 / $7.50.

Why does Gemini 3.7 Flash cost more than the sticker price suggests?

Because of reasoning tokens, and this is the cost trap that surprises teams migrating from a non-thinking model.

Gemini 3.7 Flash is a thinking model with low, medium and high thinking modes. The tokens it generates while reasoning internally — tokens you never see in the response body — are billed at the output rate. Google's pricing page states this plainly: output pricing is "including thinking tokens." OpenRouter's API exposes the same thing as a distinct internal_reasoning price of $3.75 per 1M, identical to the completion price.

The practical consequence: a request that returns a 400-token answer may bill for several thousand output tokens. Your effective cost per task is driven by thinking volume, not response length, and a naive "average response is short, so this will be cheap" estimate can be off by an order of magnitude.

Two defences. First, use the thinking-mode control — drop to low thinking for extraction, formatting, routing and classification work that does not need deliberation. Second, benchmark on cost-per-completed-task using real traffic, never on the headline per-token rate. Artificial Analysis found 3.7 Flash burned 64 million output tokens running its full Intelligence Index suite, at a total cost of $484.73. That is the number shape to reason about.

How good is Gemini 3.7 Flash? Benchmarks, with provenance

Vendor benchmarks and independent benchmarks are different evidence classes and should never be mixed into one table. Here they are separately.

Independent: Artificial Analysis

Artificial Analysis scores Gemini 3.7 Flash (high thinking) at 56 on its Intelligence Index, currently ranking #28 of 196 models tracked. Measured output speed is 279.4 tokens per second with a time-to-first-token of 12.01 seconds at high thinking — fast generation, slow start, which is exactly the signature of a model doing a lot of upfront reasoning. Blended price works out to $0.58 per 1M tokens on AA's 7:2:1 cache-hit / input / output assumption.

For reference, AA scores Gemini 3.8 Flash (high) at 59, ranking #17. Three index points is a real but not enormous gap.

Vendor self-reported: Google DeepMind

Google's own Flash model page publishes a head-to-head table. These numbers are Google-selected and Google-run — read them as a best case, not neutral ground truth. The 3.7 Flash column is what matters here.

Benchmark [Google self-reported]3.7 Flash3.8 FlashClaude Opus 5GPT-5.6 TerraClaude Sonnet 5
Vals Finance Agent v259.0%61.4%58.6%54.4%53.9%
Harvey's Legal Agent Benchmark8.8%10.0%6.7%0.8%5.0%
HLE-Verified53.6%54.9%54.4%51.1%31.0%

The honest reading of Google's own table: on the agentic knowledge-work benchmarks Google chose to publish, 3.7 Flash already beat Claude Opus 5, GPT-5.6 Terra and Claude Sonnet 5, and 3.8 Flash improves on it by 1.2 to 2.4 points. That is an incremental generation, not a step change. Anyone telling you 3.7 Flash is now obsolete is overselling the gap.

Production signal: OpenRouter

OpenRouter's live stats for google/gemini-3.7-flash show best-provider throughput of 197 tokens/second, best P50 latency of 1.40 seconds, and 99.79% availability over a rolling three-day window. Top consuming applications include Hermes Agent (632B tokens) and Portkey AI (544B tokens) — real agentic production volume, not demo traffic.

How does Gemini 3.7 Flash actually differ from 3.8 Flash?

This is the surprising part: on paper, they are the same model. Every published specification is identical.

AttributeGemini 3.7 FlashGemini 3.8 Flash
ReleasedAug 13, 2026Sep 2, 2026
Standard price / 1M (to Dec 31, 2026)$0.75 / $3.75$0.75 / $3.75
Cached input / 1M$0.075$0.075
Batch / 1M$0.375 / $1.875$0.375 / $1.875
Context window1,048,5761,048,576
Max output65,53665,536
Input modalitiesText, image, video, audio, PDFText, image, video, audio, PDF
Thinking modesLow / medium / highLow / medium / high
AA Intelligence Index56 (#28)59 (#17)
Output tokens to run AA suite64M120M
Cost to run AA suite$484.73$825.83

Price, context, output ceiling, modalities and the tuneable parameter surface are byte-for-byte the same. The entire difference is answer quality and verbosity.

And verbosity is where the interesting cost story lives. Because the per-token price is identical, but 3.8 Flash spent 120 million output tokens running Artificial Analysis's Intelligence Index against 3.7 Flash's 64 million, 3.8 Flash cost roughly 70% more to run the same evaluation suite ($825.83 vs $484.73). "Same price" per token does not mean same bill. If you are a high-volume shop with a workload 3.7 Flash already handles, the newer model can quietly cost you more per task for a three-point index gain.

That trade-off deserves its own analysis rather than a paragraph — we break it down properly in Gemini 3.8 Flash vs 3.7 Flash, and cover what the newer model actually adds in the Gemini 3.8 Flash complete guide.

How do you access Gemini 3.7 Flash?

Four routes, all live today.

  1. Google AI Studio — free tier for prototyping. Google no longer publishes flat rate-limit tables; your actual limits appear in the AI Studio dashboard.
  2. Gemini API — model ID gemini-3.7-flash.
  3. Vertex AI / Gemini Enterprise Agent Platform — same model ID, with enterprise controls, regional routing and the Batch, Flex and Priority consumption tiers.
  4. OpenRoutergoogle/gemini-3.7-flash, plus a google/gemini-3.7-flash:batch variant at the batch rate. Useful if you want one API surface across vendors.
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{"parts": [{"text": "Summarise this contract in 5 bullets."}]}],
    "generationConfig": {"thinkingConfig": {"thinkingLevel": "low"}}
  }'

Setting thinkingLevel to low on work that does not need deliberation is the cheapest single optimisation available on this model, for the reasons in the reasoning-token section above.

Should you stay on 3.7 Flash or move to 3.8?

A decision rule rather than a recommendation, because the right answer genuinely depends on your workload.

  • Stay on 3.7 Flash if your evals already pass, your volume is high, and cost-per-task matters more than the last few points of quality. It is stable, has no retirement date, and is measurably less verbose. Switching costs you nothing in price per token but may cost you real money in tokens consumed.
  • Move to 3.8 Flash if you run long-horizon agentic loops, multi-step software engineering, or complex enterprise workflows — the categories Google explicitly engineered it for, and where the DeepSWE and agent-benchmark gains show up. Same sticker price, so there is no budget objection to testing it.
  • Move off both if you are still on Gemini 3.5 Flash or older. Those are the versions actually leaving value on the table.
  • Do the migration properly: run your own eval set against both, measure total output tokens per completed task rather than per-token price, and pin explicit model strings in production.

For the wider Gemini family picture — how the Flash tier relates to Pro, and where the long-delayed Pro flagship stands — see our Gemini 3.5 complete guide and the running Gemini 3.5 Pro release tracker. As of September 3, 2026 there is still no gemini-3.5-pro model ID in Google's published API model list or on OpenRouter.

FAQ

When was Gemini 3.7 Flash released?

August 13, 2026. Google's Gemini API changelog dates the general-availability launch of gemini-3.7-flash to that day, announcing improvements across software engineering, web development and agentic workflows, along with introductory pricing running through December 31, 2026. It launched straight to GA with no preview stage, and third-party providers including OpenRouter listed it the same day.

Is Gemini 3.7 Flash still available?

Yes. As of September 3, 2026 gemini-3.7-flash is listed as Stable in Google's Gemini API model documentation and is served through AI Studio, the Gemini API, Vertex AI and OpenRouter. Google even added agentic video understanding to it on September 1, 2026. It has been superseded by 3.8 Flash but is not being withdrawn.

Should I use Gemini 3.7 or 3.8 Flash?

Both cost $0.75 / $3.75 per 1M tokens, so price is not the tiebreaker. Choose 3.8 Flash for long-horizon agentic work and complex software engineering, where it scores 59 versus 56 on the Artificial Analysis Intelligence Index. Choose 3.7 Flash for high-volume, latency-sensitive work: it is materially less verbose, so cost per completed task is often lower.

How much does Gemini 3.7 Flash cost?

$0.75 per 1M input tokens and $3.75 per 1M output tokens on the standard tier through December 31, 2026, rising to $1.50 / $7.50 on January 1, 2027. Batch is half price at $0.375 / $1.875, Priority is $1.35 / $6.75, and cached input reads cost $0.075 per 1M plus $0.50 per 1M tokens per hour of cache storage. A free tier exists in Google AI Studio.

Is Gemini 3.7 Flash deprecated?

No. Google's deprecations page publishes explicit shutdown dates for many models, including gemini-3.1-flash-lite on May 7, 2027 and gemini-2.5-flash-image on October 2, 2026. It lists no shutdown date for gemini-3.7-flash. Being superseded by 3.8 Flash is not the same as being deprecated, and no retirement notice has been issued.

What is Gemini 3.7 Flash's context window?

1,048,576 input tokens, roughly 1 million, with a maximum output of 65,536 tokens. It accepts text, images, video, audio and PDF as input and returns text only. No production Gemini model currently offers a 2M-token context window — 3.8 Flash has exactly the same 1,048,576-token limit.

Does Gemini 3.7 Flash charge for thinking tokens?

Yes, and this is the most common budgeting mistake on this model. Google's pricing page states that output pricing includes thinking tokens, so internal reasoning bills at the full $3.75 per 1M output rate even though you never see those tokens in the response. Use the low thinking mode for simple tasks, and benchmark on cost per completed task rather than per-token price.

What replaced Gemini 3.7 Flash?

Gemini 3.8 Flash, released September 2, 2026, which Google describes as its most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents and complex enterprise workflows. It carries identical pricing, context window, output limit and modality support — the only differences are answer quality and verbosity.

The bottom line

Gemini 3.7 Flash shipped on August 13, 2026, is still stable and fully served, and has no retirement date. It was superseded three weeks later by Gemini 3.8 Flash at the exact same price, with the same 1M context, same 64K output ceiling and the same multimodal input surface. The only real difference is quality against verbosity: 3.8 scores three points higher on the independent Artificial Analysis index and cost about 70% more to run the same evaluation suite because it thinks more.

So the decision rule is unglamorous but correct: if 3.7 Flash passes your evals today, there is no urgency to move, and staying may be cheaper per task. If you are building long-running agents, test 3.8 — it costs nothing extra per token to try. Either way, model your 2027 budget at $1.50 / $7.50, because that increase is already published.