Quick answer. Gemini 3.5 is Google DeepMind's mid-2026 model family — Gemini 3.5 Flash launched at Google I/O on May 19, 2026 at $1.50/$9 per 1M tokens (roughly a 3x price hike over its Flash predecessor), followed by 3.5 Flash-Lite ($0.30/$2.50) and Gemini 3.6 Flash (July 21) and Gemini 3.7 Flash (August 13), the latter two at an introductory $0.75/$3.75. Gemini 3.5 Pro never shipped — delayed indefinitely as of August 2026, with Gemini 3.1 Pro ($2/$12) still the Pro flagship. Google's best available model today is Gemini 3.7 Flash. The Gemini differentiator remains native multimodality (text + image + audio + video in a single request), 1M-token context with strong recall, and tight Google Cloud + Workspace integration; it competes head-on with Claude Sonnet 5, GPT-5.6, and DeepSeek V4. This guide covers the lineup, benchmarks, pricing, the ecosystem (AI Studio, Vertex, Gemini API, Workspace, Antigravity CLI), and how to choose.
August 2026 update: 3.6 Flash, 3.7 Flash, and the 3.5 Pro delay
The Gemini family moved fast after I/O. What has changed since this guide first published:
- Gemini 3.6 Flash (July 21, 2026). The "upgraded Flash" — better coding (DeepSWE 49% vs 37%, Google self-reported), 17% fewer output tokens, and introductory pricing of $0.75/$3.75 per 1M tokens through December 31, 2026 ($1.50/$7.50 after). Launched alongside Gemini 3.5 Flash-Lite ($0.30/$2.50, ~350 tok/s) and Gemini 3.5 Flash Cyber, a cybersecurity fine-tune in a limited government/partner pilot.
- Gemini 3.7 Flash (August 13, 2026). Google's best available model, period — "our most intelligent workhorse model yet for coding and agents," Artificial Analysis Intelligence Index 56 (vs Claude Sonnet 5's 55 and GPT-5.6 Terra's 57), 1M-token context, same intro pricing. On Google's own published benchmarks it outscores the Pro line. Full breakdown: our Gemini 3.7 Flash launch guide.
- Gemini 3.5 Pro — delayed indefinitely. The June GA never happened: the date slipped, Bloomberg and CNBC reported in July that coding performance is the blocker (Alphabet stock dipped on the news), and Google's last official line (July 21) is "as soon as it's ready." Gemini 3.1 Pro ($2/$12 per 1M tokens) remains the Pro flagship. Full timeline: Gemini 3.5 Pro delay status.
- Gemini Spark went global. As of August 13, Spark is available to Google AI Pro and Ultra subscribers in 160+ countries (it was US-only at launch).
Bottom line for August 2026: Google's best available model is Gemini 3.7 Flash ($0.75/$3.75 intro until December 31, 2026); the Pro tier is still Gemini 3.1 Pro, and Gemini 3.5 Pro remains unreleased after repeated delays.
What is Gemini 3.5 and when did it launch?
Gemini 3.5 is the May 2026 evolution of Google DeepMind's Gemini family. Two variants are shipping or imminent:
- Gemini 3.5 Flash — launched at Google I/O 2026 on May 19, 2026. Replaced Gemini 3 Flash as the default in the Gemini app, in Google Search's AI Mode globally, and in the Gemini API (model id
gemini-3.5-flash). Pricing announced: $1.50 per 1M input tokens, $9 per 1M output tokens — roughly a 3x price hike over its Flash predecessor, which launch coverage led with; Google later reset the curve with 3.6/3.7 Flash intro pricing at half that rate. - Gemini 3.5 Pro — announced at I/O for a June 2026 release that never happened. Delayed indefinitely as of August 2026 (Bloomberg: coding performance is the blocker); no model ID, no pricing, no date. Gemini 3.1 Pro ($2/$12 per 1M tokens) remains the Pro flagship.
The previous generation — Gemini 3 Flash — hit Q1 2026 with strong scores (GPQA Diamond 90.4%, Humanity's Last Exam 33.7%) and rapidly became the consumer default. Gemini 3.5 is the price/quality re-pivot on top of that base, with explicit positioning against Claude Sonnet 4.6 and GPT-5.5.
Gemini 2.x and earlier are now legacy. The migration story is: Gemini 1.5 / 1.5 Pro → Gemini 2.0 (Feb 2025) → Gemini 2.5 Pro (mid-2025) → Gemini 3 (early 2026) → Gemini 3.5 (May 2026). Each step kept the API surface stable; model id changes plus pricing updates are the main migration work.
How does Gemini 3.5 compare to Claude Opus 4.7 and GPT-5.5?
The honest 2026 picture: three frontier families that trade places on different axes, with different strengths. (The anchors below reflect the spring 2026 field — Claude Opus 4.7 and GPT-5.5; the frontier has since moved to Claude Sonnet 5 and the GPT-5.6 line, per the August update above.)
| Dimension | Gemini 3.x (3.7 Flash / 3.1 Pro) | Claude Opus 4.7 | GPT-5.5 |
|---|---|---|---|
| Coding (SWE-bench) | Strong | ~80.8% SWE-bench Verified (highest, Apr 2026) | Strong |
| Long-context | 1M tokens (no 2M Gemini model exists) | 1M tokens (extended) | ~400K tokens |
| Multimodal (text+image+audio+video) | Native, single-request | Text + image (strong) | Text + image |
| Tool use / agents | Strong, Vertex / Agent Platform integration | Computer-use agent, Agent Skills, Managed Agents | Strong tool use, function calling |
| Reasoning (extended thinking) | Yes | Yes (extended thinking) | Strong (GPT-5.6 reasoning line) |
| Pricing tier (frontier) | Mid ($0.75/$3.75 intro on 3.7 Flash; 3.1 Pro $2/$12) | Higher | Higher |
| Free-tier access | Generous (AI Studio) | Limited (Claude.ai) | Free GPT-5.5 Instant |
| Workspace / Office integration | Native Google Workspace | Via Claude for Work | Microsoft 365 Copilot |
Where each one wins in 2026:
- Choose Gemini when you need native multimodal (video + audio in one request), a 1M-token context with strong recall, generous free-tier prototyping in AI Studio, or deep Google Cloud / Workspace integration.
- Choose Claude Opus 4.7 when you need the strongest coding agent, computer-use capability, or the cleanest agent-development experience (Claude Code, Agent Skills, Managed Agents). See our Claude Opus 4.7 guide.
- Choose GPT-5.5 when you need OpenAI's reasoning models (now the GPT-5.6 line), Microsoft 365 integration, or you're already deep in the OpenAI ecosystem. See our GPT-5.5 guide.
For agent workflows, Claude leads. For multimodal and ultra-long context, Gemini leads. For pure reasoning on hard problems, OpenAI's GPT-5.6 line is the reference. Most production teams in 2026 use two or three of these via a model router rather than betting on one.
What's new in Gemini 3.5 vs Gemini 3?
The headline improvements Google demonstrated at I/O 2026:
- Repriced Flash. $1.50/$9 per 1M tokens — in fact a ~3x hike over the Flash predecessor, widely seen as mispriced; the 3.6/3.7 Flash follow-ups reset it with $0.75/$3.75 intro pricing.
- Sharper coding. Google demoed live multi-file refactors in the Gemini Code Assist extension and in the Gemini CLI (since sunset in favour of Antigravity CLI). Internal benchmarks show ~10-15 point gains on SWE-bench Verified vs 3.0 Flash.
- Better extended thinking — the reasoning chain is visible to developers (similar to Claude's extended thinking) and can be capped to budget compute via the thinking config.
- Cleaner long-context recall across the 1M-token window. (The 2M-token window expected on 3.5 Pro never materialised — the model is unreleased.)
- Native multimodal in a single request. Send an image, an audio file, and a video clip with one prompt — no separate API calls. The vision-audio-video joint reasoning is genuinely differentiated.
- Agent-platform integration. Tighter coupling with the Gemini Enterprise Agent Platform, with first-class tool definitions, retrieval, evaluation, and deployment pipelines.
How do I access Gemini 3.5?
Five surfaces, depending on what you're building:
1. Gemini app (consumer). The free Gemini at gemini.google.com and the mobile apps now run Gemini 3.5 Flash as the default. Useful for quick prototyping.
2. Google AI Studio. Free developer playground at aistudio.google.com. Generous free tier — you can prototype real agents without a credit card. The best place to evaluate Gemini before committing to API.
3. Gemini API (direct). SDK in Python, Node, Go, with REST as the universal fallback. Model ids: gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash, gemini-3.5-flash-lite. (gemini-3.5-pro does not exist — the model is unreleased.) The simplest production path — no Vertex needed.
4. Vertex AI / Gemini Enterprise Agent Platform. Enterprise hosting with IAM, VPC-SC, audit logs, fine-tuning workflows for the older Gemini variants, and the full Agent Builder pipeline. Use Vertex when you need GCP-native security and governance.
5. Antigravity CLI (formerly Gemini CLI). Google sunset the hosted Gemini CLI on June 18, 2026 for Google-auth tiers; Antigravity CLI is the successor for terminal-first development. The open-source google-gemini/gemini-cli repo is still actively releasing (v0.56.0 stable, August 19, 2026) and works against your own API key, with MCP server support.
For Workspace users: Gemini 3.5 is being rolled into Google Docs, Sheets, Gmail, Meet, and the Drive search experience. The "@Gemini" experience inside docs handles drafting, refactoring, multilingual translation, and the new "audio overview" feature for documents.
What's Gemini 3.5 actually good at?
Long-context document analysis. Drop a huge input (a codebase, a legal contract, a 500-page PDF — up to 1M tokens) and ask questions across it. No other major frontier model has this context length at usable recall. The 2026 sweet-spot use case: due diligence, code-base understanding, multi-document research synthesis.
Multimodal reasoning. Send a screenshot of a UI + a screen recording + a spec doc and ask the model to write the fix. Gemini handles all three in one request, with cross-modal attention. The video-input quality genuinely beats GPT-5.5 in our testing — Google's video training data advantage shows.
Workspace-embedded productivity. "Summarise this thread", "draft a reply", "build a sheet from this PDF" — the native Workspace surface is dramatically faster than the equivalent API workflow. For knowledge workers, this is the most-used Gemini surface.
Code review and refactor at file-tree scale. With the 1M-token context, Gemini can hold a small codebase in mind and reason across it. Code Assist's "review this PR" with full repo context is meaningfully better than Gemini 3.
Real-time multilingual translation. Gemini consistently leads on lower-resource language pairs, particularly Indian languages, Thai, Vietnamese, Indonesian, Swahili. If you're shipping product into India, Southeast Asia, or sub-Saharan Africa, Gemini's multilingual tail is the strongest of the three frontier families.
Where Gemini 3.5 lags in 2026: agentic coding (Claude leads), terminal-first dev experience (Claude Code is more polished), the hardest reasoning benchmarks (OpenAI's GPT-5.6 reasoning line still wins on competition-math and the highest tiers).
How do I structure my first Gemini 3.5 integration?
Three patterns that work in 2026:
Pattern A — Direct API with the Python SDK.
pip install google-genai
from google import genai
client = genai.Client(api_key=YOUR_API_KEY)
response = client.models.generate_content(
model="gemini-3.5-flash",
contents=["Summarise this:", open("doc.pdf", "rb").read()],
config={"thinking_config": {"thinking_budget": 8192}}
)
print(response.text)
Use for: pure backend integration, no GCP dependency.
Pattern B — Vertex AI with IAM. Same SDK, but authenticated via Application Default Credentials and routed through your Google Cloud project. Use when you need audit logs, VPC-SC, or you're already in GCP.
Pattern C — a CLI agent shell (Antigravity CLI, or the OSS gemini-cli). Google sunset the hosted Gemini CLI on June 18, 2026 in favour of Antigravity CLI; the open-source gemini-cli continues to work against your own API key. Install with npm install -g @google/gemini-cli (or via brew). Authenticate with your API key. Run in a terminal — the CLI handles tool registration, MCP server discovery, multi-turn agent loops. Use for terminal-first development, code review, ops automation. See our AI coding agents guide for how it compares to Claude Code and Cursor.
What does Gemini 3.5 cost?
Gemini 3.5 Flash pricing announced at I/O 2026: $1.50 per 1M input tokens / $9 per 1M output tokens. Image and audio inputs priced as their token-equivalent. A free tier on AI Studio remains available — Google no longer publishes flat rate-limit tables, so check the AI Studio dashboard for current limits.
Gemini 3.5 Pro pricing was never announced — the model remains unreleased as of August 2026. The Pro-tier reference is Gemini 3.1 Pro at $2/$12 per 1M tokens (≤200K context; $4/$18 above). Meanwhile the newer Flash releases undercut 3.5 Flash: Gemini 3.6 and 3.7 Flash launched at an introductory $0.75/$3.75 per 1M tokens through December 31, 2026 ($1.50/$7.50 after), and 3.5 Flash-Lite costs $0.30/$2.50.
Cached inputs: Google's prompt caching (released alongside Gemini 1.5 and improved with each generation) is the cost-control hammer for agent workloads. Cached input is charged at roughly 10% of the standard input rate (about $0.15 per 1M on 3.5 Flash). For any production agent loop that re-sends the same system prompt + tool definitions, prompt caching is the single highest-leverage optimisation.
What about Gemma 4 (the open-weight side)?
Gemini 3.5 is closed-weight, API-only. Google's open-weight family is the separate Gemma line — Gemma 4 launched April 2, 2026 with four sizes (E2B, E4B, 26B MoE, 31B dense) under Apache 2.0, 256K context, multimodal. Gemma is the right answer when you need to self-host, fine-tune locally, or run on edge devices. See our Gemma 4 complete guide.
The pattern in 2026 for teams using both: Gemma 4 for self-hosted or fine-tuned workloads, Gemini 3.5 for frontier API tasks. They share lineage but serve different deployment models.
FAQ
Is Gemini 3.5 Pro actually released yet?
No. Gemini 3.5 Flash launched at Google I/O on May 19, 2026 and is generally available, but Gemini 3.5 Pro never shipped. The June 2026 target slipped, Bloomberg reported in July that Google is holding it back over coding performance, and Google's last official line (July 21, 2026) is "as soon as it's ready." As of August 2026 there is no model ID, no pricing, and no date.
What's the context window?
Gemini 3.5 Flash — like 3.6 Flash, 3.7 Flash, and Gemini 3.1 Pro — supports up to 1M tokens. No current Gemini model offers 2M tokens; that figure was pre-launch expectation for the unreleased 3.5 Pro. Long-context recall (the "needle in a haystack" benchmark) is meaningfully improved over earlier generations.
Is Gemini 3.5 free?
Yes via Google AI Studio's free tier (Google no longer publishes flat requests-per-minute limits — current quotas are shown in the AI Studio dashboard). The Gemini consumer app and AI Mode in Google Search use the current Flash models for free by default. Paid tiers exist for API usage and for the Google AI Pro and Ultra subscriptions.
Does Gemini 3.5 support function calling and tool use?
Yes — function calling and structured outputs (with JSON schema validation) are first-class. The Gemini API exposes a tools parameter; Vertex / Agent Platform adds higher-level tool orchestration. For agent workflows, the integration with the Gemini Enterprise Agent Platform is the main 2026 surface.
Can I fine-tune Gemini 3.5?
Not the 3.5 generation directly — Google offers fine-tuning on older Gemini variants (and Gemma family) through Vertex AI. For a fine-tunable Google-lineage model, use Gemma 4. See our fine-tuning guide for the broader landscape.
How does Gemini 3.5 compare to Claude Opus 4.7 for coding?
Claude Opus 4.7 leads on SWE-bench Verified (~80.8% in April 2026 — the highest in the category) and on agentic coding workflows via Claude Code. Gemini's shipped models are competitive but lag on the hardest agentic tasks. For long-context code understanding (huge monorepo, multi-file analysis), Gemini's 1M-token context with strong recall gives it a different angle.
What about Gemini's multimodal advantage?
Genuine. Gemini accepts text, image, audio, and video in a single request with cross-modal attention. Video input quality in particular leads the field — Google's training data advantage shows. For workflows involving screen recordings, lecture videos, image+audio analysis, or any cross-modal reasoning, Gemini is the strongest choice.
What's the cheapest way to evaluate Gemini 3.5?
Google AI Studio (aistudio.google.com) — free, no credit card, full access to gemini-3.5-flash. Sufficient for production-scale prototyping. Move to the Gemini API once you exceed free quotas; move to Vertex when you need enterprise governance.
Is there a Gemini equivalent to Claude Code or Cursor?
Yes — Antigravity CLI (successor to the hosted Gemini CLI, which Google sunset on June 18, 2026), the still-active open-source gemini-cli repo (v0.56.0, August 2026), and Gemini Code Assist as an IDE extension. Antigravity CLI is the closest analog to Claude Code; Code Assist competes with Cursor for IDE-integrated workflows. See our AI coding agents guide for the full comparison.
What's coming next?
Shipped since May: Gemini 3.6 Flash (July 21, 2026) and Gemini 3.7 Flash (August 13, 2026) — the latter now Google's best model. Gemini 3.5 Pro remains delayed with no date; Google says it will ship "as soon as it's ready." On the open-weight side, Gemma 4 point releases continue.
Related guides
- Gemini 3.7 Flash launch guide — Google's current best model
- Gemini 3.5 Pro delay status — why Pro never shipped
- Claude Opus 4.7 guide — the direct frontier competitor
- GPT-5.5 guide — the other major frontier
- Gemma 4 guide — the open-weight Gemini-family sibling
- AI coding agents — Gemini CLI vs Claude Code vs Cursor
- Open-source LLMs landscape
- Fine-tuning LLMs
- Apple Silicon LLMs