AI Policy Wiki
Dashboard

Gemini 3 / Gemini 3 Pro

medium confidence · updated 2026-08-14

Google DeepMind's frontier model family (launched November 2025) — briefly the benchmark leader at launch, lost ground at the coding-agent frontier, refreshed as Gemini 3.1 Pro (February 2026), then extended through the Gemini 3.5 series from May 2026 and the Flash point-releases running to Gemini 3.7 Flash (August 2026).

Gemini 3 is Google DeepMind's frontier model family, launched November 18, 2025. At launch it was widely described as the benchmark leader; within roughly two months commentators argued it had little effect at the coding-agent frontier, and Google subsequently shipped a Gemini 3.1 Pro point-update in 2026.

FieldValue
Developer[[google-deepmindGoogle DeepMind]]
ReleasedNovember 18, 2025
Model familyGemini 3.x
TypeFrontier general-purpose
VariantsGemini 3 Flash (fast), Gemini 3 Thinking (reasoning), Gemini 3 Pro (frontier), Gemini 3 Deep Think (hard problems)
Latest point-releaseGemini 3.7 Flash (August 13, 2026)
Companion productGoogle Antigravity (coding agent)

Launch and family

Google positioned the November 18, 2025 Gemini 3 launch as a step-change "new era of intelligence," with Gemini 3 Pro claimed to outperform prior models across reasoning, multimodality, and coding benchmarks, and a Gemini 3 Deep Think mode aimed at the hardest problems (Source: blog.google). The family spans Gemini 3 Flash (fast), Gemini 3 Thinking (reasoning), Gemini 3 Pro (frontier), and Gemini 3 Deep Think, paired with the Antigravity coding agent. Antigravity features an "Inbox" concept in which agents work asynchronously and ping the user when they need help.

Capabilities and benchmarks

At launch in November 2025 Gemini 3 was described as taking "a definitive benchmark lead," accompanied by a sense that "Google is back in the lead"; Kevin Roose described it as a potential "crown" reclaiming (Source: oneusefulthing.org) (Source: interconnects.ai).

Reported demonstrations included a research task in which Mollick gave Gemini 3 messy research data and asked it to "write an original paper": the model recovered corrupted data, generated hypotheses, tested them statistically, created an NLP-based novelty measure, and produced a 14-page paper (Source: oneusefulthing.org). In a multimodal demonstration, the model built a playable game from a screenshot of a 3-year-old tweet.

Mollick characterized the model's output as "PhD-level intelligence" in the sense of competent grad student work, while noting "human-like weaknesses — some statistical methods need work, theorizing goes too far, errors are more subtle" (Source: oneusefulthing.org).

Coverage after launch was more mixed at the coding-agent frontier. One assessment held that Gemini 3 was "hailed as a false king," arguing that within two months of its November 2025 "coronation" it had "effectively no impact at the frontier of coding agents" (Source: interconnects.ai). Separately, Gemini's chatbot website was described as "much less capable" than ChatGPT or Claude.ai in terms of harness and tool capabilities, despite comparable underlying model quality (Source: A Guide to Which AI to Use in the Agentic Era). By March 2026, the Stanford HAI AI Index 2026 reported Anthropic's Claude Opus 4.6 leading by 2.7% overall (Source: Stanford HAI AI Index Report 2026).

2026 refresh: Gemini 3.1 Pro

Google released Gemini 3.1 Pro on February 19, 2026, presenting it as "a smarter model for your most complex tasks" and following a Gemini 3 Deep Think update the prior week (Source: blog.google). Google DeepMind says the point-update "raises the bar across a wide range of benchmarks," reporting figures such as 80.6% on a coding-harness measure (its best self-reported harness) and a 94.3% GPQA result described in third-party tracking as among the highest of any mainstream model on that graduate-level science-and-reasoning test (Source: deepmind.google). Google presented the refresh as its response to the post-launch coding-agent-frontier criticism. Benchmark leadership at this tier is fast-decaying and contested, and these are vendor-reported, point-in-time figures.

In April 2026 Google introduced Deep Research and Deep Research Max, agents powered by Gemini 3.1 Pro that combine web search with proprietary enterprise data (Source: venturebeat.com).

Gemini 3.5 series

At Google I/O on May 19, 2026, Google introduced Gemini 3.5 as "our latest family of models combining frontier intelligence with action," beginning with Gemini 3.5 Flash (Source: blog.google). Google reported 3.5 Flash outperforming Gemini 3.1 Pro on agentic and coding benchmarks — 76.2% on Terminal-Bench 2.1, a 1656 Elo on GDPval-AA, 83.6% on MCP Atlas, and 84.2% on CharXiv Reasoning — at roughly four times the output speed of other frontier models, and made it the default model in the Gemini app and AI Mode in Search globally. The launch introduced Gemini Spark, a personal AI agent running on 3.5 Flash, rolling out first to trusted testers and then to Google AI Ultra subscribers in the US. Google stated the 3.5 series was developed under its Frontier Safety Framework with strengthened cyber and CBRN safeguards and interpretability-based safety checks (Source: blog.google).

Gemini 3.5 Pro was described at I/O as in internal use and expected "next month"; July 2026 reporting placed its target release at July 17 (Source: techtimes.com). These vendor-reported figures and dates are point-in-time claims.

On July 21, 2026, Google released Gemini 3.6 Flash — priced at $1.50/$7.50 per million input/output tokens and reported to use 17% fewer output tokens than 3.5 Flash — alongside Gemini 3.5 Flash-Lite, which runs about 350 output tokens per second at $0.30 per million input tokens, and Gemini 3.5 Flash Cyber, a cyber-focused variant available only to governments and trusted partners. In the same announcement Google said Gemini 3.5 Pro was in partner testing and that "our most ambitious pre-training run yet, for Gemini 4" had begun (Source: blog.google; techcrunch.com). On July 23, 2026, current and former DeepMind employees told Axios the 3.5 Pro slippage stemmed partly from morale problems and departures, including Noam Shazeer's move to OpenAI and John Jumper's to Anthropic, with some resignations tied to Google's April 2026 Pentagon deal (Source: axios.com).

Gemini 3.7 Flash

Google released Gemini 3.7 Flash on August 13, 2026, three weeks after Gemini 3.6 Flash, describing it as "our most intelligent workhorse model yet for coding and agents" and attributing the short interval to developer feedback and algorithmic changes it said would carry into later models (Source: blog.google). The announcement is authored by Tulsee Doshi, senior director of product management, on behalf of the Gemini team.

Google reported the following developer-assembled comparisons against Gemini 3.6 Flash. All figures are vendor-reported and point-in-time.

BenchmarkMeasures3.7 Flash3.6 Flash
FrontierCode 1.1 MainProduction-ready code generation43.6%34.4%
DeepSWE v1.1Long-horizon software engineering65.3%49.0%
WebDev Arena (Elo)Web development15881538
GDP.pdfExpert document comprehension34.0%22.0%
AutomationBenchEnterprise workflow automation30.4%17.0%

Google also reported higher first-pass code accuracy, more functional layouts and feature-complete apps in fewer prompts for web development, and design adherence to a reference input supplied as a screenshot, image, or full design system. On developer experience it said the model "better adapts to roadblocks, clarifies intent when needed, and follows instructions with greater fidelity," putting more effort into multi-step planning and tool calls, with the stated result of less manual oversight and fewer retries (Source: blog.google).

Introductory pricing is $0.75 per million input tokens and $3.75 per million output tokens — half the 3.6 Flash rate — and expires December 31, 2026, after which the model reverts to the 3.6 Flash pricing of $1.50/$7.50 per million tokens (Source: blog.google). Availability spans Google Antigravity, the Gemini API through Google AI Studio and Android Studio, the Gemini Enterprise Agent Platform and Gemini Enterprise app, and the consumer Gemini app via Spark. Gemini Spark moved to 3.7 Flash on release day; the announcement describes Spark as available to Google AI Pro and Ultra subscribers in more than 160 countries, a wider footprint than the trusted-tester and US-Ultra rollout described at its I/O launch.

On safety, Google said 3.7 Flash ships with updated safeguards against misuse in chemical, biological, radiological and nuclear domains and in cyber offense, referencing its Frontier Safety safeguards, its bioresilience approach, and the cyber programme introduced with Gemini 3.5 Flash Cyber, and published a model card for the release (Source: blog.google; deepmind.google). The announcement names no third-party evaluator and states no capability threshold or evaluation result.

Safety and evaluations

A Cisco AI Threat Research 15-model study published May 27, 2026 found that Gemini 3 Pro's multi-turn attack-success rate rose from 18.1% (single-turn) to 73.4% (multi-turn), a gap the study said was not captured by existing model cards. By comparison, xAI's Grok 4.1 Fast (non-reasoning) rose to 88.3% and Anthropic's Claude Opus 4.6 stayed lowest at 16.2% (Source: siliconangle.com). The finding places Gemini 3 Pro in the middle of the closed-frontier group on multi-turn jailbreak resistance, and connects to the AI and Cybersecurity and NIST AI Risk Management Framework 1.0 discussion of measurement gaps.

Sources

Key sources for this page: (Source: oneusefulthing.org), A Guide to Which AI to Use in the Agentic Era, (Source: interconnects.ai), Stanford HAI AI Index Report 2026, the 2026 Google DeepMind product page (Source: deepmind.google), the Cisco multi-turn-attack study (Source: siliconangle.com), and Google's Gemini 3.1 Pro and Gemini 3.5 launch posts (Source: blog.google; blog.google).

Relationships