Gemini 3.7 Flash Explained: Features, Pricing, Speed & Benchmarks

Gemini 3.7 Flash AI model by Google - fastest coding assistant 2026 with 340 tokens per second speed

By TechZila AI Research Desk | Last checked: August 2026 | Reading time: ~10 minutes
Full disclosure: TechZila runs on reader support. Some links might earn us a commission — no extra cost to you. We keep our editorial decisions independent, always.

Quick Take

  • Release: Gemini 3.7 Flash launched August 13, 2026 — just 23 days after 3.6 Flash.
  • Pricing: $0.75/1M input tokens, $3.75/1M output tokens through Dec 31, 2026 (50% off), then doubles to $1.50/$7.50 from Jan 1, 2027.
  • Speed: 340 tokens/second output — fastest among 186 tested models, nearly 3x faster than GPT-5.6 Terra.
  • Context: 1 million token input window, 65,536 token output limit.
  • Key Gains: +9.2 points on FrontierCode (43.6% vs 34.4%), +16.3 points on DeepSWE (65.3% vs 49.0%).
  • Best For: Coding, agentic workflows, long-horizon tasks, multimodal processing.

Google did something unusual in the AI race: it shipped a meaningfully better model without waiting for a new training run. Gemini 3.7 Flash went live on August 13, 2026 — only 23 days after Gemini 3.6 Flash — and the pitch is straightforward: better almost everywhere, at half the launch price.

August 2026: What's New

Gemini 3.7 Flash is Google's most intelligent Flash model to date, built for complex coding, agentic workflows, and reliable multi-step execution. It supports tunable thinking levels (low, medium, high), accepts text/image/video/audio/PDF inputs, and delivers text-only outputs with a 1M token context window.

Gemini 3.7 Flash: Key Numbers

1M
Input tokens (context window)
65,536
Output token limit
340 tok/s
Output speed (tokens/second)
56
Intelligence Index score
$0.75
Per 1M input tokens (through 2026)
$3.75
Per 1M output tokens (through 2026)

What Is Gemini 3.7 Flash?

Gemini 3.7 Flash is Google's latest Flash-tier model, optimized for coding, agentic workflows, and multimodal processing. It's the third Flash model in 2026, following 3.5 Flash-Lite (July) and 3.6 Flash (July 21). Unlike typical model releases that trade cost for quality, 3.7 Flash improves on both fronts — at least temporarily.

The model accepts text, images, video, audio, and PDF inputs, with a massive 1,048,576-token input window and 65,536-token output ceiling. It supports configurable thinking levels (low, medium, high), letting developers tune latency versus intelligence based on task requirements.

SpecificationValue
Release DateAugust 13, 2026
Context Window1,048,576 input tokens
Output Limit65,536 tokens
Input ModalitiesText, Image, Video, Audio, PDF
Output ModalityText only
Thinking LevelsLow, Medium (default), High
Speed~340 output tokens/second
Intelligence Index56 (Artificial Analysis)

Pricing: 50% Discount (But There's a Catch)

Here's where things get interesting. Gemini 3.7 Flash launched at $0.75 per million input tokens and $3.75 per million output tokens — exactly half of Gemini 3.6 Flash's original $1.50/$7.50 pricing.

Important: This introductory rate is valid through December 31, 2026. From January 1, 2027, pricing doubles to $1.50 input and $7.50 output per million tokens.

But here's the twist: on the same day 3.7 Flash launched, Google cut 3.6 Flash to the same $0.75/$3.75 introductory rate. So both models cost identical through 2026 — making the "50% off" claim technically accurate but context-dependent.

ModelThrough Dec 31, 2026From Jan 1, 2027
Gemini 3.7 Flash$0.75 input / $3.75 output$1.50 input / $7.50 output
Gemini 3.6 Flash$0.75 input / $3.75 output$1.50 input / $7.50 output
Gemini 3.6 Flash (original)N/A$1.50 input / $7.50 output

Additional pricing tiers include cached input at $0.075/1M tokens, cache storage at $0.50/1M tokens/hour, and batch inference at $0.375/$1.875 (input/output) through 2026.

Speed: 340 Tokens/Second

Speed is where Gemini 3.7 Flash genuinely stands out. Independent benchmark firm Artificial Analysis measured the model at 340.1 output tokens per second — ranking first among 186 tested models for generation speed.

For context, that's nearly 3x the output speed of GPT-5.6 Terra and GLM-5.2, which hover around 100-120 tokens/second. This makes 3.7 Flash ideal for latency-critical applications like real-time chat, incident response pipelines, and interactive coding assistants.

Real-World Impact: At 340 tok/s, a 10,000-token response takes ~29 seconds to generate. Same output on GPT-5.6 Terra would take ~90 seconds — a 3x difference in user-facing latency.

But there's a tradeoff: 3.7 Flash uses about 40% more output tokens than 3.6 Flash to complete the same tasks, averaging around 37,000 tokens per task. This puts it in the same range as Claude Fable 5 and Alibaba's Qwen3.8 Max, but it means actual cost-per-task may be higher than raw token pricing suggests.

Benchmarks: Coding & Agent Gains

Google's own benchmarks show significant improvements over 3.6 Flash, particularly in coding and agentic workflows:

Benchmark3.7 Flash3.6 FlashImprovement
FrontierCode 1.1 Main (Production code quality)43.6%34.4%+9.2 points
DeepSWE v1.1 (Long-horizon software engineering)65.3%49.0%+16.3 points
Code Arena — Web Development (Elo rating)15881538+50 Elo
Terminal-Bench 2.1 (Terminal tasks)85.8%78.0%+7.8 points
AutomationBench (Business workflows)30.4%17.0%+13.4 points
GDP.pdf (Document processing)34.0%22.0%+12.0 points
OSWorld-2.0 (OS automation)47.9%33.8%+14.1 points
128K Long Context97.0%91.8%+5.2 points

The biggest gains are in long-horizon tasks (DeepSWE: +16.3 points) and business workflow automation (AutomationBench: +13.4 points). This suggests 3.7 Flash is particularly well-suited for multi-step coding projects and agentic workflows that require sustained reasoning.

Minor Regression: On CharXiv Reasoning (chart comprehension), 3.7 Flash actually scores 84.5% without tools, down from 85.2% for 3.6 Flash — a small but notable dip in one specific capability.

Key Features & Capabilities

Gemini 3.7 Flash supports the full suite of Gemini API features, including:

Supported Capabilities

  • Thinking modes: Low, Medium, High (minimal not supported)
  • Caching: Supported for cost optimization
  • Code execution: Supported
  • Computer use: Supported (Preview)
  • File search: Supported
  • Function calling: Supported
  • Grounding with Google Maps: Supported
  • Search grounding: Supported
  • Structured outputs: Supported
  • URL context: Supported
  • Batch API: Supported
  • Flex inference: Supported
  • Priority inference: Supported

Not Supported

  • Audio generation: Not supported
  • Image generation: Not supported
  • Live API: Not supported

Input modalities include text, images (up to 3,000 per prompt), video (up to 45 minutes with audio, 1 hour without), audio (up to 8.4 hours or 1M tokens), and PDFs (up to 3,000 pages per file).

vs GPT-5.6 Terra & Claude Sonnet 5

On Artificial Analysis's composite Intelligence Index, all three models land within two points:

ModelIntelligence IndexSpeed (tok/s)Price (through 2026)
GPT-5.6 Terra57~110~$2.50/$10.00
Gemini 3.7 Flash56340$0.75/$3.75
Claude Sonnet 555~100~$3.00/$15.00

Gemini 3.7 Flash trades a single point of intelligence for 3x the speed and roughly 1/3 the blended cost of competitors. For coding and agentic workflows where speed and cost matter more than marginal intelligence gains, this is a compelling tradeoff.

Should You Switch from 3.6 Flash?

Quick Decision Guide

Switch to 3.7 Flash if: You need better coding performance (especially long-horizon tasks), faster response times, improved document processing, or agentic workflow capabilities. The identical pricing through 2026 makes this a no-brainer for new projects.

Stick with 3.6 Flash if: You're already in production with stable results, your tasks don't benefit from the coding/agent improvements, or you're sensitive to the 40% higher token usage per task. Both models cost the same through December, so there's no rush to migrate.

TechZila Analysis

The Core Truth

Gemini 3.7 Flash is Google's clearest statement yet: the Flash tier is for coding and agents, not just cheap inference. The 9-16 point benchmark gains in coding and automation aren't incremental — they're meaningful enough to change which tasks you'd trust the model with.

The Hidden Context

Releasing 3.7 Flash just 23 days after 3.6 Flash — and cutting 3.6 to the same price — suggests Google is racing to establish Flash as the default for coding before competitors catch up. The temporary 50% discount is a classic land-grab tactic: get developers hooked now, raise prices later.

TechZila Opinion

The 340 tok/s speed is the real differentiator here. Most benchmarks focus on intelligence, but in production, latency often matters more than marginal quality gains. For coding assistants, real-time chat, and interactive agents, 3.7 Flash's speed advantage is more valuable than a single point on an intelligence index. Just budget for the 40% higher token usage per task.

Frequently Asked Questions

When was Gemini 3.7 Flash released?
Gemini 3.7 Flash launched on August 13, 2026 — just 23 days after Gemini 3.6 Flash.

How much does Gemini 3.7 Flash cost?
Through December 31, 2026: $0.75 per million input tokens and $3.75 per million output tokens. From January 1, 2027: $1.50 input and $7.50 output (pricing doubles).

What is the context window for Gemini 3.7 Flash?
1,048,576 input tokens (1M context) with a 65,536 token output limit.

How fast is Gemini 3.7 Flash?
Approximately 340 output tokens per second — fastest among 186 models tested by Artificial Analysis, nearly 3x faster than GPT-5.6 Terra.

What thinking levels does 3.7 Flash support?
Low, Medium (default), and High. The "minimal" thinking level is not supported and returns an error.

Is Gemini 3.7 Flash better than 3.6 Flash?
Yes, particularly for coding (+9.2 points on FrontierCode), long-horizon software engineering (+16.3 points on DeepSWE), and business workflows (+13.4 points on AutomationBench). Both models cost the same through 2026.

What input types does 3.7 Flash accept?
Text, images (up to 3,000), video (up to 45 min with audio), audio (up to 8.4 hours), and PDFs (up to 3,000 pages). Output is text-only.

About TechZila AI Research Desk

TechZila AI Research Desk covers AI models, developer tools, and emerging technology with an emphasis on practical performance analysis and real-world use cases.

Sources: Google official documentation, Gemini API pricing page, Artificial Analysis benchmarks, independent testing (August 2026).

Last updated: August 16, 2026

Post a Comment

0 Comments