Zilmac Blog
← Back to Tech Practice

What Is Gemini 3.8 Flash? Features, Performance, Pricing, API vs GPT-6 Astra (2026)

API & Agents ·~14 min read

Developer debugging code on a multi-monitor workstation, illustrating Gemini 3.8 Flash API and Agent coding workflows
$0.75 / $3.75
3.8 Flash launch price (per million tokens, through 2026-12-31)
1M / 64K
Context window / max output
~13×
Lower token unit price than GPT-6 Astra

Gemini 3.8 Flash is Google's latest Flash-tier workhorse, shipped on September 2, 2026. It targets long-horizon software engineering, autonomous Agents, and multi-step enterprise workflows. It is not the Gemini 3 family's flagship Pro. Instead it packs near-frontier reasoning and coding into Flash speed and launch pricing: $0.75 per million input tokens and $3.75 per million output tokens, through December 31, 2026.

The same week, OpenAI released flagship GPT-6 Astra (API model ID gpt-6-astra) at a standard rate of $10 per million input and $50 per million output. Both pitch Agents, coding, and long context—but the bills differ by an order of magnitude. This article uses Google's official blog, Gemini API docs, the DeepMind model card, and OpenAI's model page to unpack features, performance, pricing, and API, then give a selection guide you can actually ship. Vendor-reported scores are not treated as final acceptance.

Three audiences: Agent / coding-assistant teams picking a default model; platform and finance teams estimating the monthly bill; engineers building an adapter between Gemini Interactions API and OpenAI Responses API.

Three things to remember first
  • 3.8 Flash is a GA stable model. The ID is simply gemini-3.8-flash—no preview suffix.
  • Launch pricing ends 2026-12-31. From 2027-01-01 the standard rate becomes $1.50 / $7.50.
  • Thinking tokens are billed as output. Higher thinking levels can raise the per-task bill even if the unit price does not change.

What Gemini 3.8 Flash is

Per Google's official announcement, 3.8 is the third Flash iteration in the Gemini 3 family—about three weeks after 3.7 Flash. It matches 3.7 on speed and launch price, but is clearly stronger on software engineering, Agent tasks, and multi-step reasoning in specialist domains. Google shipped two variants at once:

  • Gemini 3.8 Flash: a general-purpose workhorse for developers and enterprises, available via Gemini API, Google AI Studio, Gemini Enterprise, Antigravity, and Gemini app / Search AI Mode / Sheets (Google AI Pro or Ultra required).
  • Gemini 3.8 Flash Cyber: a defensive model for vulnerability discovery and automated patching. Access is limited to trusted governments, critical infrastructure, and software maintainers through the Fairwind program. This article covers only the publicly available Flash and does not unpack unpublished defensive interfaces.

The DeepMind model card states that the knowledge cutoff is mostly March 2026, with some domains still at January 2025 (same as the Gemini 3 family). High-effort thinking can be slower, more likely to time out, and more token-hungry. Those are not footnotes—they are constraints you must include when you estimate latency and cost.

Google also made 3.8 Flash the default model for Antigravity Agent and the Antigravity SDK. New hosted Agent projects call it unless you change the config.

Features: context, thinking levels, and built-in tools

Per the latest Gemini model docs, the public capabilities collapse into one spec table:

Spec Gemini 3.8 Flash
Model ID gemini-3.8-flash (stable / GA)
Input limit 1,048,576 tokens (~1M)
Max output 65,536 tokens (docs also write 64K)
Thinking levels low / medium (default) / high; minimal errors
Input modalities Text, images, video, audio, PDF
Output modalities Text
Billing modes Standard, Batch, Flex, Priority

How to choose a thinking level

The core knob on 3.8 Flash is thinking_level, not a second model name. Official guidance maps roughly like this:

  • low: latency-sensitive work, drafts, classification, simple rewrites. Lowest token overhead.
  • medium (default): most complex code and Agent loops. First-pass correctness is usually better.
  • high: deep reasoning, math, hard multi-step orchestration. The model takes more reasoning steps and calls more tools.

Google's own phrasing is that 3.8 Flash “works harder.” On complex tasks it takes extra steps and retries tools—especially at the high level. Independent evals show high can emit about 30% more output tokens per task than 3.7 Flash. The unit price is unchanged; the per-task cost still rises. Production routing should lock the thinking level by task type, not set global high.

Built-in tools and input modalities

3.8 Flash inherits the Gemini 3 tool surface: context caching, code execution, file search, Function Calling, Structured Output, Search grounding, Maps grounding, URL context. Computer Use remains Preview. For Agents, that means you can get “model + hosted tools” running first, then decide which actions must return to your own functions.

Multimodal input is a fit when you drop a mockup, error screenshot, meeting recording, or PDF spec into the same Agent task. Output is still text, so deliveries into business systems still need Structured Output or your own schema validation—do not expect the model to “also” draw a UI.

If you are still moving from generateContent to Interactions API, split the interface layer from the state layer before you change the model name. For the migration order, see After the 2026 Gemini API update, which Agent layer to change first.

Performance: official benchmarks and the cost of “working harder”

Google stresses that 3.8 Flash approaches—or beats—more expensive frontier models on several public benchmarks, while staying at Flash prices:

  • DeepSWE v1.1 (long-horizon software engineering): Google says it outperforms most larger frontier models. Third-party comparison tables commonly list about 73.7%–73.8%. OpenAI's self-report for GPT-6 Astra is 74.1%. On the same mini-swe-agent harness they are extremely close; a 0.3-point gap does not decide a winner.
  • HLE-Verified (cross-disciplinary multi-step reasoning): official number 54.9%.
  • Specialist Agents: Vals Finance Agent V2, Harvey Legal Agent, and similar. Google says 3.8 Flash beats 3.7 Flash and other frontier models. These are vendor-chosen scenarios—treat them as a direction signal, not your repo's regression set.

Artificial Analysis's Intelligence Index puts GPT-6 Astra a tier higher (low-thinking around 57 vs 52; high-thinking composite scores are commonly reported in the 59–67 range, and the methodology is not consistent). The more useful takeaway for engineering teams: coding-Agent pass rates are already squeezed into the same narrow band; what actually splits the bill is token burn and unit price.

Developer debugging a Gemini API and Agent workflow in a code editor
3.8 Flash's edge usually shows up on long tasks that rewrite a repo and retry tools—not on single-turn Q&A.
Don't paste the leaderboard into your SLA

DeepSWE, HLE, and GPQA all depend on the eval harness, thinking level, and tool switches. Change the repo or the timeout policy and the ranking moves. Before you ship, compare 3.7 Flash, 3.8 Flash, and Astra on your own 20–50 regression tasks. Record pass rate, thoughtsTokenCount, wall-clock time, and dollar cost.

Pricing: launch window, the 2027 hike, and billing traps

The table below follows Google's published 3.8 Flash rates, all in USD / million tokens. The launch window runs through December 31, 2026. Standard rates double on January 1, 2027.

Mode Input (through 2026-12-31) Output (through 2026-12-31) Input (from 2027-01-01) Output (from 2027-01-01)
Standard $0.75 $3.75 $1.50 $7.50
Batch / Flex $0.375 $1.875 $0.75 $3.75
Priority $1.35 $6.75 $2.70 $13.50

Three easy-to-miss line items:

  1. Thinking tokens = output tokens. Bill thoughtsTokenCount in the response at $3.75/M (launch window). A high-level task that “only wrote 2K words” may already have produced tens of thousands of thinking tokens.
  2. Context caching and Search / Maps grounding are billed separately. Cache has write and storage fees, and those change in 2027 too. Grounded retrieval is per-request after the free allowance.
  3. 3.7 Flash is still available. If latency and token overhead matter more than accuracy, Google explicitly suggests staying on 3.7—or locking 3.8 to low.

Rough math for one Agent coding session: 80,000 input tokens (repo summary + history + tool returns), 20,000 output tokens including thinking. Launch-window Standard is about $0.06 + $0.075 = $0.135. The same session on GPT-6 Astra standard pricing (≤272K input) is about $0.80 + $1.00 = $1.80—roughly 13×. That still ignores Astra's whole-request hike after 272K input.

For the monthly bill that stacks subscriptions, API overage, and Cloud Mac, see How much AI programming really costs per month. Markup rules on multi-model gateways are in OpenRouter API pricing comparison.

API: model ID, Interactions, and generateContent

3.8 Flash can use two interfaces. New projects are officially pointed at Interactions API (POST /v1beta/interactions); most existing code is still on generateContent. Both accept the same model ID.

Interactions API (example)
curl -X POST https://generativelanguage.googleapis.com/v1beta/interactions \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash",
    "input": "Summarize the failing test and propose a patch.",
    "generation_config": { "thinking_level": "medium" }
  }'

The existing path is POST /v1beta/models/gemini-3.8-flash:generateContent. Do not only swap a string when you migrate: Interactions continues a run with previous_interaction_id; generateContent usually means you assemble history yourself. Tool loops, streaming events, and call_id returns all need regression.

Four things to lock down in production:

  • Pin gemini-3.8-flash in config. Do not depend on a “latest Flash” alias and get a silent upgrade.
  • Set thinking_level by route: support drafts low, repo edits medium, math / planning high.
  • Log both usage and thoughtsTokenCount, or the monthly bill will not reconcile.
  • Structured Output owns the final JSON; Function Calling owns intermediate actions. Do not mash them into one schema.

Consumer entry points are the Gemini app and Search AI Mode. Production Agents should use Gemini API or Gemini Enterprise, with keys, quotas, and data retention managed separately. Server-side state on Interactions has a retention window; audit logs still belong in your own storage.

How it compares to GPT-6 Astra

GPT-6 Astra is OpenAI's early-September 2026 flagship, aimed at “the hardest end-to-end work”: complex reasoning, coding, Computer Use, research, and long documents. Specs are from the OpenAI model page.

Dimension Gemini 3.8 Flash GPT-6 Astra
Positioning Strongest Flash workhorse / default Agent model Flagship, hardest tasks
Model ID gemini-3.8-flash gpt-6-astra
Release 2026-09-02 (GA) Announced 2026-09-03, API ~09-04
Context / max output ~1.05M / 64K–65K 1,050,000 / 128,000
Thinking control thinking_level: low / medium / high reasoning.effort: low–max (no none)
Standard unit price (≤ long-context threshold) $0.75 / $3.75 (launch window) $10 / $50; cached input $1, write $12.50
Long-context surcharge Public table does not list a separate 272K threshold Input >272K: entire-request input/cache 2×, output 1.5×
Primary interface Interactions API + generateContent Responses API (broader tool surface)
Input modalities Text / image / video / audio / PDF Text + image
Knowledge cutoff Mostly 2026-03 (some 2025-01) 2026-04-30
DeepSWE v1.1 ~73.7%–73.8% (comparison tables) 74.1% (OpenAI self-report)

Astra's tool surface leans more “host-managed”: web search, file search, image generation, Code Interpreter, hosted Shell, Apply Patch, Computer Use, Tool Search. 3.8 Flash's advantages are multimodal input, launch price, and grounding into Google Search / Maps / Workspace. Astra also hikes the entire request on very long input: stuff a whole repo or a million-token docket in one go and $10/$50 becomes $20/$75.

Batch / Flex is about half of standard on both sides. Astra also has Fast (~2×); Gemini's counterpart is Priority (~1.8×). When you compare, align the same service class, or you are comparing “cheap but queueable” with “expensive but preempting.”

How to choose and how to route

Do not make “latest flagship” the only default. A stabler approach is to tier by task:

  • Default to 3.8 Flash medium: everyday code edits, tickets, RAG, most Agent loops. Best value in the launch window.
  • Latency-sensitive: 3.8 Flash low or stay on 3.7 Flash: completion, classification, short replies. Look at p95 first, then pass rate.
  • Hard tasks escalate to Astra or 3.8 Flash high: cross-repo refactors, Computer Use, needle retrieval in long documents, reports that need 128K output. Compare pass rate and dollars on shadow traffic before you cut the main path.
  • Re-budget 2027 at $1.50 / $7.50. After launch pricing ends, 3.8 Flash is still far below Astra, but it is no longer “almost-free frontier.”

The real Agent cost is often not the model—it is the execution environment. When you need Xcode, signing, simulators, and local tool loops on a Mac, the model bill and the machine bill must be counted separately. Zilmac Cloud Mac is a fit for keeping a stable Apple toolchain on a fixed node, then calling 3.8 Flash or Astra over API—so keys, signing, and GPU inference do not pile onto the same machine.

FAQ

Are Gemini 3.8 Flash and Gemini 3.8 Pro the same model?

No. 3.8 Flash is the Flash-tier workhorse. Google stresses same speed and price as 3.7, with stronger reasoning. As of this writing (2026-09-08), the public GA push is gemini-3.8-flash. Do not apply Flash launch pricing to an unverified Pro price list.

Can I call 3.8 Flash Cyber with a regular Gemini API key?

Not under ordinary developer quota. Cyber goes through Fairwind, for trusted defensive parties. Everyday production and consumer use should stay on 3.8 Flash.

Is upgrading from 3.7 Flash to 3.8 just a model-name change?

You can change the model ID first, but you must regress tool loops, Structured Output, and thinking levels. 3.8 may call more tools and spend more thinking tokens, so latency and the bill will move. Compare on 5%–10% of traffic before a full cutover.

Which is stronger—Gemini 3.8 Flash or GPT-6 Astra?

There is no single answer. They are nearly tied on DeepSWE; Astra leads on some composite intelligence indexes and hosted tools; 3.8 Flash leads on unit price, multimodality, and the launch window. Trust your regression set and dollars per task, not a keynote slide.

The short version

  • 3.8 Flash: high-value default Agent / coding model for fall 2026. Lock medium first; watch thinking tokens.
  • Astra: reserve for genuinely hard work that is worth $10/$50. Watch the 272K threshold and cache write fees.
  • 2027-01-01: rebuild the budget on doubled Flash rates. Do not write launch pricing into an annual contract.

Models via API, builds on Cloud Mac

Gemini 3.8 Flash and GPT-6 Astra handle inference; iOS / macOS pipelines still need Xcode, signing, and a stable node. Zilmac Cloud Mac is a fit for pinning the Agent's execution environment.

Bill models and machines separately—and scale them separately. — View Cloud Mac offers

Limited Offer

Zilmac

Cloud Mac, remote dev, and Mac VPS—Apple toolchain support for iOS and cross-platform teams.

Back to Home
Limited Offer View Plans