Skip to main content
Google

Gemini 3.5 Flash

Google

Google May 2026 Flash model, the first of the Gemini 3.5 series, for MCP agents, tool use, and financial analysis, with a 1M-token context window.

google/gemini-3.5-flash

Documentation

Conversation

Google

Start a conversation

Google May 2026 Flash model, the first of the Gemini 3.5 series, for MCP agents, tool use, and financial analysis, with a 1M-token context window.

Enter to send · Shift+Enter for a new line

Uses POST /v1/chat/completions with your Unifically API key. Supports system and user prompts, tools, streaming, and thinking when available.

What is Gemini 3.5 Flash?

Gemini 3.5 Flash is Google's May 19, 2026 Flash model, announced at Google I/O as the first release of the Gemini 3.5 series. It is built on Gemini 3 Flash and is the base the newer 3.6, 3.7, and 3.8 Flash models grew from. At launch Google billed it as its strongest agentic and coding model to date, with output about four times faster than other frontier models, and made it the default model of the Gemini app and AI Mode. Inputs are text, images, audio, and video. The context window is 1M tokens and output runs to 64K tokens. It runs on Unifically as google/gemini-3.5-flash with five reasoning effort levels (minimal, low, medium, high, max), one more than the Flash models that followed it.

Key features of Gemini 3.5 Flash

Top of its table on MCP Atlas

83.6% on MCP Atlas, the multi-step workflow benchmark built on the Model Context Protocol. That beats Claude Opus 4.7 at 79.1%, Gemini 3.1 Pro at 78.2%, GPT-5.5 at 75.3%, and Claude Sonnet 4.6 at 69.5%, and it is a 21-point jump over Gemini 3 Flash at 62.0%. Agents that chain many tool calls are where this model wins.

The best financial analysis score in its table

57.9% on Finance Agent v2, six points clear of GPT-5.5 at 51.8%, Claude Opus 4.7 at 51.5%, and Claude Sonnet 4.6 at 51.0%. The independent Vals AI run agrees: 57.86%, third overall behind Muse Spark 1.2 at 60.60% and Claude Opus 5 at 58.63% as of September 1, 2026, at about half the cost of Opus 5. Statement reconciliation and filings analysis run better here than on the bigger Gemini.

Fast output with a solid independent score

Artificial Analysis measures 189.3 output tokens per second, rank #11 of 196 models, and scores it 52 on the Intelligence Index at high effort, rank #42 of 196. Gemini 3.6 Flash scores the same 52 on that index; Gemini 3.7 Flash scores 56.

Five reasoning effort levels, including minimal

`reasoning_effort` takes `minimal`, `low`, `medium`, `high`, and `max`. At `minimal` the response carries no reasoning tokens, which the newer 3.7 and 3.8 Flash models do not allow. Route cheap classification calls at `minimal` and hard agent steps at `max` on one model ID.

Best for

MCP-driven agent workflows

83.6% MCP Atlas, the best score in its launch table; long tool chains stay on track.

Financial analysis agents

57.9% Finance Agent v2, ahead of GPT-5.5, Claude Opus 4.7, and Claude Sonnet 4.6.

Real-world tool use

56.5% Toolathlon, above GPT-5.5 at 55.6% and Gemini 3 Flash at 49.4%.

Chart and image reasoning

84.2% CharXiv Reasoning and 83.6% MMMU-Pro, both top of the table without tools.

High-volume output

189 tokens per second, rank #11 of 196 on Artificial Analysis.

Zero-reasoning calls

`minimal` effort returns no reasoning tokens; the cheapest way to route and classify.

Use cases

Build agents that live inside MCP servers: ticket triage that reads three systems, deploy helpers that call a dozen tools in order, and research loops that search, fetch, and summarize. The 83.6% MCP Atlas and 56.5% Toolathlon scores show up as fewer dropped steps. Run finance pipelines that pull filings, reconcile statements, and draft analyst notes, where the 57.9% Finance Agent v2 score leads every model in the table. Point it at chart-heavy reports and slide decks: 84.2% on CharXiv Reasoning and 83.6% on MMMU-Pro cover charts and images in one call. For desktop automation, 78.4% on OSWorld-Verified sits within half a point of Claude Opus 4.7 and GPT-5.5. Coding agents get 76.2% on Terminal-Bench 2.1 and 55.1% on SWE-Bench Pro at Flash cost.

Limitations

Time to first answer token is slow at high effort: 13.89 seconds in Artificial Analysis measurements, even though output speed after that ranks #11 of 196. Drop to minimal or low for anything user-facing.

It is verbose. The model spent 75M output tokens to complete the Artificial Analysis index, and thinking tokens bill as output tokens, so watch token counts at high and max.

Long-context recall trails the bigger models. On MRCR v2 with eight needles at 128k tokens it scores 77.3%, against 84.9% for Gemini 3.1 Pro and 94.8% for GPT-5.5, and 26.6% on the 1M pointwise variant. Reasoning is not frontier either: 40.2% on Humanity's Last Exam against 46.9% for Claude Opus 4.7, 72.1% on ARC-AGI-2 against 84.6% for GPT-5.5, and 1656 Elo on GDPval-AA against 1769.

It is no longer the newest Flash. Gemini 3.6 Flash, 3.7 Flash, and 3.8 Flash have followed it, Artificial Analysis lists it as superseded by 3.6 Flash, and the newer Flash models are cheaper on Unifically right now. On the LMArena text leaderboard it holds rank #22 with 1479 Elo from 33,931 votes as of September 3, 2026, two places behind Gemini 3.6 Flash at 1480 and well behind Gemini 3.8 Flash at #7 with 1494. Output is text only. Most benchmark numbers come from Google's launch table; Artificial Analysis is the independent check.

Gemini 3.5 Flash vs Gemini 3.6 Flash

Gemini 3.6 Flash was a real step up. OSWorld-Verified rose from 78.4% to 83.0%, DeepSWE v1.1 from 37% to 49%, MLE-Bench from 49.7% to 63.9%, and 1M-token MRCR retrieval roughly doubled from 26.6% to 54.0%. Gemini 3.6 Flash also finishes the same work in about 17% fewer output tokens. The two models tie at 52 on the Artificial Analysis Intelligence Index. What 3.5 Flash keeps is the minimal effort level and its launch-table leads on MCP Atlas and Finance Agent v2. For anything else, 3.6 Flash or a newer Flash is the better and cheaper choice today.

When to use Gemini 3.5 Flash

Use Gemini 3.5 Flash when a pipeline is already tuned and regression-tested against it, when you need the minimal effort level for zero-reasoning calls, or when MCP-heavy tool work and financial analysis are the whole job. For new builds, start on Gemini 3.8 Flash: it is newer, stronger on coding, and cheaper on Unifically right now. Step up to a frontier model like GPT-5.5 or Claude Opus 4.7 for long-context recall and the hardest reasoning.

API examples

Call Gemini 3.5 Flash from any language by POSTing to /v1/chat/completions, the OpenAI-compatible endpoint shared by every language model on the platform. Full parameter docs live at docs.unifically.com/models/llm/google/gemini-3.5-flash.

curl -X POST https://api.unifically.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "model": "google/gemini-3.5-flash",
    "reasoning_effort": "high",
    "messages": [
      { "role": "user", "content": "Reconcile these two quarterly statements and list every line item that does not match, with the difference." }
    ]
  }'

The response comes back synchronously with the completion. Set "stream": true to receive tokens as they generate, and set reasoning_effort to minimal for calls that need no reasoning tokens at all.

FAQs

People also ask

google/gemini-3.5-flash, called through the OpenAI-compatible POST /v1/chat/completions endpoint with your Unifically API key. The same model ID also works on /v1/responses and /v1/messages.

Text, image, audio, and video input, with text output up to 64K tokens per response. The context window is 1M tokens.

Five, since reasoning_effort accepts minimal, low, medium, high, and max. At minimal the model returns no reasoning tokens at all, a level the newer Gemini 3.7 Flash and 3.8 Flash reject. Thinking tokens are billed as output tokens.

It is the best model in Google's launch table on MCP Atlas at 83.6% and on Toolathlon at 56.5%, and it scores 78.4% on OSWorld-Verified for computer use. Terminal-Bench 2.1 lands at 76.2% and SWE-Bench Pro at 55.1%.

57.9% on Finance Agent v2, the top score in its launch table, ahead of GPT-5.5 at 51.8%, Claude Opus 4.7 at 51.5%, and Claude Sonnet 4.6 at 51.0%. Vals AI's independent leaderboard puts it third at 57.86% as of September 1, 2026, behind Muse Spark 1.2 and Claude Opus 5, and first in the Precedents category at 36.4%.

Artificial Analysis measures 189.3 output tokens per second, rank

Rank

No, it is the May 2026 generation, followed by Gemini 3.6 Flash in July 2026, Gemini 3.7 Flash on August 13, 2026, and Gemini 3.8 Flash on September 2, 2026. The newer Flash models are cheaper on Unifically right now, so pick 3.5 Flash for pipelines locked to its behavior or that need the minimal effort level.