Skip to main content

Google Launches Gemini 3.5 Flash, Shifting Focus to Agentic AI Despite Rising Costs

24 JUNE 2026·3 MIN READ·1 SOURCE·Official source

Google has released Gemini 3.5 Flash, a high-speed model optimized for multi-step agentic workflows and tool use. While it undercuts the base price of Gemini 3.1 Pro, heavy token consumption in automated tasks makes the new model significantly more expensive to operate than its predecessor.

Google Launches Gemini 3.5 Flash, Shifting Focus to Agentic AI Despite Rising Costs

Key takeaways · 4

  • 01

    Gemini 3.5 Flash delivers 280 output tokens per second, making it four times faster than similar frontier models.

  • 02

    Token prices tripled compared to the previous Flash generation, now costing $1.50 per million input tokens.

  • 03

    Google integrated computer use directly into 3.5 Flash via the Gemini API and Enterprise Agent Platform.

  • 04

    The new Gemini Omni model allows conversational video editing using text, audio, and visual prompts.

The Gemini 3.5 Flash Launch

On May 19, 2026, Google Deepmind released Gemini 3.5 Flash, positioning it as the new default model for production developers. [13]

The model supports a one-million token context window and a maximum of 65,000 output tokens. [13] Google Chief Executive Officer Sundar Pichai stated this launch marks the beginning of the "agentic Gemini era." [12] Gemini 3.5 Flash delivers more than 280 output tokens per second, making it the fastest model in its intelligence class. [1]

According to Google's internal measurements, this latency reduction makes the model four times faster than comparable frontier models. [3] The company also announced that the Gemini 3.5 Pro variant is expected to launch publicly in June 2026. [3][4]

Agentic Capabilities and Tool Use

Gemini 3.5 Flash is designed to execute multi-step actions on external tools, including file systems, browsers, and third-party APIs. [3] The model scored 76.2 percent on Terminal-Bench and 83.6 percent on MCP Atlas for tool-use reliability. [13]

Google also noted that computer use is now a built-in tool supported in Gemini 3.5 Flash via the Gemini API and Gemini Enterprise Agent Platform. [16] As part of this agentic focus, Google is introducing Search agents that monitor the web in the background to provide real-time information. [14]

The company also introduced Gemini Spark, a new agentic assistant aimed at helping users complete tasks across Search, Android, YouTube, and connected apps. [4]

The Pricing Paradox

Google set the price for Gemini 3.5 Flash at $1.50 per million input tokens and $9.00 per million output tokens. [1][3] While this base price is roughly 40 percent less than the Gemini 3.1 Pro model, it represents a substantial increase over its direct predecessor. [1][3]

Token prices for the new Flash model have tripled compared to Gemini 3 Flash, which charged $0.50 for input and $3.00 for output tokens. [1] An analysis by Artificial Analysis found that Gemini 3.5 Flash costs 5.5 times more to run in benchmark testing than Gemini 3 Flash. [1]

Because agent-based tasks burn through so many more tokens, total costs can end up 75 percent higher than Gemini 3.1 Pro. [1]

Multimodal Expansion and Subscriptions

Alongside the new Flash model, Google unveiled Gemini Omni, an AI model built to simulate physical environments and generate video. [4][6] The first version, Gemini Omni Flash, allows users to generate and edit videos using combinations of text, images, audio, and video prompts. [6]

Google also announced a new $100-a-month Ultra plan that includes 20TB of cloud storage and priority access to Google Antigravity. [5] The original Ultra plan is now cheaper at $200-a-month. [5]

The company is switching to compute-based usage limits, allowing plans to adapt based on actual usage instead of fixed quotas. [5]

What it means

Google’s transition to Gemini 3.5 Flash highlights a strategic shift toward speed and autonomous capabilities over raw reasoning power. By prioritizing multi-step tool execution and multimodal workflows, the company is attempting to establish Gemini as a foundational operating layer rather than just a chatbot. However, the economics of this transition present a complicated reality for developers. While the advertised token price undercuts Gemini 3.1 Pro, the massive token consumption required for agentic loops makes actual operating costs significantly higher than earlier Flash models. This mirrors broader industry pricing trends seen with Anthropic’s Opus 4.7 and OpenAI’s GPT 5.5. As organizations evaluate these new models, the key metric will shift from basic prompt cost to total task efficiency. What the sources don't address: How Google plans to manage the compute infrastructure demands if agentic web monitoring and continuous background tasks become widely adopted by its billion monthly users.

Google is aggressively pivoting its AI strategy toward automated, multi-step execution frameworks. While the capabilities are expanding rapidly, developers must carefully audit their token budgets as agentic loops drastically increase consumption.

Why it matters
Story quiz

Test yourself on this story — 3 questions.

Create a free account to take the quiz, earn XP, and get a daily session built for your industry.

Take the quiz

How this developed

  1. 25 June 2026

    Event evidence refreshed from source cluster.

  2. 24 June 2026

    Event evidence refreshed from source cluster.

  3. 21 May 2026

    Event evidence refreshed from source cluster.

  4. 20 May 2026

    Event created from source cluster.

Sources

AI fluency, one session a day, built for your work.