Table of Contents
Google I/O always generates a long list of announcements — this year’s official summary runs over 100 items. But what’s worth reading carefully isn’t the feature list. It’s the direction all those features point in: Google is moving AI from the application layer to the infrastructure layer.
TL;DR
Google I/O 2026’s core shift is from AI assistant to AI agent: not “do things faster for you” but “do things on your behalf.” Gemini 3.5 Flash is the fastest frontier-class model; Gemini Omni handles multimodal generation; Gemini Spark is a 24/7 enterprise agent running autonomously in the background; Antigravity 2.0 is the developer workspace for orchestrating multiple agents. The entire ecosystem is converging on this direction.
What Happened
Gemini 3.5 Flash: Flagship Intelligence at Flash Speed
Google released Gemini 3.5 Flash with a bold positioning: “frontier intelligence at Flash-series speeds.” In Google’s benchmarks, Gemini 3.5 Flash outperforms Gemini 3.1 Pro on coding and agentic tasks.
This combination didn’t previously exist — you had to choose between flagship models (slow, expensive) or Flash versions (fast, but limited capabilities). If 3.5 Flash’s claimed performance holds up under independent evaluation, the pricing implications for the API market are significant.
Gemini Omni: Generate Anything from Any Input
Gemini Omni is Google’s multimodal generation model, emphasizing “start from video input, generate output in any format.” Google calls it a “leap forward in world understanding.”
API access begins in late Q2 2026, with enterprise tier through Google Cloud in Q3.
Gemini Spark: AI from Assistant to Full-Time Employee
This is the most important announcement for enterprise engineers at I/O 2026. Gemini Spark is a 24/7 personal AI agent that runs autonomously in the background, executing tasks in the direction users specify — no manual triggering required for each action.
Google says it’s for Gemini Enterprise and Workspace customers. This is Google’s explicit signal that it’s selling AI agent capabilities directly to enterprises.
Antigravity 2.0: Developer Command Center for Agent Orchestration
For developers, Antigravity 2.0 is a standalone desktop app for “steering, customizing, and orchestrating” multiple AI agents from a single workspace. This directly addresses a real problem: when your system runs five or more AI agents, you need a way to manage their dependencies and conflicts.
Engineering Perspective
Gemini as Infrastructure
Google’s most important strategic transformation isn’t in any single feature — it’s in the architecture direction: Gemini is becoming Google’s underlying inference infrastructure, not just a callable app or API.
This is fundamentally different from how Google promoted AI before. Previously: “you can use AI to help you write emails in Gmail.” Now: “Gemini will autonomously handle the emails in your inbox you haven’t had time to respond to — no manual trigger needed.”
Search Becoming Agentic
Google Search 2026 adds “Information agents” — the search engine doesn’t just answer questions; it monitors topics on your behalf, regularly aggregates updates, and proactively notifies you. This is search’s transition from “you ask, I answer” to “I know what you want to know before you ask.”
Google’s E-commerce Bet
Google released Universal Cart — “a truly intelligent shopping cart” integrating cross-site product comparison, inventory checking, and checkout. This is Google’s most direct push into e-commerce transaction layers in years.
Why This Matters
I/O 2026’s announcement volume is high, but the strategy isn’t complicated: Google is using Gemini to upgrade all its moats (Search, Gmail, Drive, Maps, Android) into AI-native services — while OpenAI and Anthropic are still primarily API providers, Google already has complete end-to-end deployment.
For developers, Google Cloud’s AI capabilities (Gemini 3.5 Flash API, Gemini Omni, Vertex AI’s agent framework) are rapidly becoming a complete toolchain option — not just an OpenAI API alternative.
What to Watch
-
Independent performance validation of Gemini 3.5 Flash: Google’s own benchmarks have historically been contested. Waiting for LMSYS, Hugging Face, and other neutral evaluations.
-
Gemini Spark’s actual autonomy level: “24/7 autonomous execution” sounds strong in marketing; how it balances autonomy with user control needs real enterprise customer feedback.
-
Universal Cart merchant adoption: This feature benefits users, but merchants who don’t want Google intermediating transactions may push back.
References
Answers come from this article only. Click any prompt below or open the chat at the bottom right.
🇺🇸 English
Google I/O always drops a firehose of announcements. This year the official recap runs to over a hundred items. But if you sit there reading the feature list, you're going to miss the actual story. Because the story isn't in any single feature — it's in the direction all of them are pointing. And that direction is this: Google is moving AI down the stack. Out of the application layer, into the infrastructure layer.
Here's the one-sentence version. The big shift at I/O 2026 is the move from AI assistant to AI agent. Not "let me do things faster for you," but "let me do things on your behalf." And once you have that lens, all four of the headline releases suddenly line up.
Let's walk through them.
First, Gemini 3.5 Flash. Google put a pretty bold flag in the ground here — they're calling it "frontier intelligence at Flash-series speeds." And in Google's own benchmarks, 3.5 Flash actually beats the bigger Gemini 3.1 Pro on coding and on agentic tasks. Now stop and think about why that's a big deal. Up until now, you always had to pick a lane. You either went with a flagship model — smart, but slow and expensive — or you went with a Flash model — fast and cheap, but capped on capability. Google is claiming you no longer have to choose. And if that claim survives independent testing — that's the key "if" — then the pricing math for the entire API market gets shaken up. Fast, cheap, and frontier-smart, all at once, changes what everyone else has to charge.
Second, Gemini Omni. This is Google's multimodal generation model, and the pitch is "take video as input, and generate output in any format you want." Google's framing it as a leap forward in world understanding. The practical timeline: API access opens late in the second quarter of 2026, with the enterprise tier coming through Google Cloud in Q3.
Third — and if you work in an enterprise, this is the one to circle — Gemini Spark. Spark is a personal AI agent that runs around the clock, autonomously, in the background. You point it in a direction, and it just executes toward that goal. You are not sitting there triggering each individual action. Google's aiming this at Gemini Enterprise and Workspace customers, and that placement is the tell. This is Google saying, out loud, "we are selling agent capability directly to businesses." Not a chatbot you open. A worker that's always on.
And fourth, Antigravity 2.0. This one's for developers. It's a standalone desktop app for steering, customizing, and orchestrating multiple AI agents from a single workspace. And it's solving a genuinely real, near-future headache: the moment your system is running five or more agents at once, you need something to manage them — their dependencies, their conflicts, who's stepping on whom. Antigravity 2.0 is the command center for that.
Okay, so those are the products. Now let's zoom out to the engineering picture, because this is where it gets interesting.
The most important strategic move isn't any one feature — it's an architecture decision. Gemini is becoming Google's underlying inference infrastructure. Not an app you call. Not just an API endpoint. The layer everything else runs on. And you can feel that shift in how Google talks. The old pitch was: "Hey, you can use AI to help you write emails in Gmail." The new pitch is: "Gemini will just handle the emails sitting in your inbox that you haven't gotten to — no trigger needed." See the difference? One is a tool you reach for. The other is an agent that's already working.
And it's spreading into the crown-jewel products. Take Search. Google Search 2026 adds what they're calling "information agents." So the search engine doesn't just answer your question anymore — it monitors a topic for you, keeps pulling together updates, and proactively pings you. That's Search moving from "you ask, I answer" to "I know what you'll want to know before you even ask."
Then there's the e-commerce play — Universal Cart. Google's billing it as a genuinely intelligent shopping cart: compares products across different sites, checks inventory, and handles checkout, all in one place. This is the most direct swing Google has taken at the transaction layer of e-commerce in years. They don't just want to send you to the store. They want to be the cart.
So why does all of this matter? Honestly, the volume is loud but the strategy is simple. Google is using Gemini to upgrade every one of its moats — Search, Gmail, Drive, Maps, Android — into AI-native services. And here's the competitive edge: while OpenAI and Anthropic are still, for the most part, API providers, Google already owns the full end-to-end pipeline. From the model, to the product, to the billion users. For developers, that means Google Cloud's stack — the Gemini 3.5 Flash API, Gemini Omni, Vertex AI's agent framework — is quietly turning into a complete toolchain. Not just a fallback for when you don't want the OpenAI API. A real, full alternative.
Now, let's keep our skeptic hat on, because there are three things genuinely worth watching before you buy the whole story.
One: independent validation of Gemini 3.5 Flash. Google's own benchmarks have been contested before — more than once. So the honest answer is, wait for the neutral scoreboards. LMSYS, Hugging Face, that crowd. Numbers from the company selling the model are a starting point, not a verdict.
Two: how autonomous Gemini Spark actually is. "24/7 autonomous execution" is a fantastic marketing line. But the real question — the one only actual enterprise customers can answer — is how it balances that autonomy against user control. An agent that acts on its own is powerful right up until it does the wrong thing on its own.
Three: whether merchants actually adopt Universal Cart. Users love it, sure. But merchants who don't want Google sitting in the middle of their transactions have every reason to push back. And without merchants, an intelligent cart is just an empty one.
So let me leave you with the three things that actually matter here. First: the real headline of I/O 2026 isn't a product, it's a phase change — Google moved from AI as assistant to AI as agent, from doing things faster to doing things for you. Second: Gemini is now infrastructure, the inference layer underneath Search, Gmail, all of it — and that vertical integration, model to product to user, is the thing OpenAI and Anthropic can't easily match. And third: hold your applause until the independent numbers land and real enterprises report back. The vision is coherent and, frankly, kind of formidable. Whether it delivers is a question the benchmarks and the customers still have to answer.
🇹🇼 中文
2026 年 5 月 22 號,Google I/O 落幕。Sundar 跟 Demis 在台上畫了一個野心很大的軟體未來,而那個未來講白了就一句話——Gemini 藏在每一個產品裡。
這次的產品路線圖幾乎可以這樣總結:拿 Gemini,後面接一個名詞,然後出貨。Gemini Spark、Gemini Omni、Gemini Flow,清單一路往下。Google 自己給這個階段取的名字叫「代理人化的 Gemini 時代」,agentic Gemini era。意思是:搜尋是 AI agent、Gmail 是 AI agent、Android 是 AI agent,連你的眼鏡都是 AI agent。
看完整場 keynote,你會意識到一件事:Google 已經不想再用「藍色超連結」來組織全世界的資訊了。在它的新敘事裡,傳統搜尋引擎正在變成一種過時的技術。Google 真正想做的,是趕在 Anthropic 跟 OpenAI 之前,成為「通往現實本身的介面」。
先講 Google 真正的護城河,就是規模。不管你喜不喜歡 Google,它最讓人印象深刻的能力始終是規模化。它不只把核心產品服務給數十億日活用戶,過去兩年的推論量成長更是誇張。兩年前,每個月大概 9.7 兆個 token;現在呢,每個月大概 3.2 千兆個 token。從兆跳到千兆,而且這數字還會繼續加速。相對的,Alphabet 的資本支出也跟著爆炸性成長,蓋新的基礎設施來撐住這一切,包括你們用 nano banana 生的那一大堆 AI 圖。
那撐住這種規模的關鍵之一,就是 Google 自家的 TPU,張量處理器。今年 I/O 有個底層變化:Google 把 TPU 拆成了兩種各司其職的晶片。一種叫 TPU-T,專門優化訓練;另一種叫 TPU-I,專門優化推論。換句話說,一顆晶片負責「教模型怎麼思考」,另一顆負責「把結果大規模地跑出來」。這種訓練跟推論分家的做法,背後的邏輯是:當你的推論量已經衝到千兆 token 這個等級,與其用一顆通用晶片全包,不如針對兩種截然不同的工作負載分別優化,這樣更划算。
接下來是這場 I/O 的頭條,Gemini Omni。這是一個能吃下任何輸入的模型——文字、影片、聲音,然後產出任何輸出。Demis Hassabis 看起來是徹底的「world model 信徒」。這類模型不再只是生成像素而已,而是理解語言、理解物理、理解運動,還有你世界裡的其他一切,理解到足以「按需模擬現實」的程度。
跟 Omni 一起來的,是 Gemini app 全新的設計系統,叫 Neural Expressive。乍看之下它就是換了新 icon、漸層變好看的介面改版。但它真正特別的地方在於,它是為了「按需生成 UI 元素」而優化的——像是圖表、時間軸,甚至是在你下 prompt 之前根本還不存在的迷你 app。
再來講核心模型這條線,Google 發表了 Gemini Flash 3.5。這裡要先講清楚:這不是那顆最強的大腦,這是那顆「快」的模型。根據 Google 自家那種「trust me bro」的 benchmark,Flash 3.5 的表現逼近 Opus 4.7 跟 GPT-5.5,但跑起來快得多。在它秀出來的那張速度對智慧的象限圖裡,Flash 自己獨佔了一個象限。不過再提醒一次,這不是 Google 的頂規模型。真正的旗艦 Gemini 3.5 Pro 目前還沒揭曉,預計要今年夏末才會發表——這一點讓網路上不少人相當失望。
然後是開發工具,Antigravity IDE,還有那場 Doom demo。不是每個人都對 Google 這個新方向買單。Antigravity 的前身是 Windsurf,跟 Cursor 一樣是主打 AI coding 的工具。而它最新版本,又一次跟著 Cursor 的腳步,看起來更像是 OpenAI Codex 的複製品:重心從「寫程式」偏向「管理一群 agent」。老派工程師大概不會太開心這個轉向。但現場 demo 確實有點狠——他們用這個工具,從零打造了一整個完整的作業系統,過程花了大約 12 小時,燒掉數十億個 token。接著他們想在上面跑 Doom,結果因為缺 driver 失敗了;於是他們當場讓 Gemini 把缺的 driver 寫出來,幾秒鐘後,Doom 就跑起來了。
那如果要抓出這場 I/O 背後的策略主軸,其實不複雜:Google 想把 Gemini 變成所有產品的底層,讓 agent 直接替使用者行動,而不只是提供一個你可以呼叫的 app 或 API。搜尋、Gmail、Android、眼鏡,這些原本的入口,全被重新包裝成「代理人」。
最後幫你收三個重點。第一,這場 I/O 的核心策略就是「Gemini 無所不在」,把每個入口都變成會自己行動的 agent。第二,硬體上 TPU 拆成訓練跟推論兩顆晶片,模型上主打的 Flash 3.5 賣的是速度而不是最強腦,而真正的旗艦 3.5 Pro 還要等到夏末。第三,這套「代理人化」到底是真本事,還是只是把 Gemini 這個名詞接到更多產品後面出貨,就得等這些功能真正落地,用實際體驗來檢驗了。
Tags
Related Articles
After a 1,000,000x AI Compute Leap: What Jeff Dean Sees Next
Jeff Dean breaks down where the million-fold AI compute gains actually came from — specialized hardware, distributed training systems, and architecture efficiency — and where the next phase is headed.
RAG's Five Stages: From Pipeline to Reasoning Retrieval, and the Naive RAG on My Own Site
Over the past two years RAG evolved from a 'linear pipeline' to 'loop-based reasoning'. It maps cleanly to five stages: Naive, Advanced, Modular, Graph, Agentic. The real inflection point is control moving from pipeline to agent — a System 1 → System 2 shift. Looking back at engineer-news's own RAG stack, it's stuck at the Naive edge — so this post also lays out what to fix next.
Building a Real RAG: 5 Infra Lessons from InfiniFlow's 2024 Year-in-Review
The previous post zoomed out for a five-stage panorama of RAG. This one zooms in on the five infra lessons any real RAG has to face: document ingestion, contextualized chunking, three-lane hybrid search, tensor reranker, and GraphRAG's semantic gap. Each lesson is checked against engineer-news's current stack, ending with a priority list for a personal site.