Google I/O 2026

A 128K Window Removes the Main Reason to Reach for a Hosted Model

Original speaker(s): Olivier, Product Lead · Google DeepMind

Verified sourceSession date not verifiedpresentation47:46EN2 min read

Open weights compete not by outperforming the frontier but by offering properties it cannot sell at any price — and extending the context window removes the most common practical reason teams reached for a hosted model regardless.

The specification change worth extracting from this session is the context window: from 32,000 tokens to 128,000 for smaller models, and up to 256,000 for larger ones (2:59).

For open models this matters more than a benchmark improvement, because it changes which problems are solvable without infrastructure. A 32,000-token window means retrieval — chunking documents, building an index, managing what gets included. At 128,000, a great many tasks fit whole, and everything you would have built to work around the limit becomes unnecessary.

Where these models are expected to run

The efficiency work is explicitly aimed at the edge (3:50), and the deployment options describe the range this is intended to cover.

At one end, a browser running the model with zero ongoing server cost, or a local runtime through an interface compatible with the standard API (22:09). At the other, one-click deployment producing a hosted endpoint, with the larger variant available as a service (10:17, 10:38).

The compatibility detail is the strategically significant one. A local model that speaks the same interface as the hosted APIs is substitutable without changing application code — which means the decision about where a model runs becomes an operational choice rather than an architectural commitment. That is precisely the property that makes open weights competitive: not that they are better, but that switching costs nothing.

The hybrid argument

The case they make for combining local and hosted is framed around developer effort: doing it by hand is difficult, and having a supported path simplifies it (20:12).

The underlying economics are straightforward and rarely stated. Most requests in most applications are routine, and a small model handles them at a fraction of the cost — or at no marginal cost at all when it runs on the user's device. A minority need more capability. Routing between them captures most of the quality at a fraction of the spend.

What makes this hard by hand is not the routing logic. It is deciding which requests belong where, and detecting when a local model is failing quietly rather than obviously. Neither is solved by having a supported path, though a supported path is a precondition for solving them.

What open weights are actually for

Reading the announcements together, the argument for these models is not that they outperform the frontier. It is a set of properties the frontier cannot offer at any price.

They run where the data is, which resolves a category of compliance problem rather than mitigating it. They cost nothing per request on a device the user already owns. They do not change underneath you when a provider ships an update. And they can be adapted to a domain in ways a hosted API does not permit.

The context window extension matters because it removes the most common reason teams reached for a hosted model anyway. A model you can run yourself that also holds an entire document in view is a different proposition from one that needs a retrieval pipeline to be useful — and it is the version that competes.

Key numbers

32K → 128K/256K
context window extension across the open model family 2:59

Talk chapters

Key takeaways

  1. 01

    The context window moves from 32,000 tokens to 128,000 for smaller models and up to 256,000 for larger ones, removing the need for retrieval on many tasks. 2:59

  2. 02

    Efficiency work is aimed explicitly at edge deployment, where the model runs on hardware the user already owns. 3:50

  3. 03

    A browser or local runtime speaking the standard API interface makes the model substitutable without changing application code. 22:09

  4. 04

    One-click deployment produces a hosted endpoint, with larger variants available as a service — the same model across the whole range. 10:17

  5. 05

    The hybrid case is made on developer effort: routing between local and hosted by hand is difficult enough that a supported path is a precondition. 20:12

Entities mentioned

Related talks

Jeff Dean on Why Tools, Not Models, Are the Next Bottleneck (Google I/O 2026)
Jeff Dean on Why Tools, Not Models, Are the Next Bottleneck (Google I/O 2026)

Four of Google's model, product and search leads on what changes once agents run for hours rather than seconds, and the most quotable argument comes from Dean: the constraint is moving out of the model and into the tools around it. By Amdahl's law, an agent spending half its time in tools built for human-speed interaction cannot gain more than a doubling however fast the model becomes — which reframes a great deal of current infrastructure work as latency debt. Their internal response is concrete: rewriting Python tooling into Go, framed as a fully specified translation task rather than an open prompt, produced order-of-magnitude speedups overnight. Reid supplies the counterweight from Search, where acceptable latency turns out to scale with how much work is being taken off the user rather than being a fixed budget. Woodward's detail is the quietest and perhaps the most telling: teams that have stopped writing product documents for humans and now write context files for models to act on directly.

panel

"Pick Up the Extinct Animal": Where Robotics Actually Stands
"Pick Up the Extinct Animal": Where Robotics Actually Stands

The anecdote that opens the panel does the work: a robot asked to pick up the extinct animal selected a dinosaur toy, with nothing in its training data connecting the phrase to the object. That transfer from language models into machines with hands is the premise of the current wave. What the practitioners then describe is where it stops. Physical intelligence is about exerting force and using a body to do it, which is knowledge about consequences — the one thing a corpus of internet images contains almost nothing about. The humanoid question gets an honest treatment: not that human shape is optimal, but that the world is already built for it, plus a development-loop argument about collecting data and deploying on the same hardware. The most useful passage is scepticism about the field's favourite shortcut: generated video looks realistic and does not hold up for dexterous manipulation, because looking right and being physically consistent are different properties.

panel

When Developers Stop Opening the Editor, Chat Becomes an Interrupt Handler
When Developers Stop Opening the Editor, Chat Becomes an Interrupt Handler

The observation that organises this session is not about capability but about attention. Engineers increasingly file a ticket rather than opening an editor, and the code comes back — which changes what the surrounding tools are for. If the agent works while you do something else, the conversation between you is no longer a workspace; it is the mechanism by which the agent surfaces a question it cannot resolve alone. Interfaces built for continuous conversation optimise for flow, and interfaces built for interruption should optimise for the opposite. A runtime constraint follows immediately: an agent that starts a long-running job cannot block until it finishes, which turns out to be a workflow-engine problem rather than a model one. The panel's closing formulation — that deciding what to build is the hard skill and always was — reads as reassurance and functions as a warning, since that judgement is downstream of exactly the work now being delegated.

panel

The Moment It Stops Being Single Player
The Moment It Stops Being Single Player

The most honest moment here is an aside about how the presenters have tracked their own projects: plans in documents, plans in spreadsheets, plans in bug comments, and once a plan written on a receipt. That describes the actual category being addressed — not software nobody has built, but the small internal tool every team improvises badly because building it properly was never worth the effort. The demo turns on a single question: the generated app is strictly single player, so what happens when you want to share it with the team? That boundary is where improvised tools historically died, because it is where accounts, shared storage and access rules begin. Here it is crossed in one step, with the access rules generated and deployed automatically — which is convenient, and is also the moment the application acquires obligations nobody reviewed.

session

Why Google Dropped Chat Turns for Steps: The Interactions API at I/O 2026
Why Google Dropped Chat Turns for Steps: The Interactions API at I/O 2026

The clearest statement at I/O of how an agent API differs from a chat API, and the reasoning behind each departure is stated rather than assumed. Three changes matter. Conversation state moves to the server: a call returns an identifier, and passing it back continues the thread, retiring the client-side history array. The data model abandons alternating user and model turns for discrete steps, on the argument that a trace containing reasoning, tool calls, environment responses and compaction was never really a conversation and modelling it as one distorted it. And agents receive their own persistent remote environment rather than acting on the caller's machine — addressable by identifier, and shareable, so a research agent's output files become an application builder's input without passing through the context window. Schmid is explicit that scaffolded environment files are deliberately not model input, which is what keeps large artefacts out of the context budget. Schaeff's first half covers the real-time voice path, where the notable property is speech-to-speech across ninety languages with transcription of both directions.

presentation

Pichai Calls Google a Buffer Between People and the Raw Internet
Pichai Calls Google a Buffer Between People and the Raw Internet

Pichai's framing of Google as the buffer between people and the raw internet is offered as continuity — search did it, browsers did it, agents do it more — and it is also the most contested claim in the industry, because a buffer decides what passes through. He reaches immediately for the counterweight, the connection people feel to creators they follow, which is precisely the tension the company is currently managing without resolving. Two answers are sharper than the format usually produces. On competition he describes participants running on different pre-training and release cadences rather than at different speeds in one race, which is a more honest account than the leaderboard framing and comes from someone with an interest in leaderboards. On security he acknowledges models improving at cyber work, which is the one domain where better capability does not obviously net out positive, since an attacker needs one vulnerability and a defender needs all of them.

fireside