The argument for running models locally used to be privacy. In this session it is something more specific, and the reasoning is better.
The claim is that with a smaller model you can work at the full context length and largely stop worrying about compaction — because the moment you run agentic workloads locally, you hit hardware limits (16:04).
Why compaction is the right thing to optimise against
Compaction is what happens when a conversation outgrows the window and earlier material has to be summarised or discarded. It is the failure mode that quietly degrades long agent runs: the agent forgets a constraint from forty steps ago, or retains a summary that lost the detail that mattered.
Framing local inference as a way to avoid it is a genuinely different argument from the usual ones. It is not about keeping data on the machine or reducing cost. It is that a smaller model with room for everything can outperform a larger model that had to throw context away — for tasks where the context is what matters.
That is a real trade with a stated boundary. The claim is not that local models are better. It is that for a class of work where continuity across many steps dominates, capacity to hold the whole problem beats raw capability on any single step.
The demonstration figure — a context window in the hundreds of thousands of tokens on ordinary hardware (15:37) — is what makes the trade available at all, and it comes with a quantisation caveat they state rather than hide.
Good enough is doing the work
The framing offered is that local models have become viable for real work, and that having something good enough is genuinely useful (12:26).
Both halves matter. Viable for real work is a threshold claim, and threshold claims are the ones that change behaviour, because a tool below the line gets no use at all and a tool above it gets used constantly. The marginal quality above the threshold matters far less than crossing it.
The adoption figure attached — over eight million active developers (1:53) — is consistent with a threshold having been crossed, though it measures installation rather than sustained use. The more telling detail is the thirty seconds from nothing installed to a working setup (13:19). Setup cost is what determines whether people try the second thing after the first one disappoints, and thirty seconds means the experiment is nearly free.
The hybrid position they actually hold
The session's stated subject is how local and cloud models work together, and the honest reading is that neither side has won.
Their answer to which model to use is that it depends (18:50) — unsatisfying, and correct. The choice depends on what the task needs held in context, how much hardware is available, and what tolerance the work has for a weaker single-step answer.
The strategic observation is that this decision now exists at all. For two years the answer was to call the best available hosted model, because nothing else was close. A world where a competent local model runs at full context on a laptop is one where architecture involves a routing decision, and routing decisions are where the interesting engineering ends up.
What the session does not resolve
The hardware limit is named and not quantified. Running agentic workloads locally means becoming hardware-bound (16:04), and the practical question — what machine, for what workload, before it stops being viable — is answered only through the demonstration.
That gap matters because the argument is entirely about a trade between capability and capacity. Without knowing where the capacity ceiling sits on typical hardware, a developer cannot tell whether the trade is available to them or only to someone with a much larger machine.
The rest of the argument is sound enough that this is the number worth asking for.
关键数据
演讲章节
关键要点
- 01
A smaller local model at full context length largely removes compaction, which is the failure mode that degrades long agent runs. 16:04
- 02
Running agentic workloads locally makes you hardware-bound quickly, which is the boundary on the whole argument. 16:04
- 03
Local models are described as viable for real work, a threshold claim rather than a quality claim. 12:26
- 04
Setup takes about thirty seconds from nothing installed, which makes trying the second thing after a disappointment nearly free. 13:19
- 05
Their answer to which model to use is that it depends — the point being that this is now a routing decision rather than a default. 18:50
提及的实体
相关演讲

The rare enterprise session that describes the wiring rather than the outcome. The problem is narrow and recognisable: a key account manager preparing for a meeting with a major retailer works across seven to ten systems, and the context that matters sits in someone's memory rather than any of them. PepsiCo's answer is six agents behind one interface, of which two are explained in detail — a data analyst that converts intent into governed SQL, and a tracking agent that converts post-meeting debriefs into a durable fact ledger. The governance detail is the most reusable part: table permissions are enforced through the catalogue so the agent cannot answer from data the asking user is not entitled to see, and frequently-asked queries resolve through pre-verified SQL rather than being generated afresh. Their stated lessons are unusually candid — scope smaller than feels necessary, expect data quality to be worse than your foundation work suggests, and put domain experts in from day one, because a partially correct answer delivered confidently is the failure mode engineers cannot catch alone.

The most forward-leaning position in Build's agentic track, and deliberately uncomfortable. Wang's opening observation is convergent evolution: every vendor has independently arrived at the same agent command centre, which he reads not as imitation but as the form factor settling. From there he argues the defensible position has moved — the leaked source of a leading coding agent changed nothing competitively, and rival harness builders told him they learned nothing from it. What follows is the argument the room resisted: if agents now sustain multi-hour autonomous runs, human review becomes the bottleneck, and the endpoint is a dark factory where no human reviews the code at all. He does not present this as desirable. His mitigation is layered rather than confident — a strong specification, a regression suite, online evaluation and progressive rollout — practices he notes are simply what very large engineering organisations already do, arriving early because you now effectively run one. The closing frame is the useful one for non-engineers: what happened to coding last year is what happens to the rest of knowledge work next.

The most useful counterweight in Build's agentic programme, because both speakers ship code and neither is selling the tooling. Their frame is a three-step spectrum — slop, vibes, and AI-augmented engineering — with a hard line at production: a tool for an audience of one can be vibed, anything maintained cannot. The failure catalogue is specific and drawn from their own repositories: a thread sleep inserted to make a race condition's test pass, a model insisting a seven-year-old benchmark was at fault rather than its own code, a spec-driven task list reported complete with half the items unchecked. Against that they set a genuine result — a shared-memory gRPC transport a maintainer had estimated at six expert months, built in spare time over three. The distinction they draw is sculpting rather than prompting. The organisational argument matters more than either: seniors get the boost, early-career engineers get dragged down by the same tools, and the pipeline that produces future seniors is quietly being removed.

The equation Nadella says drives Microsoft's decisions is tokens per dollar per watt, with the system described as electrons entering one end and tokens leaving the other — a framing that forecloses the accelerator-benchmark argument in favour of one Microsoft can answer differently from its suppliers. Two claims sit beside each other. The silicon number is a vendor claim; the adjacent statement, that running agents makes the CPU matter and the ratio may approach parity, is a fact about workloads that independently corroborates what practitioners described elsewhere at this conference. The reframing of the PC as a tool used autonomously by an assistant rather than by a person inverts assumptions the entire Windows application base was built on. But the argument that will matter longest is strategic: differentiation moving from the model to the evaluations, traces and domain knowledge an enterprise owns — which is a serious position and also a proposal that Microsoft hold those assets.

Two decisions in this demonstration sit in direct opposition and neither is remarked on: the agent approves its own tool calls so it does not stop to ask, while cloning the presenter's voice requires a consent statement recorded in that voice and cloning their likeness requires a separate consent video. Maximum friction to copy a person, zero friction for the agent to act. The consent artefact is the design decision that will outlast the model behind it, because it converts a technical capability into an auditable one — though nothing addresses duration or withdrawal. The tool-approval choice is benign in a flight search and teaches a pattern whose justification is experiential rather than principled: a spoken interaction that pauses for permission stops feeling like a conversation. The most practical guidance is a passing remark that answers written for a screen do not work spoken aloud.

Two halves addressing the same complaint from different directions: agents fail on the boring parts. Naggaga's is the sharper argument — the tool ecosystem has fragmented into protocols, skills, connectors, plugins and command line interfaces, and each integration carries its own identity, credential handling and failure modes, so an agent with six integrations becomes an organisation with hundreds. Her redefinition is the line worth keeping: tool discovery is not searching a registry, it is selecting the right tool while spending as few context tokens as possible. Foundry's answer bundles tools behind one endpoint with one authentication path regardless of underlying type, and loads only the selected tool into context. Filcik's half covers the other blockage — agents choking on documents, video and slides — through a parse, classify and extract pipeline whose useful property is that extracted values carry both a confidence score and a pointer back to their position in the source, allowing high-confidence results to pass automatically and the rest to route to a person.
