Great talks. Insightful articles.

Summit talks

Source-backed talks, claims and takeaways from the world's most consequential summits.

PepsiCo's Six-Agent System for Account Managers, and What It Cost to Build41:40
PepsiCo's Six-Agent System for Account Managers, and What It Cost to Build

The rare enterprise session that describes the wiring rather than the outcome. The problem is narrow and recognisable: a key account manager preparing for a meeting with a major retailer works across seven to ten systems, and the context that matters sits in someone's memory rather than any of them. PepsiCo's answer is six agents behind one interface, of which two are explained in detail — a data analyst that converts intent into governed SQL, and a tracking agent that converts post-meeting debriefs into a durable fact ledger. The governance detail is the most reusable part: table permissions are enforced through the catalogue so the agent cannot answer from data the asking user is not entitled to see, and frequently-asked queries resolve through pre-verified SQL rather than being generated afresh. Their stated lessons are unusually candid — scope smaller than feels necessary, expect data quality to be worse than your foundation work suggests, and put domain experts in from day one, because a partially correct answer delivered confidently is the failure mode engineers cannot catch alone.

Krunalkumar Patel / Microsoft Build

Verified source
Where Agentic Coding Actually Breaks: Russinovich and Hanselman at Build 202646:52
Where Agentic Coding Actually Breaks: Russinovich and Hanselman at Build 2026

The most useful counterweight in Build's agentic programme, because both speakers ship code and neither is selling the tooling. Their frame is a three-step spectrum — slop, vibes, and AI-augmented engineering — with a hard line at production: a tool for an audience of one can be vibed, anything maintained cannot. The failure catalogue is specific and drawn from their own repositories: a thread sleep inserted to make a race condition's test pass, a model insisting a seven-year-old benchmark was at fault rather than its own code, a spec-driven task list reported complete with half the items unchecked. Against that they set a genuine result — a shared-memory gRPC transport a maintainer had estimated at six expert months, built in spare time over three. The distinction they draw is sculpting rather than prompting. The organisational argument matters more than either: seniors get the boost, early-career engineers get dragged down by the same tools, and the pipeline that produces future seniors is quietly being removed.

Mark Russinovich / Microsoft Build

Verified source
The Dark Factory Argument: swyx on Agent Supervision at Build 202635:43
The Dark Factory Argument: swyx on Agent Supervision at Build 2026

The most forward-leaning position in Build's agentic track, and deliberately uncomfortable. Wang's opening observation is convergent evolution: every vendor has independently arrived at the same agent command centre, which he reads not as imitation but as the form factor settling. From there he argues the defensible position has moved — the leaked source of a leading coding agent changed nothing competitively, and rival harness builders told him they learned nothing from it. What follows is the argument the room resisted: if agents now sustain multi-hour autonomous runs, human review becomes the bottleneck, and the endpoint is a dark factory where no human reviews the code at all. He does not present this as desirable. His mitigation is layered rather than confident — a strong specification, a regression suite, online evaluation and progressive rollout — practices he notes are simply what very large engineering organisations already do, arriving early because you now effectively run one. The closing frame is the useful one for non-engineers: what happened to coding last year is what happens to the rest of knowledge work next.

Shawn Wang / Microsoft Build

Verified source
Nadella's Argument: Enterprises Stop Consuming the Frontier and Join It142:46
Nadella's Argument: Enterprises Stop Consuming the Frontier and Join It

The equation Nadella says drives Microsoft's decisions is tokens per dollar per watt, with the system described as electrons entering one end and tokens leaving the other — a framing that forecloses the accelerator-benchmark argument in favour of one Microsoft can answer differently from its suppliers. Two claims sit beside each other. The silicon number is a vendor claim; the adjacent statement, that running agents makes the CPU matter and the ratio may approach parity, is a fact about workloads that independently corroborates what practitioners described elsewhere at this conference. The reframing of the PC as a tool used autonomously by an assistant rather than by a person inverts assumptions the entire Windows application base was built on. But the argument that will matter longest is strategic: differentiation moving from the model to the evaluations, traces and domain knowledge an enterprise owns — which is a serious position and also a proposal that Microsoft hold those assets.

Satya Nadella / Microsoft Build

Verified source
Jeff Dean on Why Tools, Not Models, Are the Next Bottleneck (Google I/O 2026)40:56
Jeff Dean on Why Tools, Not Models, Are the Next Bottleneck (Google I/O 2026)

Four of Google's model, product and search leads on what changes once agents run for hours rather than seconds, and the most quotable argument comes from Dean: the constraint is moving out of the model and into the tools around it. By Amdahl's law, an agent spending half its time in tools built for human-speed interaction cannot gain more than a doubling however fast the model becomes — which reframes a great deal of current infrastructure work as latency debt. Their internal response is concrete: rewriting Python tooling into Go, framed as a fully specified translation task rather than an open prompt, produced order-of-magnitude speedups overnight. Reid supplies the counterweight from Search, where acceptable latency turns out to scale with how much work is being taken off the user rather than being a fixed budget. Woodward's detail is the quietest and perhaps the most telling: teams that have stopped writing product documents for humans and now write context files for models to act on directly.

Jeff Dean / Google I/O

Verified source
Jensen Huang's GTC 2026 Keynote: Vera Rubin, the Groq Deal and the Inference Inflection138:56
Jensen Huang's GTC 2026 Keynote: Vera Rubin, the Groq Deal and the Inference Inflection

Jensen Huang used NVIDIA's 2026 GTC keynote to argue that AI has crossed an inference inflection: models that once only generated text now reason and act, and each step multiplies the compute a single task consumes. He put NVIDIA's forward demand visibility above one trillion dollars through 2027, then spent much of the keynote explaining why that is a factory-economics claim rather than a chip claim — a gigawatt of AI factory costs roughly forty billion dollars before any compute is installed, so throughput per watt is what determines revenue. The technical centrepiece was the Vera Rubin platform; the strategic surprise was NVIDIA absorbing the Groq team to cover the low-latency decode that NVLink alone cannot reach. He closed on two extensions of the agentic thesis: OpenClaw as an emerging operating system for agents, hardened for enterprises as NemoClaw, and physical AI, where four new robotaxi partners add roughly eighteen million vehicles a year.

Jensen Huang / NVIDIA GTC / 3/16/2026

Verified source
Huang's Five-Layer Cake: The Infrastructure Argument He Took to Davos35:13
Huang's Five-Layer Cake: The Infrastructure Argument He Took to Davos

Huang brings a diagram to Davos: AI as a five-layer cake running energy, chips, cloud, models, applications — with economic benefit landing at the top and every layer below it a precondition. His argument for why this is a genuine platform shift rather than a product cycle is the strongest part, and it does not rest on his commercial position: software was pre-recorded and worked on structured data, whereas a machine that reasons about unstructured input and inferred intent makes previously impossible applications possible. What the framing accomplishes is worth noticing separately. By presenting the layers as a chain rather than a portfolio, it converts infrastructure spending from a bet into a prerequisite, and the question of proportion between layer-two spending and layer-five value stops being askable. Read against the GTC keynote two months later, the same business gets two framings: one a case for choosing his product, the other a case for the category existing at the scale he needs.

Jensen Huang / World Economic Forum Annual Meeting

Verified source
Visa Spent Eighteen Months Advocating AI Before Anything Changed (Davos 2026)46:11
Visa Spent Eighteen Months Advocating AI Before Anything Changed (Davos 2026)

A show of hands opens the session: nearly everyone has piloted, far fewer have scaled, and everyone who scaled hit problems they did not anticipate. What makes the panel useful is where the four answers do not point. None of the executives — running a healthcare manufacturer, a payments network, an energy producer and a consultancy — blames model capability, cost or data infrastructure. All four describe an organisational constraint. McInerney's account is the sharpest and is an account of failure: eighteen months of executive advocacy and democratised model access produced nothing, until three hundred senior leaders were put in a room for two days and made to build agents themselves. Jakobs supplies the mechanism worth copying, measuring returned clinician time against the three to seven minutes a patient currently receives rather than against cost. Nasser rejects the premise that acquiring compute produces value, and locates returns in operations rather than in the back-office functions most organisations automate first.

Julie Sweet / World Economic Forum Annual Meeting

Verified source
Harari's Question for Davos: Should an AI Be a Legal Person?35:28
Harari's Question for Davos: Should an AI Be a Legal Person?

Harari spends most of his address establishing terms before asking the question he came to ask. His preliminary work is to dismantle the word tool: a knife's use is decided by whoever holds it, whereas what is arriving decides for itself, and can also invent new kinds of knives. From there he argues that anything constituted by words — law, books, text-centred religion — is exposed, while drawing a firm line at feeling, where he says there is no evidence at all. The structural claim is that the ancient tension between letter and spirit has always run inside humanity and is about to be externalised between humans and the new masters of words. Only then does he arrive at personhood, and his handling is precise: corporations, New Zealand rivers and Indian deities hold legal personhood safely because the decisions are made by humans behind the container. An entity that decides for itself ends that arrangement. He does not answer the question; he tells the room it is coming.

Yuval Noah Harari / World Economic Forum Annual Meeting

Verified source
Hassabis and Amodei on the Day After AGI (Davos 2026)32:11
Hassabis and Amodei on the Day After AGI (Davos 2026)

A year after their first joint appearance, the heads of Anthropic and Google DeepMind returned to a shared stage and disagreed mainly about speed. Amodei held to a horizon of one to two years for systems that outperform humans across most cognitive work, resting the claim on a self-improvement loop that runs through code; Hassabis kept to the end of the decade, arguing that verifiable domains like coding and mathematics automate far earlier than natural science, and that the capacity to pose a new question rather than answer an existing one is still missing. The exchange is most useful where they converge: both accept the loop is the variable that decides everything, both are sceptical of doomerism without dismissing the risk, and both want more time than the competitive dynamic allows. Amodei's chip-export argument and Hassabis's call for minimum international safety standards are the two concrete policy asks.

Demis Hassabis / World Economic Forum Annual Meeting

Verified source
Escaping the Prompt-and-Pray Loop: Spec-Driven Development at re:Invent 202560:00
Escaping the Prompt-and-Pray Loop: Spec-Driven Development at re:Invent 2025

The practical counterpart to the argument made elsewhere this season that specification is what contains model entropy. Raval and Harris name the failure they are addressing precisely — a prompt-and-pray loop in which working code arrives with no record of what the model assumed, which requirements were fuzzy, what design was chosen or why, leaving nothing to review and nothing to iterate against when a defect surfaces months later. Their answer is three committed markdown artefacts: requirements written in a structured requirements syntax with acceptance criteria attached to each user story, a design document carrying technical decisions together with the reasoning behind them, and a task list whose entries cite the requirement numbers they satisfy. The traceability is the point — a reviewer questioning a decision in a pull request can follow it back through the task to the design to the requirement, all in the same repository. Notably they keep the human between each phase rather than after it, with the agent surfacing ambiguity as questions before proceeding.

Al Harris / AWS re:Invent

Verified source
The 10-15% Reality Check: Why AI Coding Gains Stay Small (re:Invent 2025)59:22
The 10-15% Reality Check: Why AI Coding Gains Stay Small (re:Invent 2025)

Drawn from a year of engagements with more than a hundred companies, this is the most direct challenge in the season's programme to the assumption that faster code generation produces faster delivery. Mishra and Raja open with external evidence rather than their own: an industry study putting realised velocity gains in the ten to fifteen per cent range, and a controlled experiment in which developers using AI estimated themselves roughly a fifth more productive while measurement showed them a fifth slower. Their diagnosis is that both prevailing working styles fail for opposite reasons. Handing an ambiguous problem to an agent and awaiting a finished result produces a volume of code the developer must nonetheless sign for and cannot confidently review, so it stalls before production. The senior engineer's alternative — decomposing the work personally and inserting AI into narrow slots — keeps the intellectual load exactly where it was, and leaves the surrounding process untouched, so hours saved in editing are consumed by the meetings that process still requires.

Anupam Mishra / AWS re:Invent

Verified source
Amazon's Answer to the Productivity Metrics Problem: Cost to Serve Software41:42
Amazon's Answer to the Productivity Metrics Problem: Cost to Serve Software

The most concrete attempt this conference season to answer a question the agentic coding sessions mostly leave open: if commit counts and hours saved are the wrong measures, what replaces them? Otto's account is unusually specific about why the obvious alternative fails — summing the small time savings a platform team delivers produces figures exceeding a hundred per cent of a developer's time, and a minute returned is not code in production. Their replacement borrows from Amazon's retail supply chain, where cost to serve measures what it takes to place a package on a doorstep, and applies the same shape to software: total cost divided by units of delivery, with the unit chosen to fit the team. The supporting research is the more quotable finding — across tens of thousands of developers over five years, individual velocity reverts to the team's mean, making team velocity the strongest predictor of both individual output and perceived productivity, which is the empirical case against measuring individuals at all.

Bethany Otto / AWS re:Invent

Verified source
What an Agent Actually Is: Marc Brooker on Agent Infrastructure (re:Invent 2025)48:11
What an Agent Actually Is: Marc Brooker on Agent Infrastructure (re:Invent 2025)

Brooker builds the definition from the bottom up rather than asserting it, using a deliberately absurd arithmetic task to separate three categories: what a model computes reliably as a fixed function of its input, what merely needs to arrive in the system prompt, and what genuinely requires reaching into the world. Only the third category justifies a tool, and the distinction matters because most production disappointment comes from tools built for the first two. His working definition follows — a system given a goal that loops between inference and tool calls until it reaches one — with the observation that modern agents increasingly embed code in their definitions, not for expressiveness but because replacing inference steps with deterministic code improves reliability while lowering both latency and cost. The remainder covers what production actually demands around that loop: somewhere to run, memory that persists preferences, a gateway to internal and external tools, evaluation, and formal methods applied to policy.

Marc Brooker / AWS re:Invent

Verified source
Maximum Friction to Copy a Person, Zero Friction to Act as One16:14
Maximum Friction to Copy a Person, Zero Friction to Act as One

Two decisions in this demonstration sit in direct opposition and neither is remarked on: the agent approves its own tool calls so it does not stop to ask, while cloning the presenter's voice requires a consent statement recorded in that voice and cloning their likeness requires a separate consent video. Maximum friction to copy a person, zero friction for the agent to act. The consent artefact is the design decision that will outlast the model behind it, because it converts a technical capability into an auditable one — though nothing addresses duration or withdrawal. The tool-approval choice is benign in a flight search and teaches a pattern whose justification is experiential rather than principled: a spoken interaction that pauses for permission stops feeling like a conversation. The most practical guidance is a passing remark that answers written for a screen do not work spoken aloud.

Henk Boelman / Microsoft Build

Verified source
Tool Sprawl Is the Agent Problem Nobody Priced: Foundry Tools at Build 202643:18
Tool Sprawl Is the Agent Problem Nobody Priced: Foundry Tools at Build 2026

Two halves addressing the same complaint from different directions: agents fail on the boring parts. Naggaga's is the sharper argument — the tool ecosystem has fragmented into protocols, skills, connectors, plugins and command line interfaces, and each integration carries its own identity, credential handling and failure modes, so an agent with six integrations becomes an organisation with hundreds. Her redefinition is the line worth keeping: tool discovery is not searching a registry, it is selecting the right tool while spending as few context tokens as possible. Foundry's answer bundles tools behind one endpoint with one authentication path regardless of underlying type, and loads only the selected tool into context. Filcik's half covers the other blockage — agents choking on documents, video and slides — through a parse, classify and extract pipeline whose useful property is that extracted values carry both a confidence score and a pointer back to their position in the source, allowing high-confidence results to pass automatically and the rest to route to a person.

Maria Naggaga / Microsoft Build

Verified source
An Eleven-Line Agent That Works Half the Time18:28
An Eleven-Line Agent That Works Half the Time

Bennett puts an eleven-line agent on screen and runs it repeatedly: pass, pass, pass, fail, fail, fail. It works about half the time, and nothing in the code says so. That breaks the debugging method conventional software allows, because every test run becomes a sample rather than an observation and a fix appears to work when you happen to draw three passes. His demonstration is the argument: changing the model and rerunning raises the success rate to around ninety per cent with no change to the application at all. The number matters less than the method — he knows the change worked because he measured a rate before and after, and without instrumentation the swap would have been indistinguishable from luck. The uncomfortable implication is that if model choice moves reliability that far with no application change, most published comparisons of agent patterns are reporting noise around a variable they did not control.

Jim Bennett / Microsoft Build

Verified source
Retrieval Built for People Breaks When an Agent Issues Twenty Searches44:57
Retrieval Built for People Breaks When an Agent Issues Twenty Searches

The organising observation comes from watching coding agents: they are remarkably good at local file access, navigating a repository and forming an understanding, because everything is local and cheap to read. The limit is what happens when knowledge lives in systems an agent cannot walk at volumes it cannot read. What follows is a bottleneck of rate rather than accuracy — a person issues a query, reads, refines and repeats a handful of times, while an agent may issue ten searches or several rounds of twenty because asking costs nothing and it is exploring rather than looking something up. Latency budgets calibrated to someone waiting for a page become dominant when multiplied twentyfold inside one task. The infrastructure argument generalises: agentic load is unpredictable in a way application load is not, so paying per use sidesteps a capacity decision nobody has the information to make.

Paulo / Microsoft Build

Verified source
Use the Expensive Model to Plan, the Cheap One to Build9:34
Use the Expensive Model to Plan, the Cheap One to Build

The recommendation at the end is the most immediately usable advice from this conference: use the larger model for planning and a cheaper automatic selection for implementation, based on the team analysing what their own conversations actually cost. The expensive model earns its price where a wrong decision propagates, and stops earning it once the plan is settled — a finer distinction than per-task selection and a larger saving. The candid moment is worth more than the feature. Context switching between agents is described as an unsolved problem, visible in user testing and in the team's own experience, and it burns you out. That is the cost nobody prices when demonstrating parallel agents: six concurrent sessions produce six streams of work in different states, each requiring reconstruction before you can usefully intervene, and human working memory does not multiply.

Harald Kirschner / Microsoft Build

Verified source
Four Agent Workflows Compared Live: Multi-Agent Patterns at Build 202646:44
Four Agent Workflows Compared Live: Multi-Agent Patterns at Build 2026

Structured as a timed competition rather than a talk, which turns out to be its value: four engineers build the same collaborative markdown editor in parallel using four different agent surfaces, and their divergent methods are visible rather than described. Kirschner opens with a research agent surveying existing products, then three parallel design explorations, before writing any code. Reddington splits roles across models, using one as planner and another as implementer. Kasper runs a single high-reasoning pass from a generated specification file, then switches models when the first one's interface work degrades. Running underneath is Dodds coaching the host through the same problem, and his method is the most transferable: he does not write plans, he holds a conversation, deliberately asking questions whose answers he already knows so the agent accumulates the architectural context before being told to proceed. The safety framing is worth noting too — they run in a hosted development container specifically so the agent can be given blanket permission without exposing local credentials.

Kent C. Dodds / Microsoft Build

Verified source