AI-assisted analysis. See our editorial policy.
Human editorial review not recorded
The design decision at the centre of this session is stated about halfway through, and it is the one most teams get wrong: rather than assembling a large context and handing it to the agent, extract what actually matters from the conversation (24:56).
That is the difference between a longer window and a memory. One holds more; the other decides what is worth holding.
Why bigger context does not solve this
The instinct when an agent forgets something is to give it more room. That works until it does not, and it fails in a way that is hard to see: an agent with an enormous context is not attending to all of it equally, and the useful fact from four sessions ago competes with every routine exchange since.
Extraction inverts the problem. Instead of preserving everything and hoping the relevant part surfaces, a policy decides at write time what is durable — which converts a retrieval problem into a curation problem.
The trade is real. Curation means something gets discarded, and if the policy is wrong the information is gone rather than merely buried. But an approach that keeps everything has the same failure with more steps, because information that cannot be found is functionally lost too.
The identity gap that blocks production
The most instructive moment is an admission during the demonstration. The identifier for whose memory this is happens to be hardcoded, and the stated goal is that it should be derived from who is actually logged in (52:59).
The fix shown is a token from the identity provider carried with the request, so the system determines the user rather than being told (53:57).
Trivial in a notebook, and the entire security model in production. Memory is per-actor by construction — an identifier for the person and another for the session (39:24) — which means the binding between a real human and their memory record is the boundary that keeps one user's history out of another's context.
Get it wrong and the failure is not a crash. It is an agent that helpfully recalls someone else's details, which is a data breach that presents as good service.
That is why the notebook-to-production transition (52:32) is the important part of this session rather than a closing formality. Everything demonstrated works with a hardcoded identifier. Nothing about it is safe.
The operational numbers worth noting
Two figures set expectations. Retention is configurable up to 365 days (33:30), and creating a memory store takes one to two minutes (33:30).
The retention ceiling is a design constraint people should notice early, because it defines the longest relationship the system can represent. An assistant that resets annually is fine for support and wrong for anything that accumulates across years.
The provisioning time matters for a different reason: it is slow enough that this cannot happen inside a user's first interaction, which pushes memory creation into onboarding rather than into the request path.
The detail about seeding
The demonstration seeds background information about the user directly, with an explicit note that this would ideally be loaded gradually rather than in one block (42:07).
That caveat is doing more work than it appears. A memory populated by bulk import contains facts nobody stated in conversation, and the agent will use them as though they were learned. The gradual version is not just more realistic — it produces a memory whose contents have provenance in things the user actually said.
The verification step that follows — listing the stored records to confirm the extraction actually happened (43:01) — is the right instinct and worth generalising. Memory systems fail silently. The write succeeds, the extraction produces nothing useful, and you find out several sessions later when the agent does not know something it should.
Key numbers
- 365 days
- maximum configurable retention for a memory store, which bounds the longest relationship the system can represent 33:30
Talk chapters
Key takeaways
- 01
Rather than passing a large context to the agent, extract what actually matters — which turns retrieval into curation. 24:56
- 02
The demonstration admits the actor identifier is hardcoded and should instead be derived from who is logged in. 52:59
- 03
A token from the identity provider determines the user, which is the boundary keeping one person's history out of another's context. 53:57
- 04
Retention is configurable up to 365 days, which bounds the longest relationship the system can represent. 33:30
- 05
Seeding background information in bulk is flagged as unrealistic — gradual loading produces a memory whose contents have provenance. 42:07
Entities mentioned
Organizations
Related talks

The practical counterpart to the argument made elsewhere this season that specification is what contains model entropy. Raval and Harris name the failure they are addressing precisely — a prompt-and-pray loop in which working code arrives with no record of what the model assumed, which requirements were fuzzy, what design was chosen or why, leaving nothing to review and nothing to iterate against when a defect surfaces months later. Their answer is three committed markdown artefacts: requirements written in a structured requirements syntax with acceptance criteria attached to each user story, a design document carrying technical decisions together with the reasoning behind them, and a task list whose entries cite the requirement numbers they satisfy. The traceability is the point — a reviewer questioning a decision in a pull request can follow it back through the task to the design to the requirement, all in the same repository. Notably they keep the human between each phase rather than after it, with the agent surfacing ambiguity as questions before proceeding.

Drawn from a year of engagements with more than a hundred companies, this is the most direct challenge in the season's programme to the assumption that faster code generation produces faster delivery. Mishra and Raja open with external evidence rather than their own: an industry study putting realised velocity gains in the ten to fifteen per cent range, and a controlled experiment in which developers using AI estimated themselves roughly a fifth more productive while measurement showed them a fifth slower. Their diagnosis is that both prevailing working styles fail for opposite reasons. Handing an ambiguous problem to an agent and awaiting a finished result produces a volume of code the developer must nonetheless sign for and cannot confidently review, so it stalls before production. The senior engineer's alternative — decomposing the work personally and inserting AI into narrow slots — keeps the intellectual load exactly where it was, and leaves the surrounding process untouched, so hours saved in editing are consumed by the meetings that process still requires.

Brooker builds the definition from the bottom up rather than asserting it, using a deliberately absurd arithmetic task to separate three categories: what a model computes reliably as a fixed function of its input, what merely needs to arrive in the system prompt, and what genuinely requires reaching into the world. Only the third category justifies a tool, and the distinction matters because most production disappointment comes from tools built for the first two. His working definition follows — a system given a goal that loops between inference and tool calls until it reaches one — with the observation that modern agents increasingly embed code in their definitions, not for expressiveness but because replacing inference steps with deterministic code improves reliability while lowering both latency and cost. The remainder covers what production actually demands around that loop: somewhere to run, memory that persists preferences, a gateway to internal and external tools, evaluation, and formal methods applied to policy.

The most concrete attempt this conference season to answer a question the agentic coding sessions mostly leave open: if commit counts and hours saved are the wrong measures, what replaces them? Otto's account is unusually specific about why the obvious alternative fails — summing the small time savings a platform team delivers produces figures exceeding a hundred per cent of a developer's time, and a minute returned is not code in production. Their replacement borrows from Amazon's retail supply chain, where cost to serve measures what it takes to place a package on a doorstep, and applies the same shape to software: total cost divided by units of delivery, with the unit chosen to fit the team. The supporting research is the more quotable finding — across tens of thousands of developers over five years, individual velocity reverts to the team's mean, making team velocity the strongest predictor of both individual output and perceived productivity, which is the empirical case against measuring individuals at all.

The framing statistic is organisational rather than technical: around eighty per cent of organisations expected to have platform engineering teams going into 2026, up from about forty-five per cent a couple of years earlier. The interesting part is the doubling. The problem described is teams solving the same problems separately, producing inconsistency and redundancy — dangerous not because of duplicated effort but because each independent solution has its own security properties, and the organisation's real posture is the weakest rather than the average. The most valuable content is that two financial services organisations went in diametrically opposite directions on workload identity and both are described as working, which implies the choice is determined by context rather than by a general answer. The honest note follows immediately: even with standardised patterns the result remains fragmented.

The line that explains this session comes from the customer in the last ten minutes: they are preparing for a world where metadata is how agent-based systems find the data they need and access it through the controls being built. That relocates a function — governance has spent two decades as compliance activity describing data that people locate by other means, and if agents navigate by the catalogue then the catalogue stops describing the access path and becomes it. An incomplete catalogue is a documentation problem when humans can ask a colleague; an agent has no such workaround. The most honest moment addresses the perennial failure that rules get written and ignored, with enforcement rather than publication as the argument. Generated descriptions and greyed-out classification suggestions divide the labour correctly, keeping a person accountable while removing the burden of finding candidates.
