AWS re:Invent 2025

Everything Here Is About Moving a Kernel Across One Line

原演讲者: Scott Perry, Solutions Architect, Annapurna Labs · Amazon Web Services

来源已核验演讲日期待核实workshop43:59EN3 分钟阅读

A compiler cannot decide whether your workload is memory-bound or compute-bound because that depends on your problem's shape, which is the entire reason a low-level kernel language exists.

Everything in this session is downstream of one diagram. A given accelerator has finite memory bandwidth and finite compute throughput, and where a workload sits between those two limits is decided by how many operations it performs per byte read from memory (4:33).

Land on the memory-bound side and the expensive compute units idle while data arrives. Land on the compute-bound side and you are getting what you paid for. The stated goal is to push workloads rightward, into the region where the arithmetic per byte is high enough to keep the engines busy (5:00).

Every technique that follows exists to move a kernel across that line.

Why a low-level language exists at all

The natural objection is that a compiler should handle this. The session's implicit answer is that it cannot, because the decision depends on the shape of your specific problem.

The chip described has four compute engines usable in parallel, plus on-chip memories close to the cores (2:45). Getting value from that arrangement means deciding what to keep resident, what to stream, and how to overlap the two — decisions that depend on tensor dimensions the compiler sees but whose relative importance it cannot infer.

The techniques named are the standard ones: pipelining, and splitting a large workload into chunks so multiple parts run concurrently and the engines stay busier (5:30). Neither is exotic. Both require knowing which stage is the bottleneck, which is precisely the knowledge a general compiler lacks.

The shape of every kernel

The structure they demonstrate repeats identically regardless of the operation: allocate space on the on-chip memory, copy the inputs down from high-bandwidth memory, compute, allocate space for the result, and copy back (11:26).

That pattern is the whole abstraction. Once you see it, the programming model stops being mysterious — you are writing an explicit data-movement schedule and attaching arithmetic to it.

They are candid that this feels low-level to anyone accustomed to a high-level framework, and justify it as what you do when you want the best return from the hardware (12:48). That is the correct framing, and it comes with an unstated cost: this code is tied to this memory hierarchy. A different accelerator with different on-chip capacity needs the schedule rewritten, not recompiled.

The mitigation is that the language integrates with the mainstream frameworks and emits hardware instructions directly (8:16), so the low-level work can be confined to the kernels that matter while everything else stays portable. Which is the right architecture for this kind of optimisation: make the expensive-to-write part small.

The methodology worth stealing

Two practices in the walkthrough are more valuable than the specific technology.

The first is scope reduction for measurement. They optimise a handful of layers rather than a full model, on the argument that performance should be broadly representative and scales accordingly (15:02). That turns an iteration cycle from an overnight job into something you can run repeatedly in an afternoon, which changes how many ideas you can test.

The second is verifying numerical equivalence rather than assuming it. After compiling for the accelerator they compare against a reference implementation using error and similarity measures (17:23) — a check that catches the failure mode specific to this work, where an optimised kernel is fast and quietly wrong.

That failure mode is the reason the practice matters. A performance regression announces itself. A numerical one produces a model that trains slightly worse for reasons no one can locate.

Reading the trace

The profiling step (18:18) is where the roofline stops being theory. A timeline showing what each engine was doing tells you directly whether you are waiting on memory or on arithmetic, which is the only fact that determines what to optimise next.

The reason attention blocks are singled out as the usual target (17:51) is visible in the same framing. Attention moves large amounts of data relative to the arithmetic it performs, which places it on the wrong side of the line by default — and that is why so much engineering effort across the field has gone into exactly this one operation.

The through-line is unglamorous and correct: measure where the time goes, find the engine that is idle, and rewrite the data movement until it is not. The tooling is new. The discipline is the same one performance engineers have used for forty years.

关键数据

4 engines
compute engines usable in parallel on the accelerator chip discussed 2:45

演讲章节

关键要点

  1. 01

    Where a workload sits between finite memory bandwidth and finite compute throughput is set by operations performed per byte read. 4:33

  2. 02

    The goal of every technique shown is pushing a kernel into the compute-bound region so the engines stay busy rather than waiting on memory. 5:00

  3. 03

    Every kernel has the same shape: allocate on-chip memory, copy inputs down, compute, copy the result back — an explicit data-movement schedule. 11:26

  4. 04

    They measure on a handful of layers rather than a full model, which turns an overnight iteration cycle into an afternoon one. 15:02

  5. 05

    Numerical equivalence against a reference is verified explicitly, because an optimised kernel can be fast and quietly wrong. 17:23

提及的实体

相关演讲

Escaping the Prompt-and-Pray Loop: Spec-Driven Development at re:Invent 2025
Escaping the Prompt-and-Pray Loop: Spec-Driven Development at re:Invent 2025

The practical counterpart to the argument made elsewhere this season that specification is what contains model entropy. Raval and Harris name the failure they are addressing precisely — a prompt-and-pray loop in which working code arrives with no record of what the model assumed, which requirements were fuzzy, what design was chosen or why, leaving nothing to review and nothing to iterate against when a defect surfaces months later. Their answer is three committed markdown artefacts: requirements written in a structured requirements syntax with acceptance criteria attached to each user story, a design document carrying technical decisions together with the reasoning behind them, and a task list whose entries cite the requirement numbers they satisfy. The traceability is the point — a reviewer questioning a decision in a pull request can follow it back through the task to the design to the requirement, all in the same repository. Notably they keep the human between each phase rather than after it, with the agent surfacing ambiguity as questions before proceeding.

presentation

The 10-15% Reality Check: Why AI Coding Gains Stay Small (re:Invent 2025)
The 10-15% Reality Check: Why AI Coding Gains Stay Small (re:Invent 2025)

Drawn from a year of engagements with more than a hundred companies, this is the most direct challenge in the season's programme to the assumption that faster code generation produces faster delivery. Mishra and Raja open with external evidence rather than their own: an industry study putting realised velocity gains in the ten to fifteen per cent range, and a controlled experiment in which developers using AI estimated themselves roughly a fifth more productive while measurement showed them a fifth slower. Their diagnosis is that both prevailing working styles fail for opposite reasons. Handing an ambiguous problem to an agent and awaiting a finished result produces a volume of code the developer must nonetheless sign for and cannot confidently review, so it stalls before production. The senior engineer's alternative — decomposing the work personally and inserting AI into narrow slots — keeps the intellectual load exactly where it was, and leaves the surrounding process untouched, so hours saved in editing are consumed by the meetings that process still requires.

presentation

What an Agent Actually Is: Marc Brooker on Agent Infrastructure (re:Invent 2025)
What an Agent Actually Is: Marc Brooker on Agent Infrastructure (re:Invent 2025)

Brooker builds the definition from the bottom up rather than asserting it, using a deliberately absurd arithmetic task to separate three categories: what a model computes reliably as a fixed function of its input, what merely needs to arrive in the system prompt, and what genuinely requires reaching into the world. Only the third category justifies a tool, and the distinction matters because most production disappointment comes from tools built for the first two. His working definition follows — a system given a goal that loops between inference and tool calls until it reaches one — with the observation that modern agents increasingly embed code in their definitions, not for expressiveness but because replacing inference steps with deterministic code improves reliability while lowering both latency and cost. The remainder covers what production actually demands around that loop: somewhere to run, memory that persists preferences, a gateway to internal and external tools, evaluation, and formal methods applied to policy.

presentation

Amazon's Answer to the Productivity Metrics Problem: Cost to Serve Software
Amazon's Answer to the Productivity Metrics Problem: Cost to Serve Software

The most concrete attempt this conference season to answer a question the agentic coding sessions mostly leave open: if commit counts and hours saved are the wrong measures, what replaces them? Otto's account is unusually specific about why the obvious alternative fails — summing the small time savings a platform team delivers produces figures exceeding a hundred per cent of a developer's time, and a minute returned is not code in production. Their replacement borrows from Amazon's retail supply chain, where cost to serve measures what it takes to place a package on a doorstep, and applies the same shape to software: total cost divided by units of delivery, with the unit chosen to fit the team. The supporting research is the more quotable finding — across tens of thousands of developers over five years, individual velocity reverts to the team's mean, making team velocity the strongest predictor of both individual output and perceived productivity, which is the empirical case against measuring individuals at all.

presentation

Two Banks Went Opposite Directions on Identity, and Both Worked
Two Banks Went Opposite Directions on Identity, and Both Worked

The framing statistic is organisational rather than technical: around eighty per cent of organisations expected to have platform engineering teams going into 2026, up from about forty-five per cent a couple of years earlier. The interesting part is the doubling. The problem described is teams solving the same problems separately, producing inconsistency and redundancy — dangerous not because of duplicated effort but because each independent solution has its own security properties, and the organisation's real posture is the weakest rather than the average. The most valuable content is that two financial services organisations went in diametrically opposite directions on workload identity and both are described as working, which implies the choice is determined by context rather than by a general answer. The honest note follows immediately: even with standardised patterns the result remains fragmented.

presentation

When Metadata Stops Describing the Access Path and Becomes It
When Metadata Stops Describing the Access Path and Becomes It

The line that explains this session comes from the customer in the last ten minutes: they are preparing for a world where metadata is how agent-based systems find the data they need and access it through the controls being built. That relocates a function — governance has spent two decades as compliance activity describing data that people locate by other means, and if agents navigate by the catalogue then the catalogue stops describing the access path and becomes it. An incomplete catalogue is a documentation problem when humans can ask a colleague; an agent has no such workaround. The most honest moment addresses the perennial failure that rules get written and ignored, with enforcement rather than publication as the argument. Generated descriptions and greyed-out classification suggestions divide the labour correctly, keeping a person accountable while removing the burden of finding candidates.

session