议题

Physical AI

AI systems that perceive and act in the physical world — robotics, autonomous vehicles and the simulation stacks used to train them.

7
演讲
11
嘉宾
8
机构

最新演讲

Cost and Offline Are Optimisations; Data Residency Is a Wall
Cost and Offline Are Optimisations; Data Residency Is a Wall

The case for local inference is made in three clauses representing different kinds of constraint: cost, where a local model removes an API call; availability, where the application works on a flight; and data, where requirements prevent information leaving the device. Only the third changes what is buildable rather than what is affordable. What makes this newly practical is unglamorous — instruction set extensions integrated into the runtime rather than a modelling breakthrough — with around thirty per cent improvement reported in image editing functions. The guidance on fine-tuning inverts the usual advice: adaptation matters most for the smallest models, because their capability budget is already spent and getting them to perform on your problem means spending some of it there. The two examples do the real work, since neither is a cheaper version of a cloud application.

Google I/O

Physics, Not Pixels: What an Embodied Reasoning Model Changes
Physics, Not Pixels: What an Embodied Reasoning Model Changes

The distinction Reese draws early is the one that matters: an embodied reasoning model is not a vision model bolted to a robot but the logic unit of the system, fine-tuned on robotics data for spatial understanding, and reasoning about a scene's physics rather than its pixels. That collapses the seam between perception and planning where most traditional robotics failures lived, because the planner no longer receives categories with everything uncategorisable discarded. The browser-based demonstration carries an argument about access as much as capability, since robotics has been gated on hardware and a physics engine in a browser moves the constraint from equipment to ideas. The session's sharpest moment is its last: a model that hallucinates in software produces a strange recipe, and the same error rate attached to something exerting force is a different category of event — a gap of orders of magnitude, not an increment.

Google I/O

"Pick Up the Extinct Animal": Where Robotics Actually Stands
"Pick Up the Extinct Animal": Where Robotics Actually Stands

The anecdote that opens the panel does the work: a robot asked to pick up the extinct animal selected a dinosaur toy, with nothing in its training data connecting the phrase to the object. That transfer from language models into machines with hands is the premise of the current wave. What the practitioners then describe is where it stops. Physical intelligence is about exerting force and using a body to do it, which is knowledge about consequences — the one thing a corpus of internet images contains almost nothing about. The humanoid question gets an honest treatment: not that human shape is optimal, but that the world is already built for it, plus a development-loop argument about collecting data and deploying on the same hardware. The most useful passage is scepticism about the field's favourite shortcut: generated video looks realistic and does not hold up for dexterous manipulation, because looking right and being physically consistent are different properties.

Google I/O

Demis Hassabis on AGI by 2030 and AI's Frontiers in Science (Google I/O 2026)
Demis Hassabis on AGI by 2030 and AI's Frontiers in Science (Google I/O 2026)

Four months after Davos, Hassabis put a sharper number on the same forecast: AGI around 2030, give or take a year, arriving gradually rather than as a single moment. His test for it is concrete — a model with a 1901 knowledge cutoff that could produce Einstein's 1905 insights — and by that standard current systems plainly fail. The interview is more useful than the Davos panel on two fronts. First, competitive position: he argues Google's advantage is being the only organisation holding the full stack from chips to billion-user products, citing 900 million monthly users on the Gemini app. Second, method: the AlphaFold story of choosing to fold every known protein at once rather than run a request service is his working example of what acceleration should look like. He closes on a warning aimed at the Bay Area — that direction matters more than velocity, and that the current frenetic pace is not conducive to the deep work the next advances require.

Google I/O

Jensen Huang's GTC 2026 Keynote: Vera Rubin, the Groq Deal and the Inference Inflection
Jensen Huang's GTC 2026 Keynote: Vera Rubin, the Groq Deal and the Inference Inflection

Jensen Huang used NVIDIA's 2026 GTC keynote to argue that AI has crossed an inference inflection: models that once only generated text now reason and act, and each step multiplies the compute a single task consumes. He put NVIDIA's forward demand visibility above one trillion dollars through 2027, then spent much of the keynote explaining why that is a factory-economics claim rather than a chip claim — a gigawatt of AI factory costs roughly forty billion dollars before any compute is installed, so throughput per watt is what determines revenue. The technical centrepiece was the Vera Rubin platform; the strategic surprise was NVIDIA absorbing the Groq team to cover the low-latency decode that NVLink alone cannot reach. He closed on two extensions of the agentic thesis: OpenClaw as an emerging operating system for agents, hardened for enterprises as NemoClaw, and physical AI, where four new robotaxi partners add roughly eighteen million vehicles a year.

NVIDIA GTC

More Chips Than We Can Switch On: Musk's One Falsifiable Claim at Davos
More Chips Than We Can Switch On: Musk's One Falsifiable Claim at Davos

Musk names his constraint without hedging: AI deployment is limited by electrical power, with chip production rising exponentially against electricity growing at three to four per cent a year, and a crossover he expects within the year where more chips are made than can be switched on. He then names the exception — China, building nuclear at scale and deploying solar at an order of magnitude beyond everyone else — and the room moves on, though placed against his own framing it is the most consequential thing said. The remainder describes a world where the constraint is solved: robots building robots until human wants saturate, humanoid units on sale to the public within roughly two years, and systems exceeding collective human capability around 2030. One claim deserved scrutiny it did not receive — that orbital compute becomes cheapest within three years — because it contradicts his own position that inference must sit near users.

World Economic Forum Annual Meeting

You Cannot Tell Who Owns the Tractor
You Cannot Tell Who Owns the Tractor

The hardest problem in this session has nothing to do with machine learning: you cannot reliably tell who owns a machine. Unlike vehicles, which carry an identification number and go through state registration, heavy equipment has no equivalent — someone can simply assert ownership. Everything the connected-product strategy promises depends on solving that, because every step after fault detection requires knowing who to contact. The estate explains why it was not solved earlier: millions of machines with 1.5 million connected, and around 160 dealers who are independent businesses with their own systems, holding the service history that makes telemetry meaningful. The prior state is described directly — multiple accumulated platforms, and dealers confused because the same question returned different answers, which destroys trust in all of them including the correct ones.

AWS re:Invent

如何引用本页

复制这份有来源支持的实体档案的稳定引用。