
Buried near the end is the most useful sentence in the session: three agentic products are in production, one is about to launch, and one was taken back to the drawing board — and that last one produced some of the most valuable data the team got. The technical argument builds toward verification, starting from a limitation rather than a capability: traditional testing only goes so far because these models are probabilistic, which quietly invalidates most of an enterprise QA apparatus. Their answer is to measure properties rather than check outputs, tracking relevance, completeness and tone while noting other organisations will need different measures. The distinction between hard and soft guardrails clarifies the design question of how much safety requirement can be pushed into a deterministic layer, and their red-teaming runs as a schedule rather than a gate.
