Programme topic
Evals & evidence
When output can be fluent and wrong, “it looked good” is not release evidence. These talks show how teams use traces, benchmarks, provenance and evaluation loops to decide whether a change really improved an AI system.
11 published talks
Talks about evals & evidence
From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent
Jiggy KakkadStaff AI Engineer, Quantium
Tinus WillemseExecutive Manager, AI & Data Science, Quantium
Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the…
Building a Security Agent: Model choice, Harnesses and Evals
Simon HarloffCTO, Dam Secure
In this presentation, I’ll show how we set out to give developers useful security feedback on every pull request in under three minutes. The benchmark results and methodology are published here:…
Classifiers are dead. Long live classifiers!
Charli PosnerBuilder, Stile Education
Training a classifier used to mean collecting labelled data, choosing an architecture, training a model, evaluating it, and deploying it. Today, for a surprising number of problems, you can replace most of that with a prompt. At Stile, we’ve been doing exactly…
Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals
Dean SosteSenior Machine Learning Engineer, Canva
Production AI problems are rarely observable through deterministic CI checks. They surface as user feedback, bad LLM-as-a-Judge scores, alerts, or even a vague sense that something is wrong. The evidence that explains them is scattered across traces, prompts,…
Give Every Agent a Flight Recorder
Rahul TrikhaPrincipal AI Engineer, Zendesk
Agent teams should not need to file a ticket with a central evaluation team just to learn whether a new prompt, model, or tool made their agent better. At Zendesk, we developed and deployed a trace-first evaluation platform that gives every agent a flight…
How to Change an LLM System Without Guessing
Yulia KuchinaStaff AI Engineer, Software at Scale
We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide…
Don't Fight Hallucinations. Make Them Impossible
Nadia MakarevichPrincipal Engineer, Heatseeker
The Heatseeker AI chat answers data questions for marketers who make decisions with million-dollar budgets. Wrong answers or hallucinated numbers are not an option here, as you can imagine ;) The fight against them (hallucinations, not marketers) was long and…
Trust is engineered, not granted: why we focus on verifying before background coding agents
Vivek KatialEngineering Lead, Applied AI, Heidi Health
Heidi is an AI scribe used by 130K clinicians a week. Our 150 engineers ship 100+ PRs daily into prod, and AI made writing code so cheap that review became the bottleneck: our P75 review wait was 14 hours, almost all of it queue time. A 14-hour queue is a…
The agents went rogue at 2%
James PeterCo-Founder, JustEvery
Giving an agent a skill sounds straightforward: it loads a Markdown file, follows the instructions and completes the task. Ours launched a long-running, paid process and told the agent to wait for the result. But agents repeatedly started the process, became…
Where Should the Dice Roll? Placing Non-Determinism Deliberately in Enterprise AI
Vighnesh DeshpandeAI Engineer, Vivanti Consulting
Most production AI guidance assumes you want consistency: pin the prompt, lower the temperature, eval for drift. But a whole class of enterprise use cases - idea generation, recommendation, exploration, synthesis - is worthless if the output is predictable.…
From vibes to a systematic eval flywheel: evals for high stakes AI agents
Donna ZhouSoftware engineer, Lorikeet
If you are sceptical about how evals drive real results, or looking for more depth than introductory tutorial videos, this talk is for you. We'll show you how Lorikeet, an AI customer service startup, built an eval system to rapidly raise the quality of a…