Programme topic

Evals & evidence

When output can be fluent and wrong, “it looked good” is not release evidence. These talks show how teams use traces, benchmarks, provenance and evaluation loops to decide whether a change really improved an AI system.

11 published talks

Talks about evals & evidence

From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent

Jiggy Kakkad

Jiggy KakkadStaff AI Engineer, Quantium

Tinus Willemse

Tinus WillemseExecutive Manager, AI & Data Science, Quantium

Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the…

Read more

Classifiers are dead. Long live classifiers!

Charli Posner

Charli PosnerBuilder, Stile Education

Training a classifier used to mean collecting labelled data, choosing an architecture, training a model, evaluating it, and deploying it. Today, for a surprising number of problems, you can replace most of that with a prompt. At Stile, we’ve been doing exactly…

Read more

Give Every Agent a Flight Recorder

Rahul Trikha

Rahul TrikhaPrincipal AI Engineer, Zendesk

Agent teams should not need to file a ticket with a central evaluation team just to learn whether a new prompt, model, or tool made their agent better. At Zendesk, we developed and deployed a trace-first evaluation platform that gives every agent a flight…

Read more

How to Change an LLM System Without Guessing

Yulia Kuchina

Yulia KuchinaStaff AI Engineer, Software at Scale

We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide…

Read more

The agents went rogue at 2%

James Peter

James PeterCo-Founder, JustEvery

Giving an agent a skill sounds straightforward: it loads a Markdown file, follows the instructions and completes the task. Ours launched a long-running, paid process and told the agent to wait for the result. But agents repeatedly started the process, became…

Read more

Explore another topic

See the programme story