# Evals & evidence — AI Engineer Sydney 2026

> When output can be fluent and wrong, “it looked good” is not release evidence. These talks show how teams use traces, benchmarks, provenance and evaluation loops to decide whether a change really improved an AI system.

Canonical page: https://webdirections.org/ai-engineer/topics/evals-evidence/
Program status: The speaker lineup and talk descriptions are public. Session days, times, rooms and the full timetable have not yet been published.

## Published talks

- [From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent](https://webdirections.org/ai-engineer/speakers/jiggy-kakkad/) — Jiggy Kakkad & Tinus Willemse
  Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the chart renders, the prose is fluent, and an incorrect figure reaches a decision-maker unchallenged.…
- [Building a Security Agent: Model choice, Harnesses and Evals](https://webdirections.org/ai-engineer/speakers/simon-harloff/) — Simon Harloff
  In this presentation, I’ll show how we set out to give developers useful security feedback on every pull request in under three minutes. The benchmark results and methodology are published here: https://docs.damsecure.ai/blog/pr-review-security-benchmark-update/. It has become our most-cited research to date. We initially expected to compare models using…
- [Classifiers are dead. Long live classifiers!](https://webdirections.org/ai-engineer/speakers/charli-posner/) — Charli Posner
  Training a classifier used to mean collecting labelled data, choosing an architecture, training a model, evaluating it, and deploying it. Today, for a surprising number of problems, you can replace most of that with a prompt. At Stile, we’ve been doing exactly that: using multimodal generative models to classify handwritten marks on scanned worksheets. In…
- [Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals](https://webdirections.org/ai-engineer/speakers/dean-soste/) — Dean Soste
  Production AI problems are rarely observable through deterministic CI checks. They surface as user feedback, bad LLM-as-a-Judge scores, alerts, or even a vague sense that something is wrong. The evidence that explains them is scattered across traces, prompts, content, documentation, support tickets and logs. At Canva, we treated that evidence as a graph and…
- [Give Every Agent a Flight Recorder](https://webdirections.org/ai-engineer/speakers/rahul-trikha/) — Rahul Trikha
  Agent teams should not need to file a ticket with a central evaluation team just to learn whether a new prompt, model, or tool made their agent better. At Zendesk, we developed and deployed a trace-first evaluation platform that gives every agent a flight recorder: a versioned, safe trail from execution to release decision. The stakes are real. Our agents…
- [How to Change an LLM System Without Guessing](https://webdirections.org/ai-engineer/speakers/yulia-kuchina/) — Yulia Kuchina
  We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide their regressions. This talk is how we went from operating on faith to changing the system on…
- [Don't Fight Hallucinations. Make Them Impossible](https://webdirections.org/ai-engineer/speakers/nadia-makarevich/) — Nadia Makarevich
  The Heatseeker AI chat answers data questions for marketers who make decisions with million-dollar budgets. Wrong answers or hallucinated numbers are not an option here, as you can imagine ;) The fight against them (hallucinations, not marketers) was long and painful. We started with a "naive" approach, which we all tried at some point, I imagine: "Hey, AI,…
- [Trust is engineered, not granted: why we focus on verifying before background coding agents](https://webdirections.org/ai-engineer/speakers/vivek-katial/) — Vivek Katial
  Heidi is an AI scribe used by 130K clinicians a week. Our 150 engineers ship 100+ PRs daily into prod, and AI made writing code so cheap that review became the bottleneck: our P75 review wait was 14 hours, almost all of it queue time. A 14-hour queue is a reliability problem — it batches changes, delays fixes, and pushes people toward the "just approve it"…
- [The agents went rogue at 2%](https://webdirections.org/ai-engineer/speakers/james-peter/) — James Peter
  Giving an agent a skill sounds straightforward: it loads a Markdown file, follows the instructions and completes the task. Ours launched a long-running, paid process and told the agent to wait for the result. But agents repeatedly started the process, became sidetracked and rebuilt the output themselves. In one test, only 1 of 16 runs produced a fully…
- [Where Should the Dice Roll? Placing Non-Determinism Deliberately in Enterprise AI](https://webdirections.org/ai-engineer/speakers/vighnesh-deshpande/) — Vighnesh Deshpande
  Most production AI guidance assumes you want consistency: pin the prompt, lower the temperature, eval for drift. But a whole class of enterprise use cases - idea generation, recommendation, exploration, synthesis - is worthless if the output is predictable. The engineering problem flips: how do you make variance useful, keep it safe, and convince a…
- [From vibes to a systematic eval flywheel: evals for high stakes AI agents](https://webdirections.org/ai-engineer/speakers/donna-zhou/) — Donna Zhou
  If you are sceptical about how evals drive real results, or looking for more depth than introductory tutorial videos, this talk is for you. We'll show you how Lorikeet, an AI customer service startup, built an eval system to rapidly raise the quality of a complex AI agent. Coach is our AI assistant that helps customers configure their AI concierges that…

The timetable is not public. This page does not imply a day, time, room or track.

- [Explore the whole programme](https://webdirections.org/ai-engineer/program/)
- [Conference overview](https://webdirections.org/ai-engineer/index.md)
