# AI engineering — AI Engineer Sydney 2026

> AI engineering is the work of turning models into dependable systems. This collection covers context, tools, evaluation, observability, infrastructure and the production decisions between a promising model and a useful product.

Canonical page: https://webdirections.org/ai-engineer/topics/ai-engineering/
Program status: The speaker lineup and talk descriptions are public. Session days, times, rooms and the full timetable have not yet been published.

## Published talks

- [The Elephant and the Goldfish: Architecture Patterns for Cutting 70% of Agent Token Costs in Production](https://webdirections.org/ai-engineer/speakers/tanya-dixit/) — Tanya Dixit
  LLM providers sell you a 2-million-token context window like an elephant that never forgets. If you actually build production agents that way, your latency explodes, your retrieval drifts, and your CFO will shut you down in 90 days. In production, the best agents think like elephants, but operate like goldfish. Details: While frontier models offer…
- [From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent](https://webdirections.org/ai-engineer/speakers/jiggy-kakkad/) — Jiggy Kakkad & Tinus Willemse
  Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the chart renders, the prose is fluent, and an incorrect figure reaches a decision-maker unchallenged.…
- [The guess never becomes a fact - provenance-governed, model-free memory for LLMs](https://webdirections.org/ai-engineer/speakers/kexuan-xin/) — Jade Xin
  Long context solves within-session coherence, not cross-session persistence. Once an assistant starts storing memories, a quieter failure appears: its own inferences can be written back, retrieved in later sessions, and presented as user facts. Hallucinations then compound over time. We encountered this while building EDN, TensorPRO's external memory…
- [Office-Brain and the Night Shift: What Happens When You Turn Off the Meter](https://webdirections.org/ai-engineer/speakers/mark-pesce/) — Mark Pesce
  "Good enough" AI now runs on a reasonably beefy laptop. Qwen3.8-27B benchmarks at an Intelligence Index of 52 - roughly the best money could buy back in February - and inferences happily on consumer kit. When tokens are minted on your own machines, the meter goes away, and a different way of working becomes economic: batch processing returns. I call it the…
- [Building Agentic Memory for an AI SOC: Why Our Semantic Cache Was the Wrong Answer](https://webdirections.org/ai-engineer/speakers/mukesh-singh/) — Mukesh Singh
  We run a fleet of AI agents against production security detections at Atlassian. One tunes noisy detection rules, another reviews new ones, and more are coming for alert triage and hunting. Every one of them started blind and stateless, re-deriving the same context from Jira tickets, a Databricks lake, a rule repo and ATT&CK, or drowning in a raw dump of all…
- [Building a Security Agent: Model choice, Harnesses and Evals](https://webdirections.org/ai-engineer/speakers/simon-harloff/) — Simon Harloff
  In this presentation, I’ll show how we set out to give developers useful security feedback on every pull request in under three minutes. The benchmark results and methodology are published here: https://docs.damsecure.ai/blog/pr-review-security-benchmark-update/. It has become our most-cited research to date. We initially expected to compare models using…
- [Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals](https://webdirections.org/ai-engineer/speakers/dean-soste/) — Dean Soste
  Production AI problems are rarely observable through deterministic CI checks. They surface as user feedback, bad LLM-as-a-Judge scores, alerts, or even a vague sense that something is wrong. The evidence that explains them is scattered across traces, prompts, content, documentation, support tickets and logs. At Canva, we treated that evidence as a graph and…
- [Query before mutation: guardrail patterns for agents that touch production infrastructure](https://webdirections.org/ai-engineer/speakers/jeffrey-aven/) — Jeffrey Aven
  Covered in the session: - Why plan-and-apply was always a human safety mechanism: a person reads the plan, notices something is off, and stops the apply - and why this quietly disappears the moment an agent is the operator - Query before mutation as the foundational pattern: requiring agents to establish current state from the live environment before any…
- [Give Every Agent a Flight Recorder](https://webdirections.org/ai-engineer/speakers/rahul-trikha/) — Rahul Trikha
  Agent teams should not need to file a ticket with a central evaluation team just to learn whether a new prompt, model, or tool made their agent better. At Zendesk, we developed and deployed a trace-first evaluation platform that gives every agent a flight recorder: a versioned, safe trail from execution to release decision. The stakes are real. Our agents…
- [How to Change an LLM System Without Guessing](https://webdirections.org/ai-engineer/speakers/yulia-kuchina/) — Yulia Kuchina
  We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide their regressions. This talk is how we went from operating on faith to changing the system on…
- [Don't Fight Hallucinations. Make Them Impossible](https://webdirections.org/ai-engineer/speakers/nadia-makarevich/) — Nadia Makarevich
  The Heatseeker AI chat answers data questions for marketers who make decisions with million-dollar budgets. Wrong answers or hallucinated numbers are not an option here, as you can imagine ;) The fight against them (hallucinations, not marketers) was long and painful. We started with a "naive" approach, which we all tried at some point, I imagine: "Hey, AI,…
- [Running 100K+ Concurrent Agents: It's Never the Prompt](https://webdirections.org/ai-engineer/speakers/harry-nguyen/) — Harry Nguyen
  Your Agents Aren't Failing at Prompting Every agent demo works. Then you run ten thousand at once, and the failures have nothing to do with the model: dead tuples piling up, a sync HTTP client freezing the event loop, requests hanging forever with no total timeout. This talk is a look inside how we build and run agents at massive scale in production. It…
- [Adaptive Learning: Encoding Clinician Edits as Memory](https://webdirections.org/ai-engineer/speakers/vlad-gavrilov/) — Vlad Gavrilov
  Clinicians edit their AI-generated notes for many reasons: adding or removing content, fixing spelling, and preferences for structure, wording, or terminology. Some of these edits are contextual, applying in some situations but not others. The data is noisy, but inside it is useful signal that can be used to predict and make these edits before a clinician…
- [Ontologies: AI’s Operating Manual For Your Business](https://webdirections.org/ai-engineer/speakers/gareth-williams/) — Gareth Williams
  Tools churn. Factories commoditise. When everyone has the same models, the advantage goes to whoever gives their agents the clearest blueprint of how the business works - what exists, how it connects and the rules it runs on. That blueprint is an ontology. In data.world's benchmark, an LLM querying enterprise SQL directly answered 16% of questions correctly.…
- [The BAML Programming Language](https://webdirections.org/ai-engineer/speakers/vaibhav-gupta/) — Vaibhav Gupta
  Whether you like it or not, it’s no longer possible to read all the code that’s generated by a model. The solution cannot be, "Just ship it," or you end up with slop everywhere. Neither can the solution be "Read everything," because that’s impractical and doesn’t take advantage of one of the greatest inventions in history. Just as TypeScript enabled…
- [Classifiers are dead. Long live classifiers!](https://webdirections.org/ai-engineer/speakers/charli-posner/) — Charli Posner
  Training a classifier used to mean collecting labelled data, choosing an architecture, training a model, evaluating it, and deploying it. Today, for a surprising number of problems, you can replace most of that with a prompt. At Stile, we’ve been doing exactly that: using multimodal generative models to classify handwritten marks on scanned worksheets. In…
- [Don't stop, won't stop. Automated verification of long-running agentic loops](https://webdirections.org/ai-engineer/speakers/aj-fisher/) — AJ Fisher
  Long-running agentic loops create different engineering problems than code generation. Once agents work unattended for many hours, iterate through 20 or more review cycles, and build stacked PRs towards a larger goal, the challenge changes to keeping that activity aligned, verifiable and useful. This talk looks at the verification systems built around…
- [Where Should the Dice Roll? Placing Non-Determinism Deliberately in Enterprise AI](https://webdirections.org/ai-engineer/speakers/vighnesh-deshpande/) — Vighnesh Deshpande
  Most production AI guidance assumes you want consistency: pin the prompt, lower the temperature, eval for drift. But a whole class of enterprise use cases - idea generation, recommendation, exploration, synthesis - is worthless if the output is predictable. The engineering problem flips: how do you make variance useful, keep it safe, and convince a…
- [The Benchmark Ends. The World Doesn’t: Building Persistent Engineering Environments for Continual Learning](https://webdirections.org/ai-engineer/speakers/theodoros-galanos/) — Theodoros Galanos
  Large language models can complete sequences of tasks—but continual learning is not just doing more tasks. Most agent environments reset after every episode: state disappears, required follow-up vanishes, delayed effects are cut off, and the next task arrives as though the previous one never happened. Real engineering projects do not reset. They deal in…
- [Extracting Tacit Knowledge for Production Agents](https://webdirections.org/ai-engineer/speakers/khali-kalpa-young/) — Khali Kalpa-Young
  An eval is a definition of "good" written down well enough for a machine to grade against. In a real business, "good" isn't written down anywhere. It lives in someone's reaction, in what past work actually did, and in how a person talks. So the eval comes second. First you get "good" out of wherever it lives, and each place needs its own method. The first…
- [We Deleted Most of Our Agents. Everything Got Faster](https://webdirections.org/ai-engineer/speakers/khang-nguyen-hoang/) — Khang Nguyen Hoang
  The default advice is to add agents: specialised roles, an orchestrator, handoffs between them. We built that. It was a genuinely useful way to explore the problem space, and under real production traffic it was slow and expensive. When an agent system is slow, the instinct is to reach for a faster model. That was the wrong lever. Our latency wasn't…
- [Does This Agent Make My Context Look Big? Right-Sizing AI Architectures for Production](https://webdirections.org/ai-engineer/speakers/hamish-songsmith/) — Hamish Songsmith
  All-in-one personal agent harnesses showcase the incredible potential of capability-rich AI assistants. But deploying a monolithic "do-it-all" agent into production often leaves teams struggling with context dilution, fragile tool calls, un-evaluable execution paths, and massive security blast radiuses. However, swinging to the opposite extreme—decomposing…

The timetable is not public. This page does not imply a day, time, room or track.

- [Explore the whole programme](https://webdirections.org/ai-engineer/program/)
- [Conference overview](https://webdirections.org/ai-engineer/index.md)
