Explore the programme
What will you learn at AI Engineer Sydney?
The programme follows the decisions teams face when AI moves from experiment to production: how to build capable systems, how to know they can be trusted, and how software work changes when agents become part of the team.
Find your way in
Start with a topic
Each topic opens a focused collection of published talks. Talks can belong to more than one topic because the useful questions rarely stay in neat boxes.
01
Build agents that work in production
Move beyond the demo: engineer context, memory, tools, interfaces and infrastructure for agents that must run reliably, economically and at scale.
The Elephant and the Goldfish: Architecture Patterns for Cutting 70% of Agent Token Costs in Production
Tanya DixitForward Deployed Engineer, Google
LLM providers sell you a 2-million-token context window like an elephant that never forgets. If you actually build production agents that way, your latency explodes, your retrieval drifts, and your CFO will shut you down in 90 days. In production, the best…
From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent
Jiggy KakkadStaff AI Engineer, Quantium
Tinus WillemseExecutive Manager, AI & Data Science, Quantium
Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the…
The guess never becomes a fact - provenance-governed, model-free memory for LLMs
Jade XinPrincipal Research Scientist, TensorPRO
Long context solves within-session coherence, not cross-session persistence. Once an assistant starts storing memories, a quieter failure appears: its own inferences can be written back, retrieved in later sessions, and presented as user facts. Hallucinations…
Office-Brain and the Night Shift: What Happens When You Turn Off the Meter
Mark PesceCo-Founder, Noops
"Good enough" AI now runs on a reasonably beefy laptop. Qwen3.8-27B benchmarks at an Intelligence Index of 52 - roughly the best money could buy back in February - and inferences happily on consumer kit. When tokens are minted on your own machines, the meter…
Building Agentic Memory for an AI SOC: Why Our Semantic Cache Was the Wrong Answer
Mukesh SinghPrincipal Detection and Response Engineer, Atlassian
We run a fleet of AI agents against production security detections at Atlassian. One tunes noisy detection rules, another reviews new ones, and more are coming for alert triage and hunting. Every one of them started blind and stateless, re-deriving the same…
Building a Security Agent: Model choice, Harnesses and Evals
Simon HarloffCTO, Dam Secure
In this presentation, I’ll show how we set out to give developers useful security feedback on every pull request in under three minutes. The benchmark results and methodology are published here:…
02
Earn the right to trust them
Replace plausible output with evidence. Trace behaviour, evaluate changes, constrain authority and design for security, auditability and real consequences.
From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent
Jiggy KakkadStaff AI Engineer, Quantium
Tinus WillemseExecutive Manager, AI & Data Science, Quantium
Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the…
Building a Security Agent: Model choice, Harnesses and Evals
Simon HarloffCTO, Dam Secure
In this presentation, I’ll show how we set out to give developers useful security feedback on every pull request in under three minutes. The benchmark results and methodology are published here:…
Classifiers are dead. Long live classifiers!
Charli PosnerBuilder, Stile Education
Training a classifier used to mean collecting labelled data, choosing an architecture, training a model, evaluating it, and deploying it. Today, for a surprising number of problems, you can replace most of that with a prompt. At Stile, we’ve been doing exactly…
Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals
Dean SosteSenior Machine Learning Engineer, Canva
Production AI problems are rarely observable through deterministic CI checks. They surface as user feedback, bad LLM-as-a-Judge scores, alerts, or even a vague sense that something is wrong. The evidence that explains them is scattered across traces, prompts,…
Give Every Agent a Flight Recorder
Rahul TrikhaPrincipal AI Engineer, Zendesk
Agent teams should not need to file a ticket with a central evaluation team just to learn whether a new prompt, model, or tool made their agent better. At Zendesk, we developed and deployed a trace-first evaluation platform that gives every agent a flight…
How to Change an LLM System Without Guessing
Yulia KuchinaStaff AI Engineer, Software at Scale
We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide…
03
Change how software gets made
Understand what coding agents demand from codebases, tests, developer tools, review systems—and from the people and organisations adopting them.
Your spaghetti code has an invoice now: what complexity does to coding agents
Artem YakimenkoEngineering Director, Site Reliability, Culture Amp
We've told engineers for decades that high complexity makes code harder for humans to reason about. It turns out it makes code measurably more expensive for agents too. Unlike human frustration, this shows up directly on your API bill. I show and discuss my…
Compounding lessons into skills: How we performed a large code migration using AI (with no new bugs!)
Tom IslesStaff Software Engineer, Canva
Canva's "Ingredient Generation" service powers all of our media generation experiences. Every image, video, audio and 3d object request for generation goes through this service, which offers 50+ AI Models. One of its core capabilities is the ability to…
The BAML Programming Language
Vaibhav GuptaCEO, Boundary
Whether you like it or not, it’s no longer possible to read all the code that’s generated by a model. The solution cannot be, "Just ship it," or you end up with slop everywhere. Neither can the solution be "Read everything," because that’s impractical and…
209 ports in six days, in a language the model barely knew
Burin ChoomnuanPrincipal Engineer/Team Lead/AI Engineer, NewsCorp Australia
Jolt and jank are two young Clojure implementations, one running on Chez Scheme and one compiling to native code through C++/LLVM. Between them they have almost no public code for a model to have learned from. Ask Claude for jank and it confidently writes JVM…
AI Janitor: Making Architecture Review Executable for Coding Agents
Dave CurrieAI Engineer, Square Peg
AI coding agents accelerate code generation, but they also accelerate architectural drift, regressions and false confidence. While building a multi-service production platform with only two of us, I found that a normal human review loop could not keep up with…
Dependency hell is back. This time it's your agent's config.
Jack RudenkoChief AI Officer (CAIO), 10X Labs
Everyone has an agentic harness now. Ours is not special. What nobody has solved is running one across a whole team without every engineer drifting into a private setup. We run Claude Code across 50 engineers and dozens of client codebases at 10xlabs. Within a…