Programme topic
AI engineering
AI engineering is the work of turning models into dependable systems. This collection covers context, tools, evaluation, observability, infrastructure and the production decisions between a promising model and a useful product.
22 published talks
Talks about ai engineering
The Elephant and the Goldfish: Architecture Patterns for Cutting 70% of Agent Token Costs in Production
Tanya DixitForward Deployed Engineer, Google
LLM providers sell you a 2-million-token context window like an elephant that never forgets. If you actually build production agents that way, your latency explodes, your retrieval drifts, and your CFO will shut you down in 90 days. In production, the best…
From Prompt Rules to Structural Guarantees: The Harness Behind a Production Analytics Agent
Jiggy KakkadStaff AI Engineer, Quantium
Tinus WillemseExecutive Manager, AI & Data Science, Quantium
Checkout AI answers open-ended questions about retail sales data in natural language. It plans, calls analytics tools over MCP, executes Python in a sandbox, and returns a written analysis with charts. In a system like this, failure is rarely a crash: the…
The guess never becomes a fact - provenance-governed, model-free memory for LLMs
Jade XinPrincipal Research Scientist, TensorPRO
Long context solves within-session coherence, not cross-session persistence. Once an assistant starts storing memories, a quieter failure appears: its own inferences can be written back, retrieved in later sessions, and presented as user facts. Hallucinations…
Office-Brain and the Night Shift: What Happens When You Turn Off the Meter
Mark PesceCo-Founder, Noops
"Good enough" AI now runs on a reasonably beefy laptop. Qwen3.8-27B benchmarks at an Intelligence Index of 52 - roughly the best money could buy back in February - and inferences happily on consumer kit. When tokens are minted on your own machines, the meter…
Building Agentic Memory for an AI SOC: Why Our Semantic Cache Was the Wrong Answer
Mukesh SinghPrincipal Detection and Response Engineer, Atlassian
We run a fleet of AI agents against production security detections at Atlassian. One tunes noisy detection rules, another reviews new ones, and more are coming for alert triage and hunting. Every one of them started blind and stateless, re-deriving the same…
Building a Security Agent: Model choice, Harnesses and Evals
Simon HarloffCTO, Dam Secure
In this presentation, I’ll show how we set out to give developers useful security feedback on every pull request in under three minutes. The benchmark results and methodology are published here:…
Tools Before Autonomy: What 200,000 Tool Calls Taught Us About Agentic Evals
Dean SosteSenior Machine Learning Engineer, Canva
Production AI problems are rarely observable through deterministic CI checks. They surface as user feedback, bad LLM-as-a-Judge scores, alerts, or even a vague sense that something is wrong. The evidence that explains them is scattered across traces, prompts,…
Query before mutation: guardrail patterns for agents that touch production infrastructure
Jeffrey AvenMaintainer, StackQL Studios
Covered in the session: - Why plan-and-apply was always a human safety mechanism: a person reads the plan, notices something is off, and stops the apply - and why this quietly disappears the moment an agent is the operator - Query before mutation as the…
Give Every Agent a Flight Recorder
Rahul TrikhaPrincipal AI Engineer, Zendesk
Agent teams should not need to file a ticket with a central evaluation team just to learn whether a new prompt, model, or tool made their agent better. At Zendesk, we developed and deployed a trace-first evaluation platform that gives every agent a flight…
How to Change an LLM System Without Guessing
Yulia KuchinaStaff AI Engineer, Software at Scale
We were running a production LLM pipeline that classified legal documents, and every change was a guess. Swap a prompt, change a model — better or worse? Nobody could say. The outputs looked plausible either way, and “plausible” is exactly how LLM systems hide…
Don't Fight Hallucinations. Make Them Impossible
Nadia MakarevichPrincipal Engineer, Heatseeker
The Heatseeker AI chat answers data questions for marketers who make decisions with million-dollar budgets. Wrong answers or hallucinated numbers are not an option here, as you can imagine ;) The fight against them (hallucinations, not marketers) was long and…
Running 100K+ Concurrent Agents: It's Never the Prompt
Harry NguyenAI engineer, insightfactory.ai
Your Agents Aren't Failing at Prompting Every agent demo works. Then you run ten thousand at once, and the failures have nothing to do with the model: dead tuples piling up, a sync HTTP client freezing the event loop, requests hanging forever with no total…
Adaptive Learning: Encoding Clinician Edits as Memory
Vlad GavrilovSenior AI Engineer, Heidi
Clinicians edit their AI-generated notes for many reasons: adding or removing content, fixing spelling, and preferences for structure, wording, or terminology. Some of these edits are contextual, applying in some situations but not others. The data is noisy,…
Ontologies: AI’s Operating Manual For Your Business
Gareth WilliamsPrincipal Engineer, Wesfarmers
Tools churn. Factories commoditise. When everyone has the same models, the advantage goes to whoever gives their agents the clearest blueprint of how the business works - what exists, how it connects and the rules it runs on. That blueprint is an ontology. In…
The BAML Programming Language
Vaibhav GuptaCEO, Boundary
Whether you like it or not, it’s no longer possible to read all the code that’s generated by a model. The solution cannot be, "Just ship it," or you end up with slop everywhere. Neither can the solution be "Read everything," because that’s impractical and…
Classifiers are dead. Long live classifiers!
Charli PosnerBuilder, Stile Education
Training a classifier used to mean collecting labelled data, choosing an architecture, training a model, evaluating it, and deploying it. Today, for a surprising number of problems, you can replace most of that with a prompt. At Stile, we’ve been doing exactly…
Don't stop, won't stop. Automated verification of long-running agentic loops
AJ FisherVP Digital Science, Tetratherix
Long-running agentic loops create different engineering problems than code generation. Once agents work unattended for many hours, iterate through 20 or more review cycles, and build stacked PRs towards a larger goal, the challenge changes to keeping that…
Where Should the Dice Roll? Placing Non-Determinism Deliberately in Enterprise AI
Vighnesh DeshpandeAI Engineer, Vivanti Consulting
Most production AI guidance assumes you want consistency: pin the prompt, lower the temperature, eval for drift. But a whole class of enterprise use cases - idea generation, recommendation, exploration, synthesis - is worthless if the output is predictable.…
The Benchmark Ends. The World Doesn’t: Building Persistent Engineering Environments for Continual Learning
Theodoros GalanosFDE @ APAC, Nomic
Large language models can complete sequences of tasks—but continual learning is not just doing more tasks. Most agent environments reset after every episode: state disappears, required follow-up vanishes, delayed effects are cut off, and the next task arrives…
Extracting Tacit Knowledge for Production Agents
Khali Kalpa-YoungFounder, Alchymie
An eval is a definition of "good" written down well enough for a machine to grade against. In a real business, "good" isn't written down anywhere. It lives in someone's reaction, in what past work actually did, and in how a person talks. So the eval comes…
We Deleted Most of Our Agents. Everything Got Faster
Khang Nguyen HoangData Scientist, Hello Clever
The default advice is to add agents: specialised roles, an orchestrator, handoffs between them. We built that. It was a genuinely useful way to explore the problem space, and under real production traffic it was slow and expensive. When an agent system is…
Does This Agent Make My Context Look Big? Right-Sizing AI Architectures for Production
Hamish SongsmithHead of Applied AI, Silicon Quantum Computing
All-in-one personal agent harnesses showcase the incredible potential of capability-rich AI assistants. But deploying a monolithic "do-it-all" agent into production often leaves teams struggling with context dilution, fragile tool calls, un-evaluable execution…