Khang Nguyen Hoang

Khang Nguyen Hoang

Data Scientist

Hello Clever

We Deleted Most of Our Agents. Everything Got Faster

See all speakers

We Deleted Most of Our Agents. Everything Got Faster

The default advice is to add agents: specialised roles, an orchestrator, handoffs between them. We built that. It was a genuinely useful way to explore the problem space, and under real production traffic it was slow and expensive.

When an agent system is slow, the instinct is to reach for a faster model. That was the wrong lever. Our latency wasn't dominated by inference — it was dominated by round trips. Every handoff re-serialised context, every planner turn burned tokens producing nothing a user would ever see, and the overhead compounded per hop. It's the N+1 query problem wearing a different hat: hundreds of small calls where one would do.

So we started deleting. Collapsing to a single agent holding multiple toolkits cut cost per resolved task by roughly two-thirds and latency by about half — same model, no infrastructure change. A caching layer over the repeated-question tail took a further ~30% off the already-reduced figures, on the same principle: the cheapest model call is the one you never make.

I'll show what we measured at each hop to identify round trips as the real cost driver, what broke when we consolidated — tool-selection accuracy, eval attribution, the loss of per-subtask model tiering — and how we handled each. The patterns for managing large toolkits are already known: retrieval over tool definitions, hierarchical grouping, progressive disclosure. What isn't documented is which ones survive production, and where each quietly stops working. Figures are directional and normalised. Includes a side-by-side trace replay of both architectures on an identical task.

Khang Nguyen Hoang

Khang Nguyen is a Data Scientist at Hello Clever, a global payments company supporting checkout in more than 25 currencies, where he owns the data and AI side of the business: the agent systems behind its AI products, the pipelines underneath them, and the evaluation and instrumentation that keep them honest in production. He has spent four years in data and AI, mostly on the unglamorous half of the work: measuring systems, finding out why they misbehave under real traffic, and rebuilding them when the architecture turns out to be wrong. Khang holds an Honours degree in Software Engineering and a Master of AI from RMIT, finishing in the top 2% of the university for both.