Theodoros Galanos
FDE @ APAC
Nomic
The Benchmark Ends. The World Doesn’t: Building Persistent Engineering Environments for Continual Learning
The Benchmark Ends. The World Doesn’t: Building Persistent Engineering Environments for Continual Learning
Large language models can complete sequences of tasks—but continual learning is not just doing more tasks. Most agent environments reset after every episode: state disappears, required follow-up vanishes, delayed effects are cut off, and the next task arrives as though the previous one never happened.
Real engineering projects do not reset. They deal in atoms, not bits: decisions end up in steel and concrete, and a missed comment cannot be patched after the pour. Evidence arrives late, decisions expire, interventions require verification, and a new agent may inherit the consequences of decisions it did not make. A project is one continuing, long-horizon world, and agents must carry work forward across tasks, revisions, and handovers.
This talk presents Persistent Task Worlds: executable environments for evaluating and training agents across continuing engineering project histories. They separate world state, observable evidence, institutional records, conversation state, and learner state, because a record is not the world: a stored session lists which file an agent read, not what it saw, so learning from history means replaying it. Time advances independently of agent turns; documents and requests arrive mid-stream; actions create requirements that persist; fresh agents inherit the same world, including the project memory earlier agents wrote; and evaluation includes effects beyond the visible scoring window, such as whether later revisions preserve earlier work.
I'll present results and early evidence from two studies, both starting from single-episode evaluation (one instruction in, one answer out) as the baseline. First, fixed models and agent harnesses traverse matched environments with different ways of carrying the past: continuous context, raw history, structured state, project memory written by the agent and harness, or only the current snapshot. We measure decision quality, missed follow-up requirements, revision behaviour, how memory is written and reused, downstream outcomes, and interaction cost.
Second, we compare frozen models with models post-trained on real agent trajectories, contrasting trajectories flattened into single episodes with continuing histories. We measure forward transfer, sample efficiency, and performance on untouched histories. External records remain available as a control, letting us separate improvement in the learner from improvement in record-keeping.
The same problem appears in coding agents, scientific workflows, infrastructure, and any system where today's actions change tomorrow's work. Persistent memory is not a persistent world—and a persistent world is not proof of continual learning.
Theodoros Galanos
Theodoros Galanos has spent more than a decade at the intersection of AI, design and engineering. He currently leads the Forward Deployed Engineering in APAC region for Nomic.ai. Previously he was the Generative AI leader at Aurecon, a tier 1 engineering firm. Through The Harness, he explores the environments and evaluations that turn model capability into reliable engineering work.