In manufacturing, "just-in-time" means parts aren't stockpiled on the factory floor—they're each pulled at the moment they're needed and the whole assembly line runs leaner in response. We've built Ask Albert on the same ideology. As a result, Ask Albert's backing agent has gotten faster and more cost-efficient, even as it grows more feature-rich with every version. In this article, we'll be discussing the story behind our agent harness architecture and what this current design enables for R&D labs.
At a high level, the path we took was:
Ask Albert v1. A fixed set of semantic search workflows over a vector database. Great at discovery, limited to a handful of pathways.
Ask Albert v2. A supervisor agent that composed workflows dynamically by spawning one of 13 specialized subagents to do the work. Each subagent was set in stone: fixed instructions, fixed tools. This pattern is usually called an "orchestrator loop" or "supervisor loop." It opened up a huge range of new use cases, and it turned Ask Albert from a retrieval tool into an assistant that could take action.
Ask Albert v3. A single agent that can handle entire workflows itself, or spawn subagents that it customizes on the spot for the task at hand. It still delegates, but only when it needs to, and the subagents it creates are dynamic rather than predefined. It covers everything v2 did, adds skills, connectors, code execution, file parsing, and more on top, and does all of it faster and cheaper.
The story from v1 to v2 to v3 is a story of steadily increasing flexibility and use case coverage. The more control we give the agent for determining how to reach the outcome, and the more resources we provide it, the more complex the workflows we can support. In v2, we gave our agent a ton of new resources: 13 specialized subagents, each carrying its own tools and domain knowledge. But those resources came pre-bundled. The agent could choose which subagent to call, but not modify their capabilities. In v3, we unbundled them. The agent now dynamically pulls together the exact tools and instructions a task calls for, and spawns a subagent built around that mix only when the work warrants it.
That is what we mean by a just-in-time agent. Nothing is stockpiled in the prompt. Tools, domain knowledge, and helpers all arrive at the moment a task calls for them, and not before. The payoff is real: compared to v2, Ask Albert v3 is 6% more accurate on hard tasks, uses 43% fewer prompt tokens, runs 39% faster, and costs 75% less per case. Getting higher accuracy, more capability, lower latency, and lower cost in the same release is rare in AI product development. Usually you pick two. The rest of this post is about how we got all four.
Ask Albert v1 was excellent at one thing: semantic search over formulations, ingredients, attachments, notebooks and related content.
That worked well for open-ended discovery, like "find me waterborne binders similar to this resin." It struggled when scientists asked precise, operational questions, such as "show all open tasks assigned to me on Project 1234" or "what tasks were completed across the last week?" Those are exact lookups with filters, the same operations a scientist performs in Albert OS every day. Semantic search alone is the wrong tool for that job.
v1 also had a structural ceiling. Every new capability meant building a new, purpose-specific tool and adding a defined workflow for using it to Ask Albert's repertoire. So adding anything beyond the existing handful of retrieval pathways was difficult and time-consuming, and each new workflow had the potential to make Ask Albert worse at picking which workflow to use.
How v1 worked under the hood: a fixed pipeline decomposed the question, ran parallel vector searches, ranked the results, and wrote a cited answer.
Ask Albert v2 replaced v1's fixed pipeline with an agentic loop. A supervisor agent read the user's question, spawned one or more of the 13 domain subagents we had built (projects, inventory, tasks, notebooks, worksheets, search, Breakthrough, and so on), and synthesized the final answer. Each subagent carried deep domain instructions and its own tool set, and neither could change at runtime. The supervisor picked from a menu; it could not edit the menu.
This meant Ask Albert v2 could compose its own workflows to answer questions instead of relying on the fixed set available in v1. Instead of only offering formula search, raw material search, notebook search, and document search (the four pathways in v1), Ask Albert could use the 13 subagents to compose many different combinations of tools on the user's behalf. It worked much better than v1 and let us support a hugely increased magnitude of use cases.
Two things changed in kind, not just in degree. First, Ask Albert could combine structured and semantic search in the same answer: filter down to the exact project, then search within it. Second, and more importantly, it could act. v2 could create tasks, update inventory, edit worksheets, and configure targets on the user's behalf, with the same permissions the user has in Albert OS. That was the moment Ask Albert stopped being a search box and became a scientific assistant.
It also came with some tradeoffs:
Cost and speed. The supervisor and subagent prompts were large, and they had to be resent on every turn of the conversation and on every delegation hop from supervisor to subagent. Answers were much more accurate, but they took far longer to reach the user and cost more.
Subagent isolation. The supervisor was constantly passing work to subagents that couldn't see each other's work. Context was lost at each hop, and on cross-domain questions a subagent often received only partial context. We were sacrificing accuracy because we had to partition the work along boundaries we drew ahead of time, not boundaries that fit the question.
Subagent hallucination. A subagent could try to call a tool that only lived in another subagent's pool, then error out or fail.
To be incredibly clear: none of these tradeoffs ever put Ask Albert's performance below v1 levels. v2 still performed miles beyond v1. But we knew there was accuracy and performance margin available for us to claw back, and we knew where it was.
v3 keeps the agentic loop but drops the fixed hierarchy. Ask Albert v3 is a single agent that can work across the entire Albert platform through dynamic tool discovery. This is where the just-in-time idea takes over. v2 stockpiled: every subagent carried its full toolset and full instructions on every call, whether the task needed them or not. v3 pulls: the agent starts each turn with a slim prompt and fetches tools, domain knowledge, and helpers as the work reveals what it needs.
Imagine one agent with a handful of blank "shell" subagents at its disposal. When the agent decides a piece of work needs a dedicated subagent, it fills a shell with the specific context, instructions, and tool access for that job. In v2, the supervisor had to pick the closest match from 13 predefined subagents. In v3, Ask Albert writes the subagent it actually needs, scoped to the task in front of it. The subagents are dynamic, not set in stone.
The single agent also relies on subagents a lot less overall. That's because it can now fetch the domain expertise that used to be locked inside v2's 13 subagents and use that context directly, rather than outsourcing the work to another agent. This solves both the performance problem and the subagent isolation problem in one move. The agent sees all the domain context it needs, at the moment it needs it, and pays for nothing it doesn't.
The rearchitecture also gave us room to grow the toolkit again. Because new capabilities plug into the same just-in-time primitives rather than requiring a new subagent, v3 shipped with a much broader set of tools than v2 had: reusable skills for common scientific workflows, connectors to the systems our customers already run alongside Albert through the Model Context Protocol (MCP), code execution for data analysis and visualization, file parsing for spreadsheets, documents, and instrument exports, and more. Each one arrives in the prompt only when a task calls for it, so the agent got considerably more capable without getting heavier.

The cleaning-robot analogy:
Imagine we give Ask Albert the task of cleaning the house.
Ask Albert v1 was a collection of four programmed robots that each did one thing: (1) wash plates, (2) wash utensils, (3) wash kitchen tools, (4) run the dry cycle.
Ask Albert v2 was a worker with 13 smarter, pre-programmed robots in the closet. One was called "run dishwasher" and could do everything v1 could do. The other 12 handled bigger jobs, like "clean the living room" and "clean floors." The worker could pick which robots to send out, but couldn't reprogram them. So when "clean the living room" and "clean floors" were both running, they couldn't talk to each other, and they'd run into each other on overlapping areas, like the floor in the living room.
Ask Albert v3 is a worker who can build a robot and program it on the spot. They walk the house, assess what needs to be done, and consult the old programming from those 13 robots when it's useful. Then they program a new robot for exactly the job at hand: "clean the floor in the living room." And most of the time, they just do the job themselves.
The nice property of a just-in-time agent is that new capabilities are nearly free. Adding a domain means adding a bundle and some SDK methods to the catalog, not a new subagent, a new prompt, and a new hop. We're already using these same primitives to extend Ask Albert into more of the design, execute, and analyze loop. If you'd like to see it in action, request a demo.