Skip to main content
Engine is the single execution kernel for all QitOS agent workflows. It runs the step loop, coordinates your AgentModule hooks, dispatches tool calls to the ToolRegistry for execution, applies critics (post-step evaluators that can approve, stop, or retry the step), checks stop conditions, and writes trace artifacts (persistent records like steps.jsonl and events.jsonl that capture what happened during the run). You interact with it directly only when you need control beyond what agent.run() provides.
QitOS enforces a single-kernel rule: there is exactly one Engine per run. Extensions like parsers, critics, memory adapters, and toolkits attach to this pipeline — they do not introduce a second execution loop.

How the loop works

Each step follows a fixed sequence:
  1. prepareagent.prepare(state) formats state into the model-ready prompt text.
  2. decideagent.decide(state, observation) is checked first; if it returns None, the Engine calls the LLM via the configured parser.
  3. act — Tool calls in the decision’s actions list are executed against the ToolRegistry.
  4. reduceagent.reduce(state, observation, decision) updates state with the new observation.
  5. critics — Any registered Critic instances (post-step evaluators that can approve, stop, or retry the step) evaluate the step; they can trigger a stop or retry.
  6. check_stop — Budget exhaustion, FinalResultCriteria, agent.should_stop(), and any custom StopCriteria are evaluated.
  7. trace — The step record and events are written to the TraceWriter.

Constructor

Prefer calling agent.run() for single-run workflows. Use Engine directly when you need to reuse an Engine across multiple runs or configure hooks dynamically between runs.

Engine.run(task)

Accepts a plain string objective or a structured Task object. Returns an EngineResult. When you pass a Task, the Engine extracts the budget from task.budget and overrides the Engine’s own budget for that run. It also orchestrates resource staging and environment lifecycle (reset, observe, close) automatically.

EngineResult

Typical usage:

Hooks

Hooks observe and react to lifecycle events without modifying Engine internals. They implement EngineHook and are called at on_before_step and on_after_step boundaries.
You can also pass hooks at construction time via the hooks parameter, or at run time via agent.run(hooks=[...]).

Budget exhaustion

When the step budget, wall-clock time limit, or token budget is exceeded, the Engine sets state.stop_reason to the appropriate value and emits an END event. The run terminates gracefully and EngineResult is still returned — inspect state.stop_reason to detect this case.

Building an Engine from AgentModule

AgentModule.build_engine() is a convenience factory that creates an Engine pre-bound to the agent:
This is equivalent to Engine(agent=agent, ...) and is what agent.run() calls internally.

AsyncEngine

AsyncEngine provides non-blocking execution for agent workflows. It wraps the same Engine loop but runs blocking calls in a thread pool, making it safe to use inside asyncio event loops.

AsyncEngine.arun(task)

Run the agent loop asynchronously. Returns the same EngineResult as Engine.run().

AsyncEngine.arun_stream(task)

Run the agent loop and yield structured EngineEvent objects as they occur — ideal for real-time UI updates or streaming progress to a client.
Events are emitted at step boundaries (step_start, step_end), phase transitions (decide, act, reduce, critic, check_stop), and multi-agent events (handoff, delegate, fanout). The stream always begins with run_start and ends with run_end.

EngineEvent

EventStream

EventStream is the async queue that underpins arun_stream(). You can also use it standalone to fan out events to multiple consumers:

Async models

When the configured model implements acall() (from AsyncModel), AsyncEngine can invoke it without blocking the event loop. Built-in async model adapters:
AsyncEngine.arun() works with any model — sync models are automatically dispatched to a thread pool. Use async model adapters when you need true non-blocking I/O, e.g. in high-concurrency web servers.