Choose the mechanism you want to study
All six use
AgentModule + Engine + Session; none embeds a second tool executor
or direct model SDK loop. The framework does not prescribe a planner, a memory
policy, a scoring rubric or a universal definition of agent success.
What changed in the framework
build_agent_composition accepts an optional agent_factory with named arguments
config, model, tool_registry, protocol, parser. Return an AgentModule
bound to these supplied objects; the resolved protocol may be bound by its ID.
The default is still ConfiguredAgent. A factory can register project tools using
the supplied registry. Composition validates bindings and required-tool policy,
owns the model/environment/runtime/Engine, and cleans its resources on failure.
Explicitly injected resources remain borrowed. Recreate the factory and resolver
in the new process when restoring; callable objects are never serialized.
The composition reference documents signatures and
ownership. The skill library reference documents durable
revisions, full-body selection and explicit Memdir deletion. A child transfer fix
also restores external snapshot components before recapturing projected input,
so independent review receives the verified artifact rather than pristine files.
These mechanisms replace repeated resource, persistence and dispatch boilerplate;
task strategies stay in the projects where researchers can change them.
Install and run without private configuration in Git
Build and install the QitOS wheel from the implementation checkout; these APIs are not yet a separately published PyPI release. Then install each desired project fromexamples/projects/. Use Python 3.10 or 3.12 and the declared provider extra
(qitos[openai]) for the live OpenAI-compatible backend. Each course includes its
own pyproject, public agent.yaml, three tasks, policy, launcher and checker.
Keep the real model configuration, credential file, database, artifacts and raw
trajectory in private directories outside any Git checkout. The public model URL
is a placeholder, not a working credential. The default output ceiling is 10,240;
raise it in private model configuration if needed. Step/request/time budgets and
retry policy are per-task guards, not a hidden total experiment quota.
validate makes no model request. run --live opts into actual execution.
resume --session reconstructs a paused Session through the same factory.
inspect reads the trajectory. Resuming a Session is not undoing external effects
or rolling back a Git working tree. The Docker component supports bounded,
artifact-backed cold file restoration; it is not a live-process checkpoint or VM.
Workspace snapshot size/file limits can block large repositories explicitly.
Learn a design, then replace or combine it
Every course begins with the original idea and the adaptation boundary. A named excerpt is generated from complete source, so a short teaching snippet cannot drift from the executable project. Follow the chapter’s task, inspect actual tool and phase events, study failure receipts, replace one module, then compare results. For horizontal composition, install PlanAct and Pi alongside QitOS and importqitos_lab_planact.with_notebook.composition_with_notebook. It binds an existing
Memdir namespace, a replaceable context budget and compactor, and Pi’s independent
weighted_summary extension. The installed consumer verifies real request memory,
plan changes and budget compaction. It does not change Engine or hide another loop.
Hermes stores procedural documents; Voyager stores checked programs. Both use the
same library seam. Catalog lookup is not loading; actual selected versions and
their complete bodies must be observed. Voyager’s bounded mastery scheduler
advances only after a persisted verified skill. Its three data-programming tasks
are not Minecraft exploration or the original paper’s learned curriculum.
Four different acceptance questions
- Framework correctness: do permissions, failures, cleanup, restoration and namespace isolation hold under deterministic counterexamples?
- Design mechanisms: did replanning, independent review, selected loading or executable reuse actually happen, rather than appear in a final message?
- Installed usability: do both Python versions run outside the source tree using only the installed framework and declared packages?
- Task effects: do three tasks per project, repeated three times, pass an independent checker with actual model requests?
scripts/qualify_agent_design_lab.py runs the opt-in matrix serially against the
installed packages. It stores raw evidence privately. A model’s own success
claim is never the checker. The code checker runs controller-selected tests inside
Docker and binds source digests; it is not a proof against malicious code designed
to subvert Python’s own test process. Public summaries require separate privacy
and redistribution review; hashes alone do not authorize publication.
Primary sources and boundaries
- ReAct: interleave reasoning, action and observation; this course does not demand or invent hidden reasoning tokens.
- Plan-and-Act: separate high-level planning from execution; this is a runtime adaptation, not planner training or benchmark reproduction.
- Pi coding agent: a small core with extensions; no TUI or package ecosystem parity is claimed.
- Claude Code design: context gathering, tools and verification; no proprietary internals are copied.
- Hermes skills: progressive loading; hub, messaging and complete Agent Skills interoperability are out of scope.
- Voyager: verified executable skill accumulation and reuse; this adaptation uses Docker data tasks, not Minecraft or embedding retrieval.
