1.0 only when its tool calls produce exactly the right state transitions and its responses contain all required output values.
QitOS ships TauBenchAdapter and a self-contained runtime (TauRuntimeEnv) so you can run Tau-Bench without installing the upstream tau_bench package. All task data, tools, wiki, and rules are vendored under qitos.benchmark.tau_bench.port.
Environments
Setup
1
Install benchmark dependencies
2
Set your model API key
Tau-Bench task data is vendored inside the QitOS package. You do not need to download any external dataset.
Loading tasks
Configuration
TauBenchAdapter accepts the following parameters:
Running the evaluation
Start with the official CLI:tau_bench_eval.py remains available as a benchmark-specific wrapper over the same official result shape and trace contract.
Run a single task:
How the runtime works
TauRuntimeEnv is a minimal drop-in for the upstream Tau environment. It exposes a reset / step / calculate_reward interface:
1.0 requires both the correct state hash and all expected output strings present in agent responses.
Task structure
Expected output
Each result line in the output JSONL file contains:tau-bench evaluation:
qita:
