- 用的到底是什么模型与解析器?
- 暴露了哪些工具?
- 这是不是同一个基准测试划分?
- 能不能回放?
- 能不能和上一版运行做差异对比?
- 官方运行契约
- 统一的基准测试结果结构
- 更完整的 qita 回放、差异对比与导出路径
Documentation Index
Fetch the complete documentation index at: /llms.txt
Use this file to discover all available pages before exploring further.
一篇简短的设计笔记:官方运行、统一基准测试输出与尽力回放为什么重要。
