Repository validation
The published revision reported 1,027 unit tests, 89 Workerd tests, and coverage over 363 expected source files. This lane can support claims about the tested contracts at that revision. It cannot establish live-provider success, user experience, production latency, or general agent performance.Terminal-Bench pilot design
The arms were not randomized or interleaved, and build-cache parity was not
established. Those limitations prevent a causal interpretation of small timing
or resource differences.
Threats to validity
Small and curated sample
Small and curated sample
Twelve tasks cannot represent every coding, business, browser, or provider
workflow. Report results as pilot observations for this sample.
Single model and configuration
Single model and configuration
The study does not establish behavior across model families, reasoning
settings, providers, or prompt variants.
Arm-order and cache effects
Arm-order and cache effects
Without randomized or interleaved order and cache parity, elapsed-time
comparisons may include environmental effects.
Aggregate metadata disagreement
Aggregate metadata disagreement
Pass-at-k was omitted because aggregate and raw metadata disagreed for one
arm. A metric with unresolved source inconsistency should not be published.
No independent production verifier
No independent production verifier
The pilot records benchmark outcomes, not independently verified external
provider effects in production.
Requirements for a fair comparison
Before comparing two Praxa runs—or Praxa with another harness—match:- Task set and sampling procedure.
- Model, provider, reasoning settings, and prompt.
- Tool availability and environment image.
- Step, token, wall-clock, and parallelism budgets.
- Attempt count, exclusions, and failure handling.
- Cache state and arm-order strategy.
- Metric definitions and raw-evidence availability.