Skip to main content
Praxa publishes repository validation and controlled agent benchmarks as different evidence lanes. Neither lane substitutes for production observation.

Repository validation

The published revision reported 1,027 unit tests, 89 Workerd tests, and coverage over 363 expected source files. This lane can support claims about the tested contracts at that revision. It cannot establish live-provider success, user experience, production latency, or general agent performance.

Terminal-Bench pilot design

The arms were not randomized or interleaved, and build-cache parity was not established. Those limitations prevent a causal interpretation of small timing or resource differences.

Threats to validity

Twelve tasks cannot represent every coding, business, browser, or provider workflow. Report results as pilot observations for this sample.
The study does not establish behavior across model families, reasoning settings, providers, or prompt variants.
Without randomized or interleaved order and cache parity, elapsed-time comparisons may include environmental effects.
Pass-at-k was omitted because aggregate and raw metadata disagreed for one arm. A metric with unresolved source inconsistency should not be published.
The pilot records benchmark outcomes, not independently verified external provider effects in production.

Requirements for a fair comparison

Before comparing two Praxa runs—or Praxa with another harness—match:
  1. Task set and sampling procedure.
  2. Model, provider, reasoning settings, and prompt.
  3. Tool availability and environment image.
  4. Step, token, wall-clock, and parallelism budgets.
  5. Attempt count, exclusions, and failure handling.
  6. Cache state and arm-order strategy.
  7. Metric definitions and raw-evidence availability.
If these conditions differ, describe the result as a separate experiment rather than a leaderboard comparison.
Last modified on August 14, 2026