Skip to main content
Praxa publishes evidence by lane so you can distinguish implementation checks from agent performance and production behavior. The public praxa-benchmarks repository contains reviewed, redacted methods and publication artifacts.

Read the preprint

Review the hypotheses, methods, results, limitations, and references.

Download the PDF

Download the versioned preprint artifact.

Evaluation examples

Adapt concrete task, event-recovery, memory, and governed-tool study designs without treating the templates as results.

Best practices

Freeze comparable conditions, retain failures, derive every visual from checked data, and publish bounded claims.

Troubleshooting

Diagnose missing trials, mismatched artifacts, rate limits, parser drift, and unsafe evidence bundles.
The available results do not establish production readiness or superiority over another agent harness. No production latency, throughput, reliability, cost, or provider-effect benchmark has been published.

Choose the evidence lane

Match package conformance, repository validation, controlled benchmarks, production observation, and user outcomes to the question you are asking.

Interpret a result

Read counts, failure classes, resources, uncertainty, and limitations before comparing systems.

Design your evaluation

Freeze the decision, sample, controls, metrics, evidence records, thresholds, stop rules, and claim boundary.

Results at a glance

Stacked bars showing 17 passed, 16 unresolved, and 3 parse-error trials in both the baseline and reliability-oriented arms.

Terminal-Bench pilot outcomes across 36 trials per arm

Both arms produced the same aggregate outcomes:
Grouped bars showing higher input tokens, output tokens, and execution steps for the reliability-oriented pilot arm.

Measured pilot resource use; values are descriptive, not production forecasts

Four horizontal bars showing 68.09 percent statement, 62.31 percent branch, 75.09 percent function, and 70.84 percent line coverage.

Repository-local coverage reported by the preprint artifact

The machine-readable values and claim boundaries are available in the local evidence summary.

Evidence available today

The preprint From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution reports two distinct empirical lanes.

Repository-local harness validation

Repository-local validation checks whether the implementation satisfies its own executable contracts. The reported source revision passed 1,027 unit tests and 89 Workerd runtime tests. The instrumented source inventory covered all 363 expected files. Measured statement, branch, function, and line coverage were 68.09%, 62.31%, 75.09%, and 70.84%, respectively. These checks support scoped claims about regression behavior, route and schema contracts, packaging, and fail-closed gates at the tested revision. They do not measure general agent intelligence, task success on unseen work, production latency, or verified external effects.

Terminal-Bench pilot

A formative Terminal-Bench Core 0.1.1 study compared a baseline agent loop with a reliability-oriented loop. The study used one model, 12 curated tasks, and three attempts per task in each arm, for 36 trials per arm. Both arms used the same model, base prompt, step limit, wall-clock limit, and token budget. The aggregate trial-accuracy result is a tie with higher token use in the reliability-oriented arm. Individual tasks moved in both directions. Pass-at-k is intentionally omitted because the aggregate summary and one arm’s raw metadata disagreed. The pilot therefore supports analysis of failure modes and harness tradeoffs, not a claim of improved accuracy or cost efficiency.

Evidence classification

Use the evidence class attached to a result before you rely on it. The current publication reports the first two classes. It also documents synthetic conformance fixtures only as protocol checks. A fixture that contains constructed outcomes is not a provider benchmark.

Results that remain unpublished

Praxa has not published measurements for these production-sensitive questions:
  • End-to-end mission latency against a deployed Integration Gateway
  • SSE throughput or tail latency under controlled load
  • Provider-effect success confirmed by an independent verifier
  • Reliability across multiple model families and repeated task samples
  • Comparative cost, latency, or accuracy against another harness
  • Longitudinal behavior in a production deployment
These results require controlled credentials, fixed bindings and budgets, redacted evidence, and independent review. A planned run or a passing local gate does not substitute for the run itself.

How to interpret future results

Before comparing a Praxa benchmark, confirm that the artifact identifies:
  1. The source revision and package versions
  2. The model, provider, reasoning settings, and prompt bindings
  3. The task suite, sampling method, and number of attempts
  4. The resource budgets, time limits, and stopping conditions
  5. The metric definitions and aggregation method
  6. The execution environment and evidence class
  7. Missing trials, failures, exclusions, and limitations
Compare results only when these conditions are compatible. Treat latency and throughput as deployment-specific measurements, not SDK constants.

Publication artifacts

Performance guidance

Learn how client configuration can affect latency, retries, streaming, and mission fan-out without treating guidance as a measured production result.

Methodology

See the sample, controls, exclusions, threats to validity, and comparison requirements.

Data and reproduction

Inspect the versioned artifacts and reproduce the published charts without turning the pilot into a broader product claim.

Reproduction workflow

Verify release lineage, regenerate checked charts, classify environment differences, and validate raw-to-aggregate parity.

Metrics glossary

Use consistent definitions for outcomes, reliability, latency, resources, coverage, and evidence classes.
Last modified on August 14, 2026