praxa-benchmarks
repository contains reviewed, redacted methods and publication artifacts.
Read the preprint
Review the hypotheses, methods, results, limitations, and references.
Download the PDF
Download the versioned preprint artifact.
Evaluation examples
Adapt concrete task, event-recovery, memory, and governed-tool study
designs without treating the templates as results.
Best practices
Freeze comparable conditions, retain failures, derive every visual from
checked data, and publish bounded claims.
Troubleshooting
Diagnose missing trials, mismatched artifacts, rate limits, parser drift,
and unsafe evidence bundles.
Choose the evidence lane
Match package conformance, repository validation, controlled benchmarks,
production observation, and user outcomes to the question you are asking.
Interpret a result
Read counts, failure classes, resources, uncertainty, and limitations
before comparing systems.
Design your evaluation
Freeze the decision, sample, controls, metrics, evidence records, thresholds,
stop rules, and claim boundary.
Results at a glance
Terminal-Bench pilot outcomes across 36 trials per arm
Measured pilot resource use; values are descriptive, not production forecasts
Repository-local coverage reported by the preprint artifact
Evidence available today
The preprint From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution reports two distinct empirical lanes.Repository-local harness validation
Repository-local validation checks whether the implementation satisfies its own executable contracts. The reported source revision passed 1,027 unit tests and 89 Workerd runtime tests. The instrumented source inventory covered all 363 expected files. Measured statement, branch, function, and line coverage were 68.09%, 62.31%, 75.09%, and 70.84%, respectively. These checks support scoped claims about regression behavior, route and schema contracts, packaging, and fail-closed gates at the tested revision. They do not measure general agent intelligence, task success on unseen work, production latency, or verified external effects.Terminal-Bench pilot
A formative Terminal-Bench Core 0.1.1 study compared a baseline agent loop with a reliability-oriented loop. The study used one model, 12 curated tasks, and three attempts per task in each arm, for 36 trials per arm. Both arms used the same model, base prompt, step limit, wall-clock limit, and token budget.
The aggregate trial-accuracy result is a tie with higher token use in the
reliability-oriented arm. Individual tasks moved in both directions. Pass-at-k
is intentionally omitted because the aggregate summary and one arm’s raw
metadata disagreed. The pilot therefore supports analysis of failure modes and
harness tradeoffs, not a claim of improved accuracy or cost efficiency.
Evidence classification
Use the evidence class attached to a result before you rely on it.
The current publication reports the first two classes. It also documents
synthetic conformance fixtures only as protocol checks. A fixture that contains
constructed outcomes is not a provider benchmark.
Results that remain unpublished
Praxa has not published measurements for these production-sensitive questions:- End-to-end mission latency against a deployed Integration Gateway
- SSE throughput or tail latency under controlled load
- Provider-effect success confirmed by an independent verifier
- Reliability across multiple model families and repeated task samples
- Comparative cost, latency, or accuracy against another harness
- Longitudinal behavior in a production deployment
How to interpret future results
Before comparing a Praxa benchmark, confirm that the artifact identifies:- The source revision and package versions
- The model, provider, reasoning settings, and prompt bindings
- The task suite, sampling method, and number of attempts
- The resource budgets, time limits, and stopping conditions
- The metric definitions and aggregation method
- The execution environment and evidence class
- Missing trials, failures, exclusions, and limitations
Publication artifacts
- Versioned preprint release
- Canonical LaTeX source
- Evidence data
- Evaluation protocols
- Pipeline improvement roadmap
- PDF preprint
- DOCX preprint
- Public SDK, CLI, and MCP contracts
- Praxa website
Performance guidance
Learn how client configuration can affect latency, retries, streaming, and
mission fan-out without treating guidance as a measured production result.
Methodology
See the sample, controls, exclusions, threats to validity, and comparison
requirements.
Data and reproduction
Inspect the versioned artifacts and reproduce the published charts without
turning the pilot into a broader product claim.
Reproduction workflow
Verify release lineage, regenerate checked charts, classify environment
differences, and validate raw-to-aggregate parity.
Metrics glossary
Use consistent definitions for outcomes, reliability, latency, resources,
coverage, and evidence classes.