Confirms that a declared interface shape responds and records basic latency/error behavior. It does not establish workload quality.
Make every result
reproducible.
This is the minimum method for a qualification result to be reviewable. It does not publish a live-provider result or promise future model quality, latency, throughput or capacity.
TWO DIFFERENT TESTS
Health is not quality.
Measures quality and performance on a frozen, representative, non-sensitive task set using pre-declared scoring and thresholds.
REPRODUCIBILITY MANIFEST
Required for every signed report.
- Mode
- SIMULATION or LIVE PROVIDER; never inferred from surrounding copy
- Workload bundle
- Named tasks, domain coverage and excluded cases
- Prompt-set version
- Immutable ID, checksum and retained source location
- Sample size
- Number of prompts, repetitions and warm-up requests
- Model configuration
- Exact model ID, parameters, tools and context bounds
- Route evidence
- Supplier, region, time window and request identifiers for live runs
- Performance
- TTFT, end-to-end latency, throughput and error rate
- Statistics
- p50, p95 and p99 plus run-to-run variation
- Quality rubric
- Task-specific score, judge/reviewer method and disagreement handling
- Acceptance rule
- Thresholds frozen before execution; no retroactive pass criteria
REPORT OUTPUT
No result without context.
A reviewable report includes the manifest, test class, run status, percentiles, quality rubric, acceptance outcome, known limitations and any excluded failures.
- SIMULATION
- Planning values; no provider request was made
- LIVE PROVIDER
- Credentialed request with route evidence
- INCONCLUSIVE
- Evidence or sample size failed the declared method
- QUALIFIED
- Passed the written acceptance rule for the tested scope only
CURRENT LIMIT
No public Computavia benchmark is represented as a production-provider result today. The customer must receive the route-specific manifest and evidence before relying on a qualification decision.