BENCHMARK METHOD · VERSION 2026.08

Make every result
reproducible.

This is the minimum method for a qualification result to be reviewable. It does not publish a live-provider result or promise future model quality, latency, throughput or capacity.

TWO DIFFERENT TESTS

Health is not quality.

SYNTHETIC HEALTH CHECK

Confirms that a declared interface shape responds and records basic latency/error behavior. It does not establish workload quality.

WORKLOAD EVALUATION

Measures quality and performance on a frozen, representative, non-sensitive task set using pre-declared scoring and thresholds.

REPRODUCIBILITY MANIFEST

Required for every signed report.

Mode
SIMULATION or LIVE PROVIDER; never inferred from surrounding copy
Workload bundle
Named tasks, domain coverage and excluded cases
Prompt-set version
Immutable ID, checksum and retained source location
Sample size
Number of prompts, repetitions and warm-up requests
Model configuration
Exact model ID, parameters, tools and context bounds
Route evidence
Supplier, region, time window and request identifiers for live runs
Performance
TTFT, end-to-end latency, throughput and error rate
Statistics
p50, p95 and p99 plus run-to-run variation
Quality rubric
Task-specific score, judge/reviewer method and disagreement handling
Acceptance rule
Thresholds frozen before execution; no retroactive pass criteria
REPORT OUTPUT

No result without context.

A reviewable report includes the manifest, test class, run status, percentiles, quality rubric, acceptance outcome, known limitations and any excluded failures.

SIMULATION
Planning values; no provider request was made
LIVE PROVIDER
Credentialed request with route evidence
INCONCLUSIVE
Evidence or sample size failed the declared method
QUALIFIED
Passed the written acceptance rule for the tested scope only
CURRENT LIMIT

No public Computavia benchmark is represented as a production-provider result today. The customer must receive the route-specific manifest and evidence before relying on a qualification decision.