Evaluation pipeline
Cacheon has two evaluation paths with different authority:
- the developer path helps a miner build and debug a proposal;
- the production referee path can create retained qualification, crown, settlement, and reward state.
The developer commands intentionally reuse parts of the ABI and evaluation machinery, but their output is not production crown evidence.
Developer path
slots -> scan -> verify -> chain-package -> host -> chain-submit -> chain-status| Step | Purpose | Authority |
|---|---|---|
slots | Inspect the live slot ABI | Informational |
scan | Run static bundle policy checks | Diagnostic |
verify | Check a single op/block against its trusted reference; collective slots use distributed verification | Diagnostic |
chain-package | Produce the deterministic hosted archive and content identity | Submission preparation |
| HTTPS host | Make the immutable archive available for validator fetch | Transport only |
chain-submit | Commit the proposal through timelock commit-reveal; optional eval-cost --pay | Chain intake |
chain-status | Inspect submission state | Informational |
Local measurements are useful for iteration. They do not select a production arena authority, reserve a finalized cohort position, retain authenticated resident speed evidence under the current v7 B/C/[B′] or v8 B/C/B′ schedule, perform the registered eager audit A and pristine-reference T stages, or satisfy independent reproduction.
Production path at a glance
1. Finalized intake
Submissions enter through native timelock commit-reveal. The validator acts only on finalized chain order. Finalized block position and commitment identity establish priority; evaluator network arrival does not.
FinalizedIntakeStore persists production authority in SQLite. It records finalized observations, fetch state, copy disposition, screen receipts, cohort reservations, qualification attempts, evidence roots, reproduction state, stack transitions, settlement, and weight-publication state.
State transitions are typed and transactional. The validator does not reconstruct production authority from console output, mutable directories, or a legacy JSON ledger.
Principal code: chain/intake.py and chain/validator_loop.py.
2. Hardened fetch and publication
The submitted URL is transport, not identity. The fetch path:
- applies the validator's HTTPS and network policy;
- writes into validator-private bounded storage;
- recomputes the deterministic bundle content hash;
- compares it with committed identity;
- derives the copy/provenance disposition;
- publishes the complete artifact into an immutable, hash-addressed worker namespace;
- reopens the publication before any screen or launch consumes it.
Partial downloads, path tricks, changed content, duplicate/malformed archives, and publication mismatches fail closed. Candidate workers receive only immutable publications; they do not fetch miner URLs themselves.
Principal code: chain/fetch.py, chain/payload.py, and bundle_hash.py.
3. Arena and target resolution
An ArenaServiceRegistry maps a public arena identifier to a closed
ArenaService. Its manifest directly binds:
- runtime, base-engine, validator-overlay, worker, model, architecture, GPU, and topology identities;
- the scored workload cells and prompt-seed scheme;
- non-crown screen policy;
- queue depth/age, cohort size, screen/qualification concurrency, and retry policy;
- the qualification-policy digest; and
- the reviewed provider implementation digest.
The target catalog, incumbent and candidate stacks, graph/engine settings,
calibration, reference, evidence, and quality identities are closed later by
the promoted candidate bindings and the provider-created typed qualification
plan. ArenaService checks the plan's policy digest and finalized reservation
order; it does not pretend all of that authority is a field of the service
manifest itself.
The proposal is resolved against the exact target catalog snapshot. A registered candidate must match its target members and permitted features; unregistered work fails resolution.
The command-line chain-validate loop can perform intake alone. Full production qualification requires the operator to inject a real ArenaServiceRegistry and select --arena-id; the repository does not manufacture a production arena provider from implicit defaults.
Principal code: arena_service.py, target_catalog.py, and stack_plan.py.
4. Non-crown screens and routing
Expensive full-engine qualification is reserved for plausible candidates. The registered arena applies a fixed sequence of non-crownable screens:
- static policy and manifest resolution;
- deterministic source closure and build planning;
- typed ABI correctness, including distributed verification for collectives;
- graphs-on capture and dynamic-input replay; and
- abbreviated serving on a small registered workload, implemented by the calibrated resident lane when the contribution is safely hot-swappable.
The resident screen keeps one stock engine alive, swaps candidate bundles into a separate resident session, recaptures graphs, and evaluates a bounded queue against shared stock brackets and canaries. Every batch is bound to its swap generation. The screen exists only to route work: a promising result advances to qualification and a clearly uncompetitive, stable result may be rejected under the registered screen policy. It cannot create a qualification PASS, crown, settlement speedup, or reward claim.
For a standing resident worker, the ordered ABI and graph screen rows can be explicit carrier deferrals: the build product is reopened in those positions, while the subsequent resident swap acknowledgement and read perform the actual all-rank registration and graph recapture without unloading the stock model. Those deferrals must never be interpreted as qualification correctness evidence. The isolated qualification audit and numerical judge remain mandatory before a candidate can PASS or crown.
Some contribution classes cannot satisfy the hot-swap contract. Direct AOT artifacts, dependency patches, native rebuilds, and setup hooks receive an explicit screen waiver and proceed to authoritative qualification. A waiver means “not screenable,” not “screen passed.” The arena caps a promoted screen cohort by both its registered policy and the provider's current capacity.
Infrastructure errors and measurements too ambiguous for an attributable rejection remain retryable rather than being converted into a loss.
Principal code: eval/oci_resident_session.py,
eval/resident_queue.py,
and eval/resident_screen_lane.py.
5. Cohort authority
At an evaluation boundary, the validator freezes:
- one incumbent
EvaluationStackManifestdigest; - a finalized chain-ordered set of candidate reservations;
- one registered arena and target-catalog context;
- deterministic candidate order derived from committed authority;
- workload, prompt, seed, role, topology, calibration, and evidence policy.
Each candidate stack is the frozen incumbent with exactly one registered target delta. Production qualification uses two isolated resident TP lanes while serializing GPU work. The primary attempt assigns the incumbent and candidate to fixed physical lanes. An eligible reproduction must bind the exact opposite physical-lane role assignment. The lane swap is part of independence authority; it is not a scheduler preference.
The routing screen may amortize a stock engine and shared brackets across candidates. Authoritative qualification does not inherit those measurements. It constructs a fresh, candidate-specific authority and retains each speed, audit, graph, and T product against that exact delta.
Principal code: eval/qualification_intake.py and stack_plan.py.
6. Version-3 adaptive qualification
The production evidence protocol is version 3. Inside it, the provider selects the resident-speed subpolicy from candidate features under the same predicate the worker uses to choose its execution substrate:
- a hot-swappable candidate uses v7 on the standing resident pair, reads B/C, and takes B′ only when the B/C ratio cannot decide under the sealed bounds;
- a non-swappable candidate (CUDA, C++, PTX, AOT, dependency-patch, or setup surface) uses v8's separate baseline and candidate processes and always reads B/C/B′ because the quality gate consumes the second stock read; and
- C′/B″ remain readable only in historical v2–v5 evidence. Nothing currently commissioned seals those reads.
Both current subpolicies retain stage and total budgets and the required physical-lane role assignment. Old evidence reopens under its own versioned arithmetic; a label change cannot upgrade it.
The authoritative work is staged. Speed is decided first; audit and pristine-reference quality run only after the speed stage remains eligible, apart from an explicitly registered calibration-observation continuation.
| Arm | Stack | Timed? | Purpose |
|---|---|---|---|
| B | Exact frozen incumbent on the assigned baseline lane/process | Yes | Opening performance read |
| C | Incumbent plus one exact target delta on the disjoint candidate lane/process | Yes | Candidate measurement and sealed trajectory |
| B′ | The same incumbent authority as B | Yes | Conditional v7 decision bookend; mandatory v8 stock-drift control |
| C′ / B″ | Historical v2–v5 candidate/baseline repeats | Yes | Reopen old evidence only; unreachable in current v7/v8 work |
| A | Candidate in a separate eager, untimed role | No | Registered sampled slot audit and typed host regrade |
| T | Pristine candidate-free reference | No | Teacher-forced semantic quality and hidden tasks |
Native build and timed execution use separate containers. Before a runtime arm starts, a disposable, no-GPU/no-network prebuild OCI parses the materialized tree, invokes only registered build patchers, and emits a sealed native publication. The resident speed lanes mount that reopened publication read-only. Candidate Python import, engine construction, and execution occur only in the runtime's positively identified scheduler ranks; runtime ranks may validate and load native products but may never compile or repair them. Both stages use read-only roots, bounded mounts and protocols, and host-owned cleanup; the trusted controller also owns timing.
For a direct-artifact row, prebuild executes the declared compiler factory only
inside a no-egress compiler child and publishes CUBIN rather than a host launcher.
After rank-local CUDA setup, the scheduler worker admits the exact CUBIN, binds its
complete driver-observed ABI to the declarative device plan by ordinal, and
materializes parameters and lifecycle storage in validator code. Qualification
requires per-member aot_loaded, aot_invoked, and normal completed coverage,
with no fallback receipt. See Sealed direct artifacts.
V7 precommits the mathematical bounds that make a B/C result invariant to every
legal B′. It stops after B/C for a clear result and collects B′ only inside the
inconclusive band. V8 precommits all three reads, so B′ is taken regardless of
the observed B/C result. Neither path can request a favorable extra read after
seeing an outcome. When later brackets drift beyond the v5+ sealed ceiling,
they are excluded and the adjacent C/B comparison decides; the drift is retained
as evidence rather than converted into an indefinite NO_DECISION.
When the registered plan requires sampled slot audit, a separate eager, untimed candidate role emits bounded raw facts. The trusted host grades exact slot × TP-rank/process coverage through a Torch-free gate and canonicalizes floating-point facts before durable receipt identity is computed. These facts cannot enter the charged speed roles. A slot or target without a registered audit requirement cannot acquire audit authority from an incidental diagnostic receipt.
After the candidate speed and audit lifetimes are destroyed, T grades the sealed trajectory under a separate pristine lifetime. T never contains the candidate and does not compete on speed. Hidden reference work, quality policy, and selected prompt identity are bound into retained evidence.
The host applies the exact versioned policy. Conceptually, a v7 clear decision uses C/B; an inconclusive v7 result and every v8 result use the registered B/B′ bookend unless the later bracket is excluded by the sealed drift rule:
v7_clear_speedup = C / B
bookended_speedup = C / mean(B, B′)
required_bar = 1 + max(margin_floor, noise_multiplier × measured_noise)The exact registered policy, not this explanatory formula, is authoritative.
Principal code: eval/crossover_runtime.py,
eval/qualification.py,
eval/qualification_runner.py,
and audit_gate.py.
7. Verdict semantics
Qualification has three outcomes:
PASS
The candidate clears the registered speed, quality, graph, evidence, and whole-stack requirements under a stable cohort authority. A first pass is retained as reproduction_pending; it does not settle alone.
FAIL
Complete evidence attributes a policy failure to the candidate under a valid authority. Examples include incorrect output, a quality regression, or a stable speed result below the registered bar.
NO_DECISION
The evaluator cannot make a valid attributable decision. Infrastructure failure, missing evidence, baseline drift, broken cohort invariants, or an invalid reference lifetime must not mint either a crown or a loss. The result retains a failure product and retry policy.
The distinction is load-bearing: treating evaluator failure as candidate failure would let infrastructure state rewrite economic truth.
Worked lifecycle: retry, reproduce, settle
Consider a hypothetical candidate for one registered singleton target:
- Intake fixes its finalized priority and immutable publication. It clears all five non-crown screens.
- Its first v7 attempt produces valid B/C work inside the inconclusive band, but
the required B′ evidence cannot be authenticated. The attempt is
NO_DECISION. The proposal remains retryable; it has neither lost nor passed. - A later fresh authority reopens the same candidate identity. This time B/C
clears the precommitted v7 pass bound, so no outcome-dependent B′ is taken.
The registered audit and T products accept the candidate. The attempt becomes
the first
PASSand state becomesreproduction_pending. - A second independently selected authority repeats the exact reproduction identity, swaps the incumbent and candidate physical-lane roles, and also passes. A pass against a newer incumbent, different target specification, or the same lane-role assignment would not count as this reproduction.
- Settlement reopens both evidence roots and makes the pair eligible for its same-authority cohort. If the candidate is selected as that cohort's current registered winner, settlement takes the lower accepted speedup, revalidates the live target transition, and atomically updates the evaluation stack; otherwise the pair is held.
- If crowned, reward projection can now see the active claim. Product integration remains a separate review; no proposal bytes have entered a release merely because settlement completed.
The values and candidate in this walkthrough are illustrative. The state transitions and failure semantics are the important part.
8. Independent reproduction
Settlement requires a second PASS for the exact same core reproduction
identity. SettlementReproductionIdentity contains exactly the arena digest,
target ID, selected-delta digest, hotkey, incumbent stack/tree digests, and
candidate stack/tree digests.
The settlement pair applies additional rules around that core identity. The two qualification rows must match the same contribution, reservation, finalized priority, manifests, members, and arm, while seven independence fields must all differ: qualification authority, plan, attempt, report, selection commitment, selection-secret commitment, and selection evidence. Authority therefore does not belong inside the equal core identity; distinct authority is a separate pair constraint.
For version-3 resident-family evidence, including current v7/v8 witnesses, the pair must also prove the exact physical-lane role swap required by the registered plan. Two nominally independent attempts that assign stock and candidate to the same physical lanes do not satisfy production reproduction.
An attributable second-attempt FAIL terminates the proposal. A
NO_DECISION follows the registered bounded retry/hold policy. In either case,
the retained first pass cannot update the incumbent by itself.
9. Settlement
Settlement reopens both recorded attempt references instead of trusting an in-memory verdict. The references may live under the same content-addressed store root. It verifies:
- both reports are complete
PASSresults; - all seven required authority, attempt, report, commitment, and selection digests are pairwise distinct across the two passes;
- reproduction identities match exactly;
- the incumbent and target transition are still current;
- target overlap, displacement, and composition remain valid;
- the requested stack update matches the measured candidate.
The planner leases one cohort whose rows share qualification authority and incumbent
state. Stale rows are held, and exactly one current registered winner is selected across
the remaining rows, including rows whose targets do not overlap. Other current pairs are
held as conflict_lost or incumbent_advanced; a stale pair is held as
stale_incumbent.
For the selected winner, the conservative settled speedup is the lower accepted speedup from the two passes. The stack transition and settlement evidence are committed transactionally; a partial write cannot expose a half-updated incumbent.
Principal code: settlement.py.
10. Incentive state and weight publication
Settlement and weight publication are separate state machines. The repository retains two explicitly fenced incentive paths:
- Legacy V1 projects active standing and discovery claims through
set-weights. During an all-uncrowned bootstrap, an operator may provide a registered burn hotkey; the burn projection becomes invalid as soon as a crown, claim, or active V2 composition exists.--watchoperates the same journaled reconciler continuously with bounded retry rules. - V2 finite debt is a retained design whose implementation was extracted from the tree on 2026-08-09 without ever being activated; see Finite-debt V2.
The publisher persists intent and later readback states. An SDK return value does not prove inclusion, and the publisher may not advance economic authority from an unconfirmed vector.
Principal code: chain/weights.py.
See the emissions policy.
11. Integration
A settled crown may enter integration review, but serving remains a separate state machine. Reviewed source is promoted to an integrated contribution; the signed chain-independent release product was removed on 2026-08-19 because no release was ever produced or consumed.
The dated State of record tracks implementation and validation limits.
Operational handoff checklist
Before treating an attempt as production authority, an operator or reviewer should be able to reopen, rather than merely observe, each handoff:
- finalized chain position, commitment, fetched content identity, and immutable worker publication;
- registered arena, target catalog, incumbent manifest, candidate transition, and screen receipt;
- lane identities and a versioned
ResidentSpeedWitnesscontaining exactly the scheduled rows (current v7 B/C with optional B′, or current v8 B/C/B′), plus retained graph/quality/pristine-T references and witnesses; richer raw session/device frames are validated in-run but are not serialized into the attempt; - frozen calibration and the exact policy that maps the retained witness/evidence products to the verdict;
- the seven digest-distinctness fields across primary and reproduction over one reproduction identity;
- transactional settlement events and resulting evaluation-stack digest;
- reward projection plus publication intent/status/chronology records; later readback vectors must be re-observed because the journal does not serialize them; and
- if shipping is proposed, the separate integration records and signed release identity.
Console output, a green local benchmark, one PASS, or a successful chain SDK return is
not a substitute for the corresponding reopenable product.
Product model
Cacheon Engine is a chain-independent, open inference-acceleration distribution. It pins the supported SGLang runtime as part of product identity and assembles accepted data-plane optimizations into a canonical engine…
Target and slot contract
The slot contract is Cacheon's narrow waist: a stable, validator-owned tensor boundary between untrusted optimization code and a pinned inference engine.