Cacheon

Threat model

Cacheon evaluates attacker-supplied Python, patches, and native GPU code. The production design assumes the candidate engine is fully hostile. Static scanning, typed manifests, and correctness tests reduce exposure; they are not the primary containment boundary.

This page describes implemented controls and the risks they do not close. It applies to the finalized-intake and OCI qualification path, not to direct development diagnostics.

Security objective

Cacheon does not attempt to make candidate code trustworthy. It tries to ensure that a hostile candidate can influence only the bounded engine computation it was admitted to, and that no economic or product transition occurs unless trusted, reopenable evidence proves the registered conditions.

That produces three independent control goals:

  1. Containment: candidate code cannot reach ambient secrets, networks, state, or later workloads through an authorized interface.
  2. Measurement integrity: the host, not the candidate, owns identity, workload, role, timing, output buffers, quality authority, and verdict construction.
  3. Promotion integrity: crown, settlement, weight publication, integration, signing, and serving are distinct authenticated transitions.

A control can satisfy one goal without satisfying the others. No-egress OCI improves containment but does not make candidate-reported throughput trustworthy. A valid crown proves registered measurement, not safe production source.

Protected assets

  • Validator wallet and chain-signing authority.
  • Model weights and exact model identity.
  • Incumbent and candidate proposal bytes.
  • Workload selection secrets and hidden tasks.
  • Host-owned timing, role assignment, and device observations.
  • Raw referee evidence, calibration, and SQLite state.
  • Evaluation-stack and reward-family state.
  • Release signing keys, reviewed source, registry identity, and serving model.

Trust boundaries

ComponentTrust position
Chain and finalized event historyExternal consensus/order authority; availability and correctness assumed according to deployment policy
HTTPS originUntrusted byte transport; URL, routing, archive, and content hash are validated
Candidate proposal and native outputHostile
Candidate OCI engineHostile process inside constrained runtime
Trusted controller and arena providerValidator authority; must be reviewed and operated securely
Pristine T workerCandidate-free quality authority, still dependent on reviewed runtime/model/reference
SQLite and evidence storeTrusted durable state, protected from unprivileged mutation; host administrator remains in scope
Private recovery object storeOff-pod availability mirror of selected digest-bound state; never live SQLite/evidence authority; credentials, access policy, encryption, retention, and rollback resistance remain operator-controlled
Shared-weight object storeMutable availability layer; authenticated offer provenance is required in push-enabled mode, while push-disabled raw mode is explicitly operator-trusted
Weights gatewayTrusted permit/authentication boundary with no chain-signing authority; it must retain push verification secrets and protect its response hotkey
Weight signer and release signerHigh-value control plane, kept outside evaluator containers
Reviewed release containerTrusted product artifact after verification; not treated as a hostile candidate sandbox

Threats and controls

ThreatImplemented controlResidual or assumption
Hostile URL, SSRF, redirect, or archiveHTTPS only; TLS 1.2 minimum; globally routable DNS answers; pinned reviewed IP with SNI/hostname verification; every redirect revalidated; bounded raw gzip/tar preflight including PAX/GNU extension payloads; strict member/size/path rules; committed hash rederivedCA/DNS/origin compromise can affect availability; network stack and TLS library remain trusted
Proposal changes after commitContent hash is checked after extraction and again across immutable publicationHash does not establish authorship, license, or safety
Candidate imports or patches trusted controllerController parses candidate as data and never imports candidate Python/native; complete engine runs in a separate OCI workerA container/kernel/runtime escape can cross the boundary
Candidate exfiltrates model/evidenceRuntime has no network, read-only root, exact read-only mounts, private cache tmpfs, bounded protocol; prebuild has no model/GPU/network/home/walletGPU/driver side channels, co-tenancy, host compromise, and undiscovered runtime flaws remain possible
Candidate tampers with timer or roleHost assigns physical lanes and the exact v7 B/C/[B′] or v8 B/C/B′ roles, owns clocks, validates bounded raw batches/token counts, observes device state, and controls teardown; a reproduction must exchange lane rolesHost clocks, firmware, driver, and provider scheduling must be trustworthy and calibrated
Candidate fakes qualityAny required audit is collected in a separate eager/untimed role and host-regraded; candidate speed lifetimes are destroyed before candidate-free pristine T teacher-forces sealed trajectoriesReference bugs, audit sampling limits, and finite hidden-work coverage remain possible
Candidate behaves only on known shapes/promptsPost-commit selection, hidden work, typed graph requirements, and registered decode/long-prefill mixtureWorkload overfitting cannot be eliminated; corpora and regimes need ongoing governance
Candidate exploits noiseFrozen v7 invariant B/C bounds with conditional B′ or precommitted v8 B/C/B′, v5+ bracket-drift exclusion, physical-lane-swapped reproduction, lower reproduced speedup, and complete per-rank execution evidence before gradingHardware drift, boot-state outliers, cross-validator variance, and candidate-process receipt forgery remain operational concerns
Candidate hangs or exhausts resourcesStage deadlines, CPU/memory/PID/file/shm/tmpfs bounds, cohort admission, retry budgets, forced container cleanup, durable leasesA GPU/driver hang may require host reset; sustained spam can still consume bounded capacity
Unpaid or replayed intake spamOptional eval-cost gate (eval_cost_tao_rao, default off): one transfer_keep_alive to the subnet owner coldkey, remarked to one hotkey + content hash + netuid, consume-once on reserved or deferred admission; a private operator may grant one audited artificial credit for one unpaid revealThe gate does not meter GPU work after admission. A stolen miner hotkey can spend that hotkey's unused pointer. There is no refund. A credit is a privileged local bypass and must be audited; it is hotkey-scoped, oldest-first, consumed only on admission, and never rescues an invalid payment pointer
Candidate persists into later workEphemeral containers, read-only mounts, private tmpfs/cache, lease-scoped resources, restart recovery, post-run quiescence checksHost/container-runtime compromise can persist beyond these controls
Copy/front-run attackNative timelock, finalized event priority, submitted-delta exact/normalized/containment fingerprintsObfuscated or independently convergent implementations may evade or collide; structural similarity is advisory
State replay or partial settlementChain-scoped single-writer SQLite, durable cursor/statuses, content-addressed evidence, lease generations, atomic settlementDisk loss, privileged database edits, and faulty backup/restore are operator risks
Recovery archive is truncated, substituted, oversized, or staleSQLite online backup; closed manifest and content digests; bounded reads; reopen-after-upload; fresh-root semantic restore of SQLite, publications, evidence, and redacted journalObject-store writers can deny service or replay an older valid snapshot; operators must pin the intended manifest, deny public access, protect credentials, configure encryption/versioning or object lock, and review any live cutover
Recovery archive leaks sensitive stateClosed automatic scope excludes wallets, credentials, models, OCI images, caches, unredacted logs, and unrelated evidence; sealed inputs require explicit names; operational journal is redactedExplicit sealed inputs and source-path metadata may still be sensitive; no client-side encryption is provided, so bucket policy and deployment encryption remain required
Weight publication ambiguitySeparate signer; live metagraph refresh; intent-before-submit journal; exact recipient-set plus fixed-tolerance normalized-value readback and last_update; held stateHotkey theft, malicious operator, chain faults, and policy disagreement remain external risks
Shared-weight forgery, rollback, or replayEval push and its fresh exact acknowledgement are HMAC-authenticated; push-enabled storage retains an HMAC envelope over credential id and offer digest which the gateway verifies before response signing; current-offer writes reject block rollback/same-block conflict; followers pin gateway authority, require live permit, bound initial staleness, verify stable UIDs, and retain a monotonic signer journalNetwork/object-store writers can deny service, and object-store writers can replay a previously valid envelope; a fresh follower may accept that replay within the configured freshness window; push-secret, gateway-hotkey, wallet, or host-root compromise remains authoritative. Push-disabled raw storage deliberately trusts its operator
Crown automatically reaches productionIntegrated-only release manifest, model seal, deterministic artifacts, SBOM/provenance, Ed25519 signature, expected key, registry reproducibility attestationIntegration review, key custody, base image/toolchain security, registry, and rollout policy remain human/operational authorities

Candidate isolation boundary

Production uses two candidate stages:

  1. A native prebuild container parses the materialized tree and invokes only validator-registered build patchers. It has no network, GPU, model, wallet, Docker socket, or ambient host mount.
  2. A GPU runtime container receives only the exact model, materialized engine tree, reopened native publication, selected GPUs, private runtime cache, and bounded session protocol.

Both use a digest-pinned local image, read-only root, dropped capabilities, no-new-privileges, reviewed seccomp, non-root UID/GID, resource bounds, private /tmp, and no network. See Isolation.

Abuse stories

Concrete attacker stories help reviewers test the composition of controls:

“I will make stock code look like my acceleration”

The candidate tries to remain ineligible or fail and rely on fallback while the server still answers. Pre-selection stock routing is allowed for ordinary availability. Direct-artifact and one-shot qualification require positive execution coverage and treat selected-path fallback as invalid evidence. Pair-native hot-swap execution counts are minted per generation, and current source holds a leg unless every expected rank fired and completed the candidate under that activation generation. These receipts close accidental non-invocation but not deliberate in-process forgery. End-to-end host timing is bound to the exact candidate launch identity.

“I will grade outputs that I know how to game”

The candidate does not choose prompts, selection entropy, role, or quality reference. Its trajectory is sealed, the candidate is destroyed, and pristine T teacher-forces the selected trajectory and hidden work. This reduces self-grading and known-prompt attacks; it does not eliminate finite-workload overfitting.

“I will persist into the next candidate”

Writes are lease-scoped and ephemeral, roots/mounts are bounded, no network is available, and the host force-cleans labeled resources then proves device/process quiescence. If it cannot prove absence, the attempt is not converted into a candidate verdict and the host must be drained. Kernel/driver escape remains residual risk.

“I will turn a crown into production code”

Evaluation manifests may name hostile proposals; release manifests cannot. Integration must preserve the selected payload identity while separately approving provenance, license, security, compatibility, and tests. The signed release, external expected key, reproducible registry identity, host authorization, and serve receipts are later gates.

Control-plane separation

Evaluator workers never receive chain keys. chain-validate does not sign weights. Legacy set-weights runs separately against durable state and live chain readback.

The preferred shared-weight split keeps eval off the chain-signing path: push-weight-offer authenticates one exact offer to serve-weights; the gateway verifies its authenticated storage envelope and serves only callers with a live validator permit; follow-weights rebinds and submits from a separate signer host. The gateway's response signature authenticates the verified stored offer, not its economic correctness. Followers still apply freshness, UID, journal, and chain-readback gates. Running the gateway without push credentials changes the boundary: its configured storage becomes trusted operator input.

Release signing is separate again. A release build context contains a public verification key, never private signing material. The serving release uses reviewed source and may use host networking; it is not the hostile candidate container.

Non-authoritative development paths

scan and verify are contributor diagnostics. Contributor-controlled matched A/B profiling may load candidate code in model workers, but its process boundary is not equivalent to the production OCI controller, evidence graph, or pristine reference. No contributor-controlled execution path is acceptable for crownable work.

Explicit nonclaims

Cacheon does not claim:

  • formal verification of Python, CUDA, OCI, the Linux kernel, GPU firmware, or drivers;
  • prevention of every side channel or cross-tenant attack;
  • complete detection of plagiarism or semantic equivalence;
  • immunity to denial of service or evaluator-capacity exhaustion;
  • universal performance or quality beyond the registered arena;
  • protection against a malicious host administrator, arena provider, reference, or release signer;
  • Byzantine agreement among validators; or
  • that signatures, SBOMs, and reproducible digests replace security and license review.

Security status should be stated as “implemented under these assumptions,” not simply “closed.”

Security-review checklist

  • Identify whether the change affects intake, prebuild, runtime, reference, settlement, signing, or serving; do not reuse controls from a different boundary by analogy.
  • List new bytes, mounts, devices, network paths, environment keys, protocol fields, patchers, subprocesses, and durable state visible to hostile code.
  • Confirm every candidate-controlled value is bounded and authenticated before allocation, import, compilation, launch, or persistence.
  • Exercise malformed, oversized, timeout, crash, partial-write, restart, and cleanup-race paths; verify ambiguous infrastructure cannot mint FAIL or PASS.
  • Prove stock fallback cannot masquerade as candidate execution in strict mode.
  • Reopen the resulting evidence and reconstruct the decision without process memory or operator narrative.
  • Check that evaluator keys, release keys, wallet state, Docker control, model provenance, and hidden-work bytes stay outside candidate interfaces.
  • State residual risk explicitly, including host root, kernel/container runtime, GPU driver/firmware, registry, reference policy, and signing-key custody as applicable.

Source anchors

On this page