Skip to main content

WAI Extension: Task Fidelity

Mirrored from the canonical text at commit 117bad22 ().

Status: Draft. Names the contract for media whose correctness is a downstream model’s behaviour, not a byte comparison — the class determinism-tiers cannot express. Sibling to that file, not a tier inside it: the tiers there are reproducibility guarantees a sink can check from the bytes alone, and a task claim is not one. No new wire format and no new receipt mechanism; a claim shape + an evidence grading over §7, modelled on energy-measurement’s acquisition class. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.

1. Why a second axis

determinism-tiers.md answers one question: does the sink reproduce the same bytes? A growing class of coded media is not answering that question at all. Three independent arrivals, none of them WAI’s:

WAI today can register these as capabilities but cannot state what conformance means for them. ReplicateClass::Behavioral is a disclaimer, not a criterion. Task-equivalent but not bit-equal is unstateable, so an implementation either overclaims a reproducibility tier it does not meet or says nothing at all.

The distinction that makes this a separate file. A reproducibility tier is checkable by the sink from the bytes it received. A task-fidelity claim is not: checking it requires a model, a corpus, and a test run. Listing the two together would blur the one property determinism-tiers.md currently keeps clean — that its tiers are verifiable without leaving the object.

2. The claim

A task-fidelity claim is a tuple, and it is meaningless with any element missing:

elementmeaning
modelcontent address of the network the claim is about — the identity under which δ was observed
testcontent address of the task, corpus and metric the claim was measured under
toleranceτ — the behavioural change the claimant asserts is acceptable, in the metric test names
observedδ — the behavioural change actually measured under test
evidencethe class in §3

A claim asserts δ ≤ τ under test, for model. It asserts nothing about any other model, any other corpus, or any metric test does not name.

3. Evidence class (REQUIRED)

What separates the rungs is provenance and scope, not accuracy — the same principle as energy-measurement.md §2, and for the same reason: a confidently formatted δ says nothing about whether anyone ran the test.

classmeaningrequires
Measuredδ was observed by running test against model on the stated corpus.retained artifacts sufficient to re-run
Transferredδ was observed under a different model or corpus and is argued to carry over.the source claim, and the argument, both named
Assertedδ is claimed without an observation.—

4. What the claim is, and is not

5. The split-computation case — named, not solved

FCM makes the origin/sink work partition a negotiated wire parameter: the cut point N decides how much of the network each side runs. That breaks an assumption in jwp-receipts.md, where joules_micro is one u64 for producing the object by one producer.

A cut-point-aware receipt needs three things WAI does not have: joules before the cut, joules attributable after it, and a network-identity hash making the halves provably compose. The third is a carriage problem — the artifact a network identity names has no normative wire form in WAI today.

This extension does not solve that. It records the gap so a task-fidelity claim over a split computation is not mistaken for a complete accounting: such a claim MUST state the cut point, and MUST NOT report a single joule figure as if one producer had done all the work.

6. Framing

Byte-comparability is becoming one fidelity contract among several rather than the only one. WAI’s position is unchanged by that — it names contracts and grades the evidence behind them, and does not adjudicate which contract a domain should want. What would be wrong is silence: a ratified class of media arriving with a conformance notion the container cannot state, described instead in the vocabulary of a guarantee it does not make.