WAI Extension: Task Fidelity
Mirrored from the canonical text at commit 117bad22 ().
Status: Draft. Names the contract for media whose correctness is a downstream model’s behaviour, not a byte comparison — the class
determinism-tierscannot express. Sibling to that file, not a tier inside it: the tiers there are reproducibility guarantees a sink can check from the bytes alone, and a task claim is not one. No new wire format and no new receipt mechanism; a claim shape + an evidence grading over §7, modelled onenergy-measurement’s acquisition class. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.
1. Why a second axis
determinism-tiers.md answers one question: does the sink reproduce the same
bytes? A growing class of coded media is not answering that question at all.
Three independent arrivals, none of them WAI’s:
- Video coding for machines. ISO/IEC 23888-2 (MPEG-AI Part 2; DIS, FDIS planned late 2026) codes video and descriptors extracted from it, jointly optimised for bitrate and for machine-task performance after decoding. Its fidelity target is a task’s accuracy, not a pixel comparison.
- Feature coding for machines. ISO/IEC 23888-4 (MPEG-AI Part 4; CD) codes intermediate feature tensors from inside a machine-vision network — the origin runs the first N layers, ships the tensor, the sink runs the remainder. No pixels are reconstructed at any point, so there is no medium to hash.
- Vendor framing. At least one vendor markets compression as safe for machine consumption when the change it induces in a downstream model’s behaviour is statistically indistinguishable from the variance that model already tolerates. No standards body, no designation — and the vendor’s own material states the industry lacks a shared framework for the claim, which is what makes the seam worth naming rather than ceding.
WAI today can register these as capabilities but cannot state what conformance
means for them. ReplicateClass::Behavioral is a disclaimer, not a criterion.
Task-equivalent but not bit-equal is unstateable, so an implementation either
overclaims a reproducibility tier it does not meet or says nothing at all.
The distinction that makes this a separate file. A reproducibility tier is
checkable by the sink from the bytes it received. A task-fidelity claim is not:
checking it requires a model, a corpus, and a test run. Listing the two together
would blur the one property determinism-tiers.md currently keeps clean — that
its tiers are verifiable without leaving the object.
2. The claim
A task-fidelity claim is a tuple, and it is meaningless with any element missing:
| element | meaning |
|---|---|
model | content address of the network the claim is about — the identity under which δ was observed |
test | content address of the task, corpus and metric the claim was measured under |
tolerance | τ — the behavioural change the claimant asserts is acceptable, in the metric test names |
observed | δ — the behavioural change actually measured under test |
evidence | the class in §3 |
A claim asserts δ ≤ τ under test, for model. It asserts nothing about any
other model, any other corpus, or any metric test does not name.
- A claim MUST carry all five elements. A tolerance without an observation is a target, not a claim, and MUST NOT be sealed as one.
modelandtestMUST be content addresses, not names. “Tested on a standard detector” is unfalsifiable; a hash is not.- A claim MUST NOT be restated for a different model or corpus without a new observation. Behavioural tolerance does not transfer by assertion.
3. Evidence class (REQUIRED)
What separates the rungs is provenance and scope, not accuracy — the same
principle as energy-measurement.md §2, and for the same reason: a confidently
formatted δ says nothing about whether anyone ran the test.
| class | meaning | requires |
|---|---|---|
Measured | δ was observed by running test against model on the stated corpus. | retained artifacts sufficient to re-run |
Transferred | δ was observed under a different model or corpus and is argued to carry over. | the source claim, and the argument, both named |
Asserted | δ is claimed without an observation. | — |
- An implementation MUST NOT promote a claim’s evidence class in transit, as
§2 of
energy-measurement.mdforbids for acquisition class. - A verifier that cannot resolve
modelandtestMUST treat the claim asAssertedregardless of the class stated, and SHOULD say so rather than discard the claim silently. TransferredMUST NOT be used to cross a modality or a task family. A detector’s tolerance is not a segmenter’s.
4. What the claim is, and is not
- It is not a reproducibility tier, and MUST NOT be presented as one. A
capability carrying a task-fidelity claim MUST NOT thereby claim
entropy-consistencyordecode-equivalence; those are separate assertions about separate properties, and a capability may hold both, one, or neither. - It is not sink-checkable from the bytes. This is the defining limitation, and stating it is the point: a receipt bearing a task-fidelity claim carries a graded assertion, not a verification, and a relying party that treats it as the latter has misread it.
- It is not a quality metric. δ is behavioural change in a named downstream task, not perceptual fidelity, and the two MUST NOT be reported in the same field.
- It does not license degradation elsewhere. A task-fidelity claim says
nothing about energy, and a payload an encoder altered to save energy, such as
one produced in answer to a decoding-operation-reduction request, is still bound
by
determinism-tiers.md§2.
5. The split-computation case — named, not solved
FCM makes the origin/sink work partition a negotiated wire parameter: the cut
point N decides how much of the network each side runs. That breaks an assumption
in jwp-receipts.md, where joules_micro is one u64 for producing the object
by one producer.
A cut-point-aware receipt needs three things WAI does not have: joules before the cut, joules attributable after it, and a network-identity hash making the halves provably compose. The third is a carriage problem — the artifact a network identity names has no normative wire form in WAI today.
This extension does not solve that. It records the gap so a task-fidelity claim over a split computation is not mistaken for a complete accounting: such a claim MUST state the cut point, and MUST NOT report a single joule figure as if one producer had done all the work.
6. Framing
Byte-comparability is becoming one fidelity contract among several rather than the only one. WAI’s position is unchanged by that — it names contracts and grades the evidence behind them, and does not adjudicate which contract a domain should want. What would be wrong is silence: a ratified class of media arriving with a conformance notion the container cannot state, described instead in the vocabulary of a guarantee it does not make.