WAI Extension: Determinism Tiers for Neural Decode
Mirrored from the canonical text at commit 117bad22 ().
Status: Draft. Names the spectrum of reproducibility guarantees a neural capability can make, so a manifest can declare which one it satisfies and a verifier knows what it may check. Defines two tiers —
entropy-consistency(the entropy decode reproduces) anddecode-equivalence(the whole decode to the final medium reproduces) — and states which oneintent = replicateover a neural capability promises, what a receipt can hash under each, which WAI capabilities are registered at each and at what cost, and what the repository demonstrates today. One optional manifest member carries a tier name: a pin’stier(SPEC.md §3.1). Otherwise a vocabulary and conformance framing over §7. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.
1. Why tiers
Neural decode reproducibility is not binary. A neural decode evaluated in float
through an inference runtime does not, in general, reproduce its reconstruction
across hardware. Each IEEE-754 basic operation (add, subtract, multiply, divide,
square root) is correctly rounded, so it gives the same result for the same
operands everywhere; the differences come from what a runtime may choose per
device — kernels and accumulation order, vector width, fused multiply-add, reduced
precision, math-library functions — and from the process’s floating-point
environment. A decoder that also derives its entropy parameters in float can,
from one diverging bit, select a different distribution from the encoder’s and
corrupt the entropy decode. A float stage whose operations, order, precision and
environment are all fixed does reproduce, within the implementation that fixes
them; §5 describes one, and §2–§3 say why that is not a tier and does not back
replicate. Three published routes to reproducibility are below. Two make the
entropy stage reproducible while leaving the reconstruction to pixels to the
implementation, within whatever criterion a conformance specification sets on it:
- The base standard. JPEG AI (ISO/IEC 6048) defines its bit-exact conformance point at the output of the entropy decoder, not at reconstruction to pixels; its entropy pipeline is integer so that stage replicates across devices, while the outputs of later processes “may diverge slightly from the reference”. One codestream may legally reconstruct through up to three synthesis transforms, which a profile or the encoder can restrict. Bit-exact picture reconstruction was left out of the first version on purpose. An overview of the standard (arXiv 2510.13867) states all of this. It also says that the conformance part, ISO/IEC 6048-4, “specifies the requirements for generating a JPEG AI-compliant reconstruction”, with a testing suite for each profile and level pair, and “considers the possibility that a decoder is conformant to the standard without having bit-exact reconstruction” — “the possibility of standard compliance with a certain leeway”. These passages point to a criterion around a reference on the reconstruction; the overview does not state one. Part 4 was a draft when the overview was written; this document has not read it, so it does not say whether Part 4 sets such a criterion, of what form, or how wide.
- Explicit parameter transmission. A published cross-platform learned video codec (arXiv 2606.28027) transmits its entropy model’s scale parameters explicitly, through the hyperprior, so the decoder reads them instead of deriving them and the entropy decode is consistent across devices without bit-exact arithmetic. The side information costs bitrate. Same target: a reproducible entropy stage, a reconstruction that may vary.
The third makes the reconstruction itself reproducible:
- A quantised decoder. A published learned image codec (arXiv 2312.11209) quantises the weights and activations of all its decoder subnetworks, without accumulator overflow, so that the reconstructed image is deterministic across devices, and reports its output as bit-exact.
WAI names the two outcomes as tiers, so a capability states which one it makes, an
envelope’s intent can demand the one it needs (replicate over a neural
capability demands decode-equivalence, §3), and a receipt or a provenance binding
(reconstruction-binding) can do the same.
2. The tiers
entropy-consistency
The entropy decode reproduces byte-for-byte across devices — the decoded symbols (the entropy-decoder output) have a portable hash — but the reconstruction to the final medium (pixels, samples) MAY differ between conforming sinks. A sink achieves this either by integer entropy arithmetic (deriving the coding parameters with fixed-point ops, as WAI’s integer decoders and JPEG AI’s entropy pipeline do) or by explicit parameter transmission (the codestream carries the parameters, so no bit-exact derivation is required).
- Conformance criterion: entropy-equivalence — a verifier decodes the entropy stage and compares the decoded symbols, NOT the pixels.
- What a receipt can hash: the codestream and the decoded symbols. The latent formed from the symbols is hashable only where every step that forms it (adding a predicted mean, say) is itself integer or transmitted. The reconstruction has no portable hash, so an energy figure at this tier binds to the codestream and symbols, never to an output every conforming sink reproduces (energy-binding §3); a reconstruction that one implementation pins is that implementation’s (§3).
- JPEG AI sets its bit-exact conformance point at this tier, and explicit
parameter transmission reaches it (§1). The overview also says JPEG AI’s
conformance part “considers the possibility that a decoder is conformant to the
standard without having bit-exact reconstruction” (§1); whether that part sets a
criterion on the reconstruction is not read here, and any such criterion is not
one of this document’s tiers. WAI registers
wai.image.jpegaihere (§5). The tier is necessary for a sound neuralreplicateunder SPEC §7 — without it the decode desyncs — but not sufficient, and it MUST NOT backreplicate:replicatepromises the reconstruction, and this tier leaves the reconstruction free.
Why transmitting parameters stops at this tier. A learned entropy decode desyncs when the decoder derives an entropy parameter (a scale, and from it the distribution the next symbol is decoded with) that selects a different distribution from the one the encoder used. Transmitting the parameter removes that derivation from the decoder: every decoder reads the same value and decodes the same symbols, whatever arithmetic its networks run in. Nothing after the entropy decoder is constrained. The steps that turn symbols into a latent and the synthesis transform still run in the arithmetic the device provides — float, reduced precision, fused or reordered accumulation — and, unless the implementation fixes that arithmetic, their results differ by device. A codestream can pin values; it cannot pin the arithmetic a network is evaluated in, and the reconstruction depends on that arithmetic. So a portable hash of the reconstruction needs every stage after the entropy decode fixed as well, which is the next tier. The two routes compose: a capability MAY transmit parameters and pin its arithmetic.
decode-equivalence
The whole decode to the final medium reproduces byte-for-byte on every
conforming sink — the reconstructed pixels (or samples, or attributes) have a
portable hash. The capability’s registration pins one reconstruction and
specifies it in integer / fixed-point arithmetic end to end, entropy and
synthesis, so the output depends on nothing a build, a device or a process
chooses. Where the base codec admits several conformant reconstructions, this
tier fixes one. This is what intent = replicate over a neural capability
promises (§3).
A tier is a property of the capability as registered, not of one implementation. An implementation that pins a reconstruction by evaluating float stages in a fixed order (§5 describes one) reproduces it only under conditions of its own build and process (§3). That is a property of the implementation, stated as such; it does not place the capability at this tier.
- Conformance criterion: the medium’s
*-equivalence(e.g. decode-equivalence forwai.neural.int_hyper, sample-equivalence forwai.audio.int_codec, splat-equivalence forwai.splat.int_codec). - What a receipt can hash: everything the lower tier can, plus the reconstructed
medium itself as the capability’s canonical buffer (canonical
RGB8for images, per reconstruction-binding; int16 PCM; the attribute tensor; a per-frame hash for video), under the descriptor that pins the reconstruction. This is the tier at which an energy figure binds to an output a verifier can reproduce (energy-binding §3). - Where a base standard sets its conformance point at the entropy decoder (JPEG
AI, §1), a capability over it reaches this tier only by registering one
reconstruction in integer arithmetic;
wai.image.jpegaihas not (§5). For any capability this tier is the precondition for a hard provenance binding over the rendered content that every conforming sink can check — a decoded-pixel hash is only portable when the reconstruction is pinned.
A capability satisfying decode-equivalence satisfies entropy-consistency by
construction; the reverse does not hold.
What a tier covers
A tier is a property of the whole path from codestream to the medium that is hashed, bound, or presented — not of the decode stage in isolation. A sink that decodes conformantly and then alters the result has not satisfied the tier, however faithful the decode itself was.
This is not hypothetical. Published and draft standards define mechanisms that change the medium after, or alongside, a conformant decode, and they arrive through the same pipe as the codestream:
-
Post-decode neural filters. ITU-T H.274 (V4, 01/2026)
nnpfc/nnpfaSEI messages activate a neural post-filter after decoding, changing output pixels. The filter may be named by URI or carried in-band. -
Energy-saving presentation and quality-recovery metadata. ISO/IEC 23001-11 (green metadata) defines metadata with which a receiver changes the decoded pictures, before display or inside the decode:
- Attenuation-map information, carried in an SEI message with the attenuation map as an auxiliary picture, describes post-processing to apply to the decoded pictures so the display emits less light. A receiver can also ask the transmitter for such maps with a display attenuation map power reduction request (DAMPR-Req); the maps it receives are applied to the decoded content in the same way.
- RGB-component statistics and quality levels drive display adaptation. The
receiver chooses a quality level, scales the RGB components of the
reconstructed frames by the peak signal over that level’s
max_rgb_component, and lowers the backlight or voltage in proportion (Annex B.2). At any level other than the no-quality-loss operating point, components above that level’smax_rgb_componentare saturated; for contrast enhancement, components at or belowlower_boundare set to zero and those at or aboveupper_boundto the peak signal (clause 7.4). - Quality metrics for cross-segment decoding: a decoder uses them to decide whether to enhance the first picture of a low-quality segment from the last picture of the previous segment, and the enhanced picture becomes the reference for the pictures that follow.
The standard notes that its metadata can also be used “to get larger energy savings, but at the expense of some QoE degradation”.
Other 23001-11 mechanisms leave the decode as it is. Three of them change the payload before it arrives:
- A decoding operation reduction request (DOR-Req) goes from the receiving device to the remote encoder in a point-to-point session. The encoder produces a less complex bitstream and can answer with a response (DOR-Resp) saying how it decided to answer.
- A display power reduction attenuated video request (DPRAV-Req) asks the remote encoder to apply an attenuation map to the video before encoding it. The receiver then gets an attenuated video, and the response (DPRAV-Resp) states the energy reduction rate and quality it can expect.
- Low-power encoding, which produces the low-quality segments that cross-segment decoding repairs, alternates high-quality and low-quality segments to reduce the encoder’s power.
In each case the sink decodes what it receives like any other bitstream, but the payload is not the one that would have been sent otherwise. The remaining two alter neither a payload nor a decode’s output. Media-selection metadata lets an adaptive-streaming client choose among the available representations by their decoder- and display-power characteristics; what it receives is one of those representations as published. Complexity metrics let a decoder vary its operating frequency, which changes what a decode costs, not what it outputs.
All of this is metadata for saving power, stated in advance; none of it reports energy that was spent. Its energy quantities are expectations: an attenuation map’s expected energy saving rate, and the rates a receiver asks for and a transmitter says the receiver can expect, each a percentage (clause 7.4); a representation’s decoder-operation reduction ratios and its maximum potential display power saving (clause 8). Complexity metrics are proportions of decoding operations over a period (clause 6.2), and a DOR-Resp acknowledges what the encoder changed. So a sink that acts on this metadata has a prescription, not a measurement, and a tier claim or receipt over its output is bound by the rules below like any other.
This section accounts for every metadata set that the Introduction of the 4th-edition CD text (WG 3 N1864, dated 2026-08-28) names. Read in that text: the Introduction, and in part clauses 6, 7, 8 and 9 and Annex B.2, B.4, B.5 and B.6. The published edition was not read.
The rules:
- A sink claiming a tier MUST evaluate it over the medium it emits. If it applies processing to the decoded medium that alters it — after decode, as attenuation maps and display adaptation do, or fed back into the decode as cross-segment decoding does — it MUST either not apply the processing, or MUST NOT claim the tier for the emitted result.
- A tier claim, a decoded-medium hash or an energy figure over a payload that an encoder altered to save energy — at a receiver’s request, such as a decoding-operation-reduction or attenuated-video request, or on its own account, as low-power encoding does — covers that payload, not the one it stands in for, and a receipt over it MUST NOT be presented as covering the payload it stands in for. Spending fewer joules by computing or emitting less is a fidelity choice; a receipt that reports the joules while concealing the degradation states a cost without its consideration.
- A capability’s registration SHOULD state whether its payload can carry post-decode processing directives at all. A capability whose payload cannot is unaffected by this section, and saying so is cheaper than leaving it open.
The general rule, which is why this section exists: a determinism claim covers what the sink emits. A hash over a decode that something later modified describes an artifact nobody saw — the hash verifies, and it attests the wrong thing.
3. Declaring and using a tier
- Declaring. A capability’s registration (SPEC §5) states its tier via its
*-equivalencecriterion, and its registry class carries it (SPEC §7):NeuralIntegerExactcapabilities are registered atdecode-equivalence,NeuralEntropyExactcapabilities atentropy-consistency, andNeuralFloatcapabilities claim neither. §5 lists WAI’s capabilities at each tier. A capability MAY declareentropy-consistencywhen it only guarantees the entropy stage. A pin MAY also state a tier in the manifest, astier(SPEC §3.1), from the closed vocabularydecode-equivalence,entropy-consistency, ornonefor neither; an encoder writes only those tokens, and a sink reads any other token asnone. A stated tier above the capability’s registered tier does not conform (tier_exceeds_class): thereplicateverdict does not rest on it. - Over a neural capability,
replicatepromisesdecode-equivalence(SPEC §7). An envelope withintent = replicateover a neural capability MUST be over a capability registered atdecode-equivalence, which SPEC §7 admits only by the integer construction: an integer / fixed-point decode path with no float operation. Such an envelope MUST pin the parameter set of that capability with a pin that declaresdecode-equivalence(SPEC §3.1, §7).entropy-consistencyMUST NOT backreplicate. A neural decode that derives its entropy parameters in float meets neither tier; one that reads them from the codestream, or derives them in integer arithmetic, and runs its later stages in float meetsentropy-consistencyonly. A reconstruction that one implementation pins through fixed-order float stages does not backreplicateeither: what makes such a stage reproducible — a build that emits no fused or reassociated operations, the default floating-point environment in the running process, library functions whose last-place error cannot reach the output — belongs to one implementation and one process, and neither the envelope nor the registry can state it. That is whywai.image.jpegaiis registered atentropy-consistencyand does not backreplicate(§5). - Over a standard codec,
replicatepromises the named standard’s decoder conformance (SPEC §7, Whatreplicatepromises over a standard codec). For some standards that is the same bytes on every sink; others hold decoders to a criterion around a reference — numeric thresholds in the clauses read for Opus and JPEG XL; for JPEG and JPEG 2000 a criterion whose bound the text read does not state (SPEC §7); and, for AV1 when film grain is signalled, a perceptual criterion the AV1 specification leaves undefined — and over thosereplicatepromises that criterion, not byte identity. The tiers in this document describe neural capabilities.wai.image.jpegaihas a named standard too, and the overview’s account of its conformance part (§1) points to a criterion around a reference on the reconstruction; if Part 4 states one, that would put it among the standards whose conformancereplicatecould promise. That part has not been read here. Holdingwai.image.jpegaitodecode-equivalence, and so refusingreplicateover it, is a conservative choice pending that reading, not a finding about the standard; SPEC §7 (wai.image.jpegaiand the standard classes) states the choice and how it differs from the choice made for the standard codecs. - What the reference impl checks.
decode_envelope_strictapplies SPEC §7 in two steps (codecs::replicate_refusal). First the capability’s registry class: onlyNeuralIntegerExactpasses among the neural classes (the three standard classes, which are not neural, are not refused by this rule). The class does not depend on anything the sender declares, so areplicateenvelope overwai.image.jpegaior a float codec is refused even when its prior declares an integer contract. Then, for an integer-exact capability, the pin: it must declaredecode-equivalence— bytier, or, for a pin withouttier, by the reading of itsdeterminismstring that SPEC §7 fixes, a substring match (a contract containingint,fixedorintegerand none off32,f64,float). Either way it verifies a declaration, not the arithmetic. Neither step inspects the decode path a sink actually runs; forwai.neural.int_synthandwai.neural.int_mlicpp, theNeuralIntegerExactcapabilities the reference neural dispatch decodes, there is only the integer path. - Receipts. A receipt that carries a decoded-medium hash as reproducible by
any conforming sink MUST come from a
decode-equivalencecapability. A hash over a reconstruction that one implementation pins (§5) is that implementation’s, and a receipt MUST NOT present it as the capability’s. - Post-decode processing. A tier claim is void if the sink alters the medium after decode (§2, What a tier covers). This applies to the emitted result, not to the decoder’s internal conformance.
- Provenance. The reproducible-reconstruction binding
— a C2PA / JPEG-Trust hard binding over decoded pixels — is checkable by every
conforming sink only over a
decode-equivalencecapability; over anentropy-consistencycapability only the entropy-stage output is bindable that way. reconstruction-binding §1–§2 also define a binding overwai.image.jpegaias a binding over its exact decoder’s pinned reconstruction; that binding is checkable only by that decoder under the conditions in §5–§6.
4. Framing
Both tiers here are reproducibility guarantees: a sink can check them from the bytes it received. A separate axis exists for media whose correctness is a downstream model’s behaviour rather than a byte comparison — see task-fidelity, which is a sibling rather than a third tier precisely because such a claim is not sink-checkable from the bytes.
The two tiers are named so that each capability states which one it makes.
entropy-consistency is the tier at which JPEG AI sets its bit-exact conformance
point, and the tier WAI registers wai.image.jpegai at. decode-equivalence
pins one full reconstruction, so the rendered result, not only the entropy
stream, is reproducible and hashable; it is what replicate over a neural
capability promises. §5 lists the WAI capabilities at each tier, with what the
tier costs them.
5. Where WAI offers each tier, and what it costs
decode-equivalence
| capability | hashable medium | how the reconstruction is pinned |
|---|---|---|
wai.neural.int_synth | RGB | integer synthesis (QAT weights) |
wai.neural.int_hyper | RGB | integer hyper-decoder, integer scale → distribution step, integer synthesis (post-training quantisation of a float model) |
wai.video.int_hyper | per-frame | an int_hyper keyframe, then integer motion-compensated P-frames |
wai.audio.int_codec | int16 PCM | integer synthesis over a factorized rANS latent |
wai.splat.int_codec | attribute tensor | integer synthesis over a factorized rANS latent |
wai.neural.int_mlicpp | RGB | integer hyper-decoder; an integer entropy network re-run before each of the 2·K channel-slice × checkerboard passes, integer scale → distribution step; integer synthesis. Stream values an exact decode cannot hold are refused, not wrapped |
These six are integer throughout, are the capabilities the registry classifies as
NeuralIntegerExact, and are the neural capabilities that back replicate.
entropy-consistency: wai.image.jpegai, and its exact decoder’s pinned reconstruction
The registry classifies wai.image.jpegai as NeuralEntropyExact and registers
it at entropy-consistency. This is a downgrade: earlier in v0.4.0 the
capability was listed in the table above and classified StandardDefined, and the reference
strict check passed a replicate envelope over it whenever the prior declared an
integer contract. It moved because its reconstruction runs float stages and the
reference sink can decode it through a float ONNX Runtime path, so
decode-equivalence never held as a property of the capability (SPEC §7).
| stage | arithmetic |
|---|---|
| entropy decode (mANS) and the scale-index derivation (an integer hyper-scale network) | integer on both reference decode paths — what the capability guarantees |
| hyper-decoder, context-model and synthesis networks | 16-bit-PTQ integer in the exact decoder; float through ONNX Runtime in the fallback |
| dequantisation, latent refinement, the post-filter chain (the eICCI post-filter is a float CNN), colour conversion | IEEE-754 float in a fixed evaluation order in the exact decoder; float in the fallback |
The exact decoder’s pinned reconstruction. WAI’s exact decoder
(jpegai_decode::decode) pins one reconstruction to canonical RGB8: the
high-operation-point synthesis, then the post-filters, in the arithmetic above. It
reproduces across machines when the conditions in §6 hold: its float stages use only
IEEE-754 basic operations in the fixed order of the source, with the gain scaler’s
exp and every rounding built from those operations (float_repro), so no platform
math library is on the path; the target evaluates at the operands’ own precision;
and the floating-point environment is the default. The decoder checks the last two
before decoding (jpegai_decode::check_scope). Those conditions belong to this
implementation. That is why the pinned reconstruction is not a tier, does not back
replicate, and is bindable only as
reconstruction-binding defines. On one codestream, the
decoded bytes, and the float values the decoder reports after each stage, have been
measured identical on four platforms (§6); values inside a stage are not compared
directly.
The exact decoder implements one configuration, which its tests exercise.
jpegai_decode::check_codestream accepts a codestream only when its headers select
exactly that configuration:
- a 256×256, 8-bit, 4:4:4 picture with no display cropping and the built-in YUV→RGB colour transform;
- the high-operation-point synthesis (
synthesis_transform_id2); model_id0, whose gain, LSBS and LEF tables and residual-scaling constants are tools_0’s;beta_displacement_log0 for luma and chroma. The exact decoder does not implement the quantiser’s β displacement, so a codestream with any other β is refused, not decoded;- 160 luma and 96 chroma latent channels, one entropy substream per plane, and no region partitioning, cube flags, 3-D gain or synthesis tiling; LSBS, residual variance scaling and GRFS enabled;
- post-filters only in the forms implemented: one single-tap linear enhancement filter per chroma plane; eICCI untiled on all three planes with the model pair the weight bundle carries (short-list indices 0 and 1 — the bundle format does not record its pair, and the exporter emits only that one); the nonlinear enhancement filter untiled and unmasked, with eight weights; LEF on a luma channel;
- both headers parse to the end of their segments, and the SOZ, SORP and SORS substreams are present with no SOQ segment.
That is narrow. Of the reference encoder’s five default rate points
(cfg/BRM/default.json), only the first selects model_id 0 with β = 0; the
others select model_id 1–3 or a non-zero β and are refused, as is any codestream
to which the encoder’s bitrate matcher assigns a non-zero β.
The float ONNX path hard-codes the same models, tables and filter selection, and
refuses the same codestreams, so the reference sink refuses a codestream outside
this configuration rather than decoding it to the wrong image. Within it, the exact
decoder runs when the sink supplies its integer weight bundle and the thread’s
floating-point environment is the IEEE-754 default. Otherwise the reference sink
decodes the same capability ID through the float ONNX Runtime path, whose
reconstruction is not pinned. Both paths return the same result type, and the
reference dispatch does not report which one ran. A decoded-pixel hash or a
reconstruction binding over wai.image.jpegai holds
only for output of the exact decoder.
What decode-equivalence costs
The tier adds nothing to the wire: nothing is transmitted beyond the capability’s own codestream. The costs fall elsewhere, and the JPEG AI exact decoder pays the same ones for its integer networks:
- Fidelity to the float model. The quantised networks do not reproduce the
float model’s output, so the pinned reconstruction differs from it. For
wai.neural.int_hyper, over one pretrained model (mbt2018_mean, quality 6) and one pinned test photograph, the integer decode of the full conditional bitstream measured 35.985 dB PSNR against 36.175 dB for the same model’s float decode, a cost of 0.19 dB (macOS arm64;wai-rs/src/int_hyper_synth.rs). Quality is image-dependent, and this compares the model only with its own float decode. For the JPEG AI exact decoder,wai-rs/tests/jpegai_e2e_full.rscompares against the reference decoder’s output on one 256×256 reference codestream and asserts max|Δ| ≤ 2 with at most 5% of samples differing; the recorded result is max|Δ| = 1 on 4.446% of samples (macOS arm64, release build, 2026-09-27). CI jobs re-check both bars:wai-sota-vectors.ymlrequires theint_hypercost to stay within 1 dB on that photograph, andwai-jpegai-e2e.ymlruns the JPEG AI comparison and requires the decoded pixels to hash to a pinned value. Both have run on the forge’s Linux runner and passed on every run since 2026-09-28 (§6). - Compute. The decoders are CPU code in integer arithmetic (i64, with i128
intermediates in places; plus fixed-order float for the JPEG AI exact
decoder’s float stages, which stay single-threaded). Their convolutions, IGDN
and the JPEG AI synthesis’s attention and element-wise steps run through
int_kernels: register-blocked tiles, NEON widening multiply-accumulate on aarch64 (portable Rust elsewhere), and worker threads on native targets (one worker on wasm). Each fast kernel gives the same integer for every output as the scalar reference kernel kept beside it, which stays selectable at run time (Policy::reference()). They cannot use the float GPU / NPU paths a float learned decoder runs on. Measured throughput, on one machine and with its limits, is inresearch/decode-throughput.md§4–§5; no claim in this document depends on those figures. - Model preparation. A model is offered at this tier only once exported into the decoder’s pinned form, sink-supplied: quantised weights and requantisation constants for the integer networks (and, for the JPEG AI exact decoder’s eICCI post-filter, float weights with BatchNorm folded in). A float-only model is not offered at this tier as it stands.
Parameter transmission pays in bitrate and leaves the decoder’s arithmetic free; the integer decode pays in quantisation, compute and model preparation and adds no bits. A capability’s tier states which trade it made.
6. Evidence status
What the repository demonstrates, claim by claim, as of 2026-09-27 (main at
afc7ece), except where a row names a later run: the rows that cite run 745 read
the forge’s runs of the landing chain’s head 7cfcecfa on 2026-09-30. A test that
skips reports as passing, and a matrix leg whose runner did not run reports
nothing; neither is evidence.
The CI rows below come from the forge’s commit statuses
(/api/v1/repos/Transaction-Science/open-standards/commits/<sha>/statuses), read
for every commit on main since 2026-08-10 and for the pre-merge commits named, not
from the workflow definitions. Two workflows run the same golden-byte suite
(wai/byte-exact-conformance, cargo test --release): the cross-arch
byte-equality job of wai-conformance.yml, and wai-byte-equality.yml. Both
declared Linux x86_64, Linux arm64, macOS arm64 and Windows x86_64 legs until
2026-09-30, when the Windows x86_64 legs were taken out of both workflows because its
runner is offline; there is no Windows leg now, and no macOS x86_64 or Windows arm64
leg. No Windows x86_64 job ran for this repository after 2026-08-27: every later
Windows status on main is “Waiting to run” or “Has been cancelled”, and the forge’s
task list has no Windows task after that day. The Windows results below are history.
| claim | evidence | where |
|---|---|---|
Cross-architecture identity of the rANS coder, int_synth, and the JPEG AI symbol decode (a reference-encoder codestream, given its distribution indices) | golden-byte assertions passed on Linux x86_64, Linux arm64 and macOS arm64 on main at afc7ece (2026-09-27, wai-conformance.yml; wai-byte-equality.yml last ran on main at 3f50097, 2026-08-29, all three legs passing), and on Windows x86_64 last on main at 3e7811f (2026-08-27). Between 3e7811f and afc7ece the suite only gained the video and keyframe goldens; these goldens, the sources they decode and the suite’s lockfile are unchanged | byte-exact-conformance/; .github/workflows/wai-conformance.yml, wai-byte-equality.yml |
Cross-architecture identity of integer inter-frame video, and of int_hyper on a small learned keyframe | the same suite, passing on Linux x86_64, Linux arm64 and macOS arm64 on main at afc7ece (2026-09-27). Windows x86_64 never decoded the keyframe clip: it was added in d4c922b (2026-08-28), whose Windows legs were cancelled. Windows decoded the integer video goldens only on the pre-merge commits c9be2ad, 256e5dd, 7dfca26 and 3d34860 (2026-08-25/26), whose video sources, vectors and golden file match main but whose byte-exact-conformance/Cargo.lock differs; the change reached main at 3f50097, after the last Windows run on main | as above |
wai.image.jpegai exact decoder: the integer derivation of the distribution indices matches the reference decoder exactly | a test on one machine that skips unless its fixtures (sink-supplied weights, normative tables, codestream) are provided; CI does not provide them | wai-rs/tests/jpegai_hyperscale_calibration.rs |
wai.image.jpegai exact decoder, whole decode through the one public decode() call: it stays within a stated bound of the reference decoder’s output (max|Δ| ≤ 2, ≤ 5% of samples); two decodes in one process are byte-identical, and a reconstruction binding taken over one verifies the other | wai-jpegai-e2e.yml fetches the pinned jpegai/v1 inputs (an extraction of the reference-software assets, published under WAI’s models root), checks each file against the SHA-256 manifest tools/jpegai-e2e-v1.sha256, and runs both tests with WAI_REQUIRE_SOTA_VECTORS=1, under which a missing input fails instead of skipping. The job also fails on any skip banner in the test log, and requires the decoded pixels to hash to a pinned SHA-256. One 256×256 codestream. The workflow has run on the forge’s Linux runner (x86_64 hardware) and passed on every run since 2026-09-28. Before that, its job passed under forgejo-runner exec on the same host (2026-09-27) with the exact decoder as it stood on main at afc7ece, which calls the platform exp; there too the decoded pixels matched the pinned hash, the value measured on macOS arm64 | wai-rs/tests/jpegai_e2e_full.rs, jpegai_reconstruction_binding.rs; .github/workflows/wai-jpegai-e2e.yml |
wai.image.jpegai exact decoder across platforms: the jpegai/v1 codestream decodes to the same RGB8 bytes, and to the same bits in the float values the decoder reports after each stage (values inside a stage are not hashed), on each platform measured | measured 2026-09-28 with rustc 1.98.0, release builds, inputs checked against their SHA-256, all four on the source of the change that added this row: macOS arm64 (native, on the development machine); wasm32-wasip2 (wasmtime 49, on the same machine); Linux arm64 (glibc 2.36, a container in an arm64 Linux VM on that machine); Linux x86_64 (glibc 2.36, the x86_64 build run in that container under user-mode emulation, QEMU 7.0, not on x86_64 hardware). All nine pinned hashes matched on each platform. On all four they also matched on an earlier state of that change, before the decode’s exp and rounding moved to float_repro. Before the move the decode called the platform exp on every target, and its roundings compiled to calls to the symbols round, roundf and rint on x86_64 and round and roundf on wasm32, which the link resolves: an x86_64 Linux test executable inspected after the move, whose other code still rounds this way, took them from Rust’s compiler builtins, not the C library (for wasm32 this was not checked). Not measured: macOS x86_64 (no x86_64 translation layer is installed on the development machine), Windows, x86_64 hardware, 32-bit targets. The test skips unless the inputs are supplied. wai-jpegai-e2e.yml supplies them and requires this test to run and pass on its Linux runner; that workflow has run on the forge’s Linux x86_64 runner and passed on every run since 2026-09-28 | wai-rs/tests/jpegai_cross_platform.rs |
float_repro, the decode’s exp and rounding: the same bits over a fixed grid of inputs; exp within 2 ULPs of the host math library (1 ULP measured); rounding equal to the standard library’s | library unit tests under the library-test step of the conformance workflow (neural_int), which has run them on its macOS, Linux x86_64 and Linux arm64 runners since 2026-09-28. The pinned-bits and rounding tests never skip. The host comparison asserts the host library’s accuracy as much as this module’s. Its gate selects by target, not by host: it runs on macOS arm64, on Linux x86_64 and arm64 with glibc, and on WASI, and is reported as ignored, with the reason, elsewhere, Windows included. On the source of the change that added this row it was run against these hosts’ exp, with a worst case of 1 ULP on each: the platform math library on macOS arm64; the WASI C library under wasmtime; glibc 2.36 on Linux arm64, and glibc 2.35 in the Ubuntu 22.04 image the workflow’s Linux arm64 runner used on 2026-09-28; glibc 2.36 on Linux x86_64 under user-mode emulation, against three of the four exp variants glibc chooses among by processor feature (the variant each run used follows from the features glibc’s loader reported active): SSE2 (QEMU 7.0, whose emulated processor reports none of AVX, FMA, AVX2 and FMA4), AVX (QEMU 7.2, -cpu max with FMA and AVX2 masked by glibc’s hwcaps tunable, which leaves AVX active) and FMA (QEMU 7.2, -cpu max). On a separate probe of 20 million inputs, the SSE2 and AVX variants returned the same bits, and the FMA variant different bits for some of them. Not compared: the FMA4 variant, which glibc chooses on a processor that has FMA4 but not both FMA and AVX2, and which neither emulator offers. Not run on x86_64 hardware or on other glibc versions, where the gate still runs it; a failure there would say that the host’s exp and this one are more than 2 ULPs apart, not that float_repro’s bits changed | wai-rs/src/float_repro.rs tests |
| The exact decode path calls no platform math library | two checks. Clippy’s disallowed_methods (wai-rs/clippy.toml), denied in each module on the decode path, rejects any use of a listed float method, including f64::exp(x) and .map(f64::exp). tests/jpegai_decode_symbols.rs lists the symbols of an executable that links only the decoder’s entry points and weight loaders. It rejects any C math-library function the executable imports, and any it defines other than the round-to-integral functions (floor, round, rint and the like), whose results IEEE 754 fixes. It allows those definitions because a definition shows what the linker kept, not what the decode calls: linked with GNU ld 2.40 for x86_64 Linux, the debug executable defined the compiler builtins’ floor, though nothing in it referred to floor. On the source of the change that added this row, both checks pass on macOS arm64, the symbol check in debug and release builds. The symbol check also passes, in debug builds, on Linux arm64 (the container above) and on Linux x86_64 under user-mode emulation, linked with GNU ld (reporting that floor as allowed) and with LLD. With a platform exp put back as f64::exp(…), both fail on macOS arm64 while the decoder’s unit tests still pass, and the symbol check fails on Linux x86_64 with GNU ld. With a float % put into the decode, the symbol check fails on macOS arm64 and clippy, which does not see operators, passes. The conformance workflow runs both on its unix legs, on the forge’s runners since 2026-09-28 | wai-rs/clippy.toml, wai-rs/tests/jpegai_decode_symbols.rs, .github/workflows/wai-conformance.yml |
wai.image.jpegai scope: the reference codestream is accepted; copies with β, model_id, the colour transform or an LSBS flag changed are refused; a flush-to-zero or round-toward-zero environment is detected. The float ONNX path calls the same check | library unit tests over a committed copy of the reference codestream, which never skip and fall under the library-test step of the conformance workflow (neural_int); the environment test runs on arm64 only. Run on the four platforms above, on the source of the change that added the cross-platform row (the environment test on macOS arm64 and Linux arm64). The float path’s refusal (β and model_id edits) is tested only with its ONNX bundle supplied (jpegai_scope_ort.rs, skips otherwise) | wai-rs/src/jpegai_decode.rs tests, wai-rs/tests/vectors/jpegai/ |
SPEC §7’s replicate rule: the registry class decides first, so an integer-declared prior over wai.image.jpegai or a float codec is refused; wai.image.jpegai is the only NeuralEntropyExact capability; the standard codecs sit in the classes SPEC §7 lists, and the tolerance and unclassified ones are not refused; an encoder that pins no prior (wai wrap) writes replicate only over the standard classes, for the capability and for any declared fallback, or, in the WAI2 companion form, for the base payload’s capability and each companion’s | library unit tests (codecs, container) that never skip and fall under the same library-test step. The strict-decode test in tests/neural_native_decode.rs needs the neural feature; wai-neural-gates.yml runs it on the Linux runner, in an Ubuntu 24.04 image because the neural test binary does not link in that runner’s default 22.04 image. Its strict-decode job ran on the forge and passed on every run from 2026-09-28; before that the job passed under forgejo-runner exec on the Linux runner’s host (Linux x86_64, 2026-09-27). The pin-aware tests since added to that workflow (its longer exact list of neural_native_decode tests, the step of model-free neural library tests, and the neural C ABI step, tests/ffi_neural_pin.rs) have no recorded forge run yet. The wai binary no longer needs the codecs feature, so the tests that run wai wrap (cli_wrap_intent.rs, and cli_registry.rs for the candidate refusal and the registry export) run in the conformance workflow’s carriage step on every leg; this row records no forge run of them yet. The C envelope packers emit through Wai::to_emit_bytes and WaiMulti::to_emit_bytes, whose candidate refusal a library test checks with no feature (the_emit_path_refuses_a_candidate, in the library-test step); the C wrappers themselves build under the codec-free ffi feature, and tests/ffi_envelope.rs (envelope_pack_refuses_a_candidate, envelope_select_never_picks_a_candidate) runs in the conformance workflow’s C ABI envelope step on every leg; this row records no forge run of that step yet. The browser writers’ refusal runs in wai-web-wasm.yml (verify_pack.mjs) | wai-rs/src/codecs/mod.rs, container.rs tests; wai-rs/tests/neural_native_decode.rs, cli_wrap_intent.rs, cli_registry.rs; wai-web/demo/verify_pack.mjs |
wai.neural.int_hyper and wai.video.int_hyper at full model scale (the pretrained mbt2018_mean, exported to integers): the Rust decode reproduces the export tool’s integer reference byte for byte (synthesis, hyper-decoder, the whole conditional bitstream, an eight-frame clip), and costs at most 1 dB against the same model’s float decode on one pinned photograph | wai-sota-vectors.yml regenerates the vectors on the runner from pinned inputs (the photograph’s pixels and the checkpoint’s bytes by SHA-256, the Python stack by version) and runs both tests with WAI_REQUIRE_SOTA_VECTORS=1, under which a missing vector fails instead of skipping. The job also fails on any skip banner in the test log, and runs require_vectors_rule.rs, which checks that without the switch a variable naming absent or incomplete inputs fails as well. Rust is checked against the export of the same run on the same machine: the vectors come from float computation, and some of them differed between machines and between runs, so none of their hashes is pinned and none is compared across machines. The workflow has run on the forge’s Linux runner (x86_64 hardware) and passed on every run since 2026-09-28 | wai-rs/tests/int_hyper_synth.rs, int_video_sota_keyframe.rs, require_vectors_rule.rs; .github/workflows/wai-sota-vectors.yml |
wai.audio.int_codec, wai.splat.int_codec | committed vectors that never skip. Their frozen hashes are the output of an independent Python integer reference, computed when the vectors were exported (tools/wai_int_audio_export.py, tools/wai_int_splat_export.py); no Python runs in CI. The conformance workflow’s fast-kernel steps enable audio_int and splat_int and, in debug and release builds, decode both vectors against those hashes under the reference kernels and every fast-kernel policy (tests/int_kernels_vectors.rs); that passed on the forge’s Linux x86_64, Linux arm64 and macOS arm64 runners for the landing chain’s head 7cfcecfa (run 745, 2026-09-30). The same steps also name tests/int_audio.rs and tests/int_splat.rs, which check each vector against the same hash and against its own metadata (sample count; attribute count, tensor size and grid side); Cargo declares each with its feature as a required feature, so a step that names it without that feature fails instead of running nothing. This row records no forge run of those two | wai-rs/tests/int_audio.rs, int_splat.rs, int_kernels_vectors.rs; .github/workflows/wai-conformance.yml |
wai.neural.int_mlicpp on its committed vector (synthetic seeded weights; a 64×64 image, 8 latent channels in 4 slices): the decode reproduces the export tool’s independent integer reference byte for byte — the RGB, and every scale bucket and symbol of the latent stream in coding order (512 symbols over all 16 buckets of the scale table, 47 of the symbols coded by the rANS escape) — under the reference kernels and under the fast kernels on 1, 2, 3, 5 and all available threads; one integer unit changed in each of four chosen parameters, one on each path that feeds a value (a hyper-decoder weight, an entropy-network weight, a mean-head bias, a synthesis weight), changes the image; a header whose shapes the model cannot produce, or on which a network would form a buffer over the cap, is refused before anything is allocated (tests/int_mlicpp_preflight.rs, which counts allocations); damaged payloads are refused or decode, and never panic in an overflow-checked build | measured 2026-09-29 with rustc 1.98.0 on the source of the change that added this row: macOS arm64 native (the wai-rs tests, with overflow checks on and off; the byte-exact-conformance golden, release) and wasm32-wasip2 under wasmtime 49.0.1 (the same tests and golden, release; on wasm the fast kernels run single-threaded as portable scalar code, and usize is 32-bit). All pinned hashes matched on both. Not measured: x86_64 — the development machine has no x86_64 translation layer installed, and an x86_64-apple-darwin build of the golden compiles but cannot execute there — nor Linux or Windows. The golden runs on every leg of the byte-exact matrix (wai-conformance.yml, wai-byte-equality.yml) and tests/int_mlicpp.rs on every leg of wai-conformance.yml in both profiles. Both passed on the forge’s Linux x86_64, Linux arm64 and macOS arm64 runners for the landing chain’s head 7cfcecfa (run 745, 2026-09-30); there is no Windows leg. The dispatch through decode_envelope_strict (the prior pinning the one-file model bundle, the strict gate) runs in wai-neural-gates.yml. The weights are synthetic, so nothing here measures compression quality | wai-rs/tests/int_mlicpp.rs, wai-rs/tests/int_mlicpp_preflight.rs, byte-exact-conformance/src/int_mlicpp_golden.rs, wai-rs/tests/neural_native_decode.rs; tools/wai_int_mlicpp_export.py |
So at entropy-consistency, WAI’s JPEG AI symbol decode is demonstrated across
architectures (x86_64 against arm64) once the distribution indices are given; the
integer derivation of those indices is demonstrated against the reference on one
machine, by a test CI does not run. At decode-equivalence, cross-architecture
identity (x86_64 against arm64) is demonstrated for int_synth, integer video and
int_hyper on committed small-scale vectors; of these, only int_synth has also
been decoded on Windows on main. wai.audio.int_codec and wai.splat.int_codec
decode their committed vectors to the hashes a separate implementation computed, on
Linux x86_64, Linux arm64 and macOS arm64 (run 745): cross-architecture identity
(x86_64 against arm64) against one pinned hash per vector, on small vectors; neither
has been decoded on Windows. For full-scale int_hyper, identity rests on
construction, not on a cross-machine comparison of decoded output: CI checks its
decode against the export’s integer reference on one machine per run.
wai.neural.int_mlicpp’s committed vector decodes to the same pinned bytes on macOS
arm64 and on wasm32-wasip2 (a 32-bit target), measured on the development machine;
it also passed on the byte-exact matrix’s Linux x86_64, Linux arm64 and macOS arm64
legs in run 745.
For the JPEG AI exact decoder’s pinned reconstruction (§5), identity across machines rests on construction, and on one comparison: a single codestream on the four platforms in the table. The construction is integer arithmetic for the integer stages and, for the float stages, only IEEE-754 basic operations, which are correctly rounded:
- evaluated single-threaded in the fixed order of the source, each at its own precision (true of x86_64, AArch64 and wasm32 targets; not of x87, which the exact decoder refuses);
- with no fused or reassociated operations, which Rust does not emit unless the source asks for them;
- in the default floating-point environment, which the exact decoder probes before decoding (§5): round to nearest even, subnormals neither flushed nor read as zero;
- with no platform math library on the path. The gain scaler needs
exp, which math libraries are not required to round correctly. Rounding to an integer is exactly specified, but on some targets it compiles to a call: with rustc 1.98,f64::round,f32::roundandf64::round_ties_evenbecome calls to the symbolsround,roundfandrinton x86_64, and the first two do on wasm32, and the link decides what answers them (in the x86_64 Linux executable inspected, Rust’s compiler builtins did, not the C library). The decoder takes both fromwai-rs/src/float_repro.rs, built from basic operations, so their bits are fixed by the source. Clippy’sdisallowed_methods, denied in each module on the decode path, and a check of the linked decoder’s symbols (tests/jpegai_decode_symbols.rs) keep math-library calls off the path (see the table above).
That the exact decoder’s gain scaler equals the reference decoder’s rests on margin,
because the two compute it in different arithmetic. The scaler is
exp((g + β) · log_k / 2⁷) rounded to 2⁻¹⁰. The exact decoder evaluates it in f64
with float_repro::exp, which was measured within 1 ULP of the host math library
on the four platforms. The reference software evaluates it in float32: the
log-domain gain as float32, times log_k rounded to float32, then the float32 exp
of its tensor library, then rounding half to even. Neither exp is guaranteed to be
correctly rounded. The inputs are exactly the normative gain tables, because the exact
decoder refuses β ≠ 0.
tools/wai_jpegai_gain_exp_margin.py computes the exact value of both evaluations
with 60-digit decimal arithmetic, over the gain tables exported from the reference
software, and gates on the result at β = 0. It found the following for every entry
of all eight tables:
- the exact value of the decoder’s f64 evaluation lies at least 4 × 10¹⁰ f64 ULPs from a rounding boundary;
- the exact value of the reference’s evaluation, from its float32 argument, lies at least 77 float32 ULPs from one;
- both round to the same integer.
So any float32 exp within 77 ULPs gives the reference the exact decoder’s scaler.
The tool also fails unless the decoder’s log_k is the value the reference computes.
This does not extend to every β. The tool’s --beta-sweep option reports, without
gating, all 5118 values of g + β that the tables and the 12-bit header field can
produce. At four of them (1115, 1922, 2285 and 2289), the exact value of the
reference’s evaluation lies within one float32 ULP of a boundary. There the
reference’s own scaler depends on the last bit of its exp, so it can differ
between the platforms the reference runs on. At 2285 and 2289 the exact value lies
just above the boundary, a correctly rounded float32 exp lands exactly on it, and
rounding half to even then gives the integer below it. At every g + β, the exact
decoder’s formula gives the scaler that a correctly rounded float32 exp gives.
The reference’s own expression, evaluated on the development machine (macOS arm64),
also gave the exact decoder’s scaler at all 5118; its float32 exp there was at
most 0.85 ULP from exact on those inputs. A decoder that implements β would still
have to settle what the reference’s scaler is at those four inputs. This one does
not implement β.