Skip to main content

WAI Extension: Integer Payloads

Mirrored from the canonical text at commit 75a74f31 ().

Status: Draft. Pins the payload formats of the WAI-native integer codecs that SPEC.md §5 registers without one (Payloads of the WAI-native integer codecs): WIV1 for wai.video.int_motion and wai.video.int_hyper, WIH1 for wai.neural.int_hyper, WIA1 for wai.audio.int_codec and WIS1 for wai.splat.int_codec. Defines each layout, its decode, the checks a sink makes and the closed list of refusal codes; names each capability’s parameter-set members; and makes the JSON the encode and export tools write a converter input, never a payload. Reference impl: wai-rs — int_payload::{wire, decode_int, resolve}, int_payload::{wiv1_from_json, wih1_from_json, wia1_from_json, wis1_from_json}, codecs::PINNED_INT_PAYLOADS, container::check_int_payload_framing. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.

1. Scope

capabilitypayloadparameter set (§9)
wai.video.int_motionWIV1 with keyframe 0 (§3)none
wai.video.int_hyperWIV1 with keyframe 1, a nested WIH1 (§3)several files: entropy.json, g_s.whm, h_s.whm
wai.neural.int_hyperWIH1 (§4)several files: entropy.json, g_s.whm, h_s.whm
wai.audio.int_codecWIA1 (§5)one file: model.wam
wai.splat.int_codecWIS1 (§6)one file: g_s.whm

Why this pin is permitted. Until this revision, these capabilities’ SPEC §5 rows named the decoder that read their payload, and no byte layout. The rows of wai.video.int_motion, wai.audio.int_codec and wai.splat.int_codec also named the encode or export tool that wrote a bitstream, a JSON document; those of wai.neural.int_hyper and wai.video.int_hyper named their decoders only. SPEC §5 called every row’s payload the bytes its codec normally emits, which for these rows was the tools’ JSON. So this revision changes what these strings carry. SPEC §8 permits it under Completing a registration: a revision may complete the registration of a WAI-native string whose row named a decoder and fixed no byte layout by pinning one, and it lists the envelopes that become non-conforming (§11; SPEC §5). SPEC §8’s rule that a registered string’s payload format is fixed forever applies to these strings from this revision on. The encode and export tools still write JSON that carries the same fields; it is a converter input (§10).

wai.neural.int_synth and wai.neural.int_mlicpp keep the payload formats SPEC §5 gives them. A sink decodes them and the five strings above (wai.video.int_motion, wai.video.int_hyper, wai.neural.int_hyper, wai.audio.int_codec, wai.splat.int_codec) through one path (int_payload::decode_int), and §8 names the codes it reports for those two.

2. Common rules

Encoding. Every multi-byte field is little-endian. Offsets are in bytes.

Framing. A payload starts with its four-byte magic and a version byte. This revision defines version 1 for every format. Every reserved field is zero. Nothing follows the last section.

rANS streams. A stream is a u32 length L followed by L bytes. L is a multiple of 4 and at least 8. It is decoded as follows (precision 16, with a bypass escape of 4-bit groups):

One clip, several payloads. These rules fix how a payload is read, not the only payload of a decoded output: tables that differ in entries the stream never selects (a symbol the clip does not use, say), or a different escape layout, code the same output in other bytes. A payload’s hash therefore identifies the payload, not what it decodes to. Where one output must have one identity, it is identified by the digest of its decoded output: for a clip, its frame-state-equivalence digests (SPEC §7, a video receipt’s clip_hash); for the other formats, their capabilities’ equivalence digests.

CDF tables. A table is n entries of u32 with 3 ≤ n ≤ 4097, cdf[0] = 0, cdf[n − 1] = 65536 and every entry greater than the one before. These rules are what keep the symbol search and the state update above in range.

Order of the checks. A parser checks, in this order, and refuses with the first that fails:

  1. the magic: the capability’s (§1), or else format_capability_mismatch if it is another of the four magics, or bad_magic;
  2. the version byte: present (stream_length), and 1 (version);
  3. the fixed header: present (stream_length);
  4. the reserved fields (reserved_nonzero);
  5. the header’s fields, in the order the format’s section lists their rules (field_range), then, for WIV1, the keyframe kind (keyframe_capability_mismatch);
  6. the caps of §7 on the sizes the header declares (too_large), in the order the format’s section lists them;
  7. each later section in order: present (stream_length), then its rules (field_range, cdf_malformed). A section of one table per channel (WIA1, WIS1) is read table by table: each table present, then checked, before the next;
  8. nothing after the last section (trailing, or nested_length for WIV1’s nested keyframe).

A decode then makes the checks of §7, and refuses a stream that runs out or does not end canonically (§2). Two conforming sinks refuse the same bytes with the same code.

3. WIV1 — integer inter-frame video

offsizefieldrule
04magic WIV1
41version1
51mv_precision1 (whole-pel) or 2 (half-pel); 1 when n_frames is 1
61keyframe0 or 1
71reserved0
84n_frames≥ 1
124h≥ 1
164w≥ 1
204block≥ 1; 1 when n_frames is 1
244q1..=65535
284offseti32
322cdf_len3..=4097
342reserved0
36cdf_len·4the residual table§2
…(n_frames − 1)·nb·4motion vectorsi16 dy, i16 dx
…4 + Lthe residual stream§2
…4 + Kkeyframe 1 only: u32 K, then a WIH1 payloadK is exactly the bytes that remain

The header’s field rules are checked in the order mv_precision, keyframe, n_frames, h and w, block, q, cdf_len, then, for a one-frame clip, mv_precision 1 and block 1: such a clip codes no motion vector, so neither field would change its decode, and they are fixed so that it has one header. keyframe 0 is wai.video.int_motion’s and 1 is wai.video.int_hyper’s; the other is keyframe_capability_mismatch. The frame h·w·3, the clip n_frames·h·w·3, the frame count n_frames and the motion-vector field 2·nb·(n_frames − 1) are capped (§7), in that order.

Motion vectors. nb = ⌈h/block⌉·⌈w/block⌉ and nbx = ⌈w/block⌉. The vectors are frames 1 to n_frames − 1 in order, each its block grid row-major: the vector of pixel (y, x) in frame f is entry (f − 1)·nb + (y div block)·nbx + (x div block), in units of 1/mv_precision pel.

Decode. One decoder reads the residual stream for the whole clip, against the residual table at offset. For each frame in order, for each pixel in raster order (y, x), for each channel c of R, G, B: decode the symbol r, and the sample is clamp(pred + r·q, 0, 255). Since |r| < 2^31 and q < 2^16, every step fits i64.

The output is n_frames RGB8 frames of h × w (rows top to bottom, pixels left to right, R, G, B bytes): the buffers of the capability’s frame-state-equivalence digests (SPEC §7 Digests).

4. WIH1 — the integer conditional image codec

offsizefieldrule
04magic WIH1
41version1
53reserved0
84m≥ 1: latent channels
124sh≥ 1: latent height
164sw≥ 1: latent width
204n_z≥ 1: hyper-latent channels
244zh≥ 1: hyper-latent height
284zw≥ 1: hyper-latent width
324 + Lzthe hyper-latent stream§2
…4 + Lythe latent stream§2

The caps (§7) apply to n_z·zh·zw and to 2·m·sh·sw.

Decode, against the set of §9 (g_s.whm and h_s.whm, two integer networks at one fixed-point scale 2^afrac; entropy.json, the tables):

  1. For each hyper-latent channel k and each of its zh·zw elements, decode a symbol s from the hyper-latent stream against channel k’s table and offset; the element is s·2^afrac + median_k.
  2. h_s on the hyper-latent [n_z, zh, zw] gives [2·m, sh, sw]: its first m channels are the scales σ, the rest the means μ.
  3. For each latent element p in order (channel, then row, then column), its table is b, where, with the scale edges e_0 … e_{N−1} and σ' = max(σ_p, e_0), b = N − 1 − #{ i < N − 1 : σ' ≤ e_i }. Decode a symbol r from the latent stream against table b and its offset; the element is r·2^afrac + μ_p.
  4. g_s on the latent [m, sh, sw] gives three channels; each value v is clamped to [0, 2^afrac] and becomes (v·255 + 2^(afrac−1)) >> afrac, as RGB8.

Every value is formed in i64 and must fit it (§7).

5. WIA1 — the integer learned audio codec

offsizefieldrule
04magic WIA1
41version1
51afrac1..=30; equal to the model’s (§7)
62reserved0
84sr≥ 1: the sample rate the PCM is presented at, in hertz
124n_ch≥ 1: latent channels
164lat_len≥ 1: latent length
204t_out≥ 1: samples presented
244offseti32: the symbol offset of every channel’s table
282cdf_len3..=4097
302reserved0
32n_ch·8mediansi64 each, `
…n_ch·cdf_len·4one table per channel§2
…4 + Lthe latent stream§2

The header’s field rules are checked in the order afrac, then sr, n_ch, lat_len and t_out, then cdf_len. The caps (§7) apply to n_ch·lat_len and to t_out.

Decode, against the set of §9 (model.wam, the integer synthesis):

  1. For each channel k and each of its lat_len elements, decode a symbol s against channel k’s table at offset; the element is s·2^afrac + median_k.
  2. The synthesis (one-dimensional integer transposed convolutions and integer inverse GDN) on [n_ch, lat_len] gives one channel of t ≥ t_out samples.
  3. The first t_out samples v become (v·32767 + 2^(afrac−1)) >> afrac, clamped to [−32768, 32767].

The output is t_out samples of i16 PCM, one channel, at sr hertz: the buffer of the capability’s sample-equivalence digest (signed 16-bit little-endian samples).

6. WIS1 — the integer learned Gaussian-splat attribute codec

offsizefieldrule
04magic WIS1
41version1
51afrac1..=30; equal to the model’s (§7)
62reserved0
84c≥ 1: latent channels
124sh≥ 1: latent height
164sw≥ 1: latent width
204s≥ 1: the attribute planes’ side
244attr≥ 1: attribute planes
284n_gaussians1..=s·s
32attr·8rangesper plane, f32 lo then f32 hi (IEEE 754 binary32): finite, lo ≤ hi
…c·2table lengthsu16 each, 3..=4097
…c·4symbol offsetsi32 each
…c·8mediansi64 each, `
…the lengths’ sum ·4one table per channel§2
…4 + Lthe latent stream§2

The header’s field rules are checked in the order afrac, the dimensions, n_gaussians. The caps (§7) apply to c·sh·sw and to attr·s·s.

Decode, against the set of §9 (g_s.whm, the integer synthesis):

  1. For each channel k and each of its sh·sw elements, decode a symbol s against channel k’s table and offset; the element is s·2^afrac + median_k.
  2. g_s on [c, sh, sw] gives [attr, s, s] (geometry_mismatch otherwise).
  3. Each value v is clamped to [0, 2^afrac] and becomes (v·255 + 2^(afrac−1)) >> afrac, a byte.

The output, plane after plane, is the buffer of the capability’s splat-equivalence digest. The first n_gaussians cells of the planes, in raster order, hold one Gaussian each. The ranges are presentation data: a sink presents a plane’s byte u at about lo + (hi − lo)·u/255. They are outside the digest, and this extension does not fix presentation.

7. Caps and the checks a decode makes

Caps, the same on every target. No tensor a payload sizes may exceed 2^26 elements: a frame h·w·3, a WIV1 motion-vector field 2·nb·(n_frames − 1), a latent, a hyper-latent, a network’s output, an attribute tensor, a clip’s t_out. A WIV1 clip may not exceed 2^30 bytes, nor 2^20 frames.

Planning the networks. Before a decode allocates a latent, it plans every network on the payload’s shapes from the layers’ geometry and parameter lengths alone. The network’s input must be within the cap (too_large); then, layer by layer in order:

  1. the layer fits the running channel count (model_mismatch): a convolution’s output channels, kernel and stride are non-zero, it has c·out·k·k weights (c·out·k for the audio synthesis) and out multipliers and biases, and its requantisation shift is under 128; an inverse GDN has c·c and c parameters and shifts under 128; a leaky ReLU’s shift is under 64;
  2. every size the layer forms is within 2^26 (too_large): a convolution’s padded input side n + 2·pad, a transposed convolution’s uncropped side (n − 1)·stride + k + opad and uncropped plane, and the output;
  3. the layer leaves an output (model_mismatch): a kernel no larger than the padded input, a crop 2·pad smaller than the uncropped side.

Every size is formed in 64-bit arithmetic with overflow checked, and a size past 64 bits is past the cap, so a 32-bit sink refuses exactly what a 64-bit one does, with the same code; once a network is planned, every size it forms fits a 32-bit index.

Memory and work. A payload’s header declares its decode’s output and the work that output takes, and a sink can read both before it decodes anything. A stream does not bound them: a table can make one symbol cost a small fraction of a bit, so a few kilobytes of stream can code a clip at the 2^30-byte cap, as any codec can describe a large output in a small file. The bound is the header’s, within the caps:

A WIA1’s latent is the one its t_out needs and no longer: the synthesis of lat_len samples gives at least t_out samples and that of lat_len − 1 fewer (geometry_mismatch), so a payload cannot declare a large latent to present a few samples.

A decode refuses a stream that runs out at its first missing symbol (§2), so a short stream stops early, but a stream that carries its symbols is decoded in full: a sink that will not spend what a header declares refuses before decoding, by its own size policy, applied to the declared cost (int_payload::declared_cost and int_payload::check_cost, with the set; neural::declared_cost over the set a registry resolved; the C ABI’s wai_int_payload_declared, given the set’s path; wai-web’s intPayloadCost; the header alone through container::declared_size), cost_over_limit (§8). The reference neural sink applies one when it is given one (neural::ModelRegistry::with_cost_limit); its C ABI does not decode integer video, PCM or splat attributes, which it has no output kind for. This extension fixes the caps; a sink’s policy may be stricter.

The parameter set must fit the payload (model_mismatch):

The shapes the payload declares must be the ones the decode gives (geometry_mismatch): h_s’s output and [2·m, sh, sw]; a WIV1 keyframe and h × w; the audio synthesis’s length and t_out (it may be longer, by less than one latent sample’s worth: the synthesis of lat_len − 1 samples gives fewer than t_out); g_s’s output and [attr, s, s].

The order of a decode’s checks. After the parse (§2), a decode checks in this order and refuses with the first that fails:

  1. the set’s members, in the order of §9’s list, each present (set_incomplete) and parsed (model_mismatch);
  2. the parameter set against the payload (model_mismatch): one fixed-point scale, the payload’s afrac, the entropy tables, the hyper-latent channels;
  3. every network planned, in decode order (h_s then g_s; the audio synthesis; the splat g_s), each layer by the three rules above (model_mismatch, too_large): the caps come before any shape is compared;
  4. what the networks give against the payload: for WIH1, h_s’s shape (geometry_mismatch), then g_s’s three channels (model_mismatch); for WIV1 with keyframe 1, its keyframe’s as for WIH1, then the keyframe’s size against h × w (geometry_mismatch); for WIA1, one channel (model_mismatch), then at least t_out samples, then a latent no longer than that needs (both geometry_mismatch); for WIS1, g_s’s shape (geometry_mismatch);
  5. each stream in decode order (WIH1: the hyper-latent’s, then, after h_s runs, the latent’s; WIV1 with keyframe 1: the keyframe’s streams, then the residual stream), symbol by symbol: a symbol that runs its stream out (stream_exhausted) before its value’s range (field_range); after its last symbol, its canonical end (stream_noncanonical);
  6. each network as it runs, layer by layer (field_range, below).

So a run-out is reported before the value it would have given, and a size past the caps before a shape that disagrees. A sink that applies a cost limit (Memory and work) compares the declared cost with it after step 4, once everything the cost reads has been checked, and before any symbol is decoded (cost_over_limit, §8).

Exactness. A latent rebuilt as s·2^afrac + base must fit i64; before each network layer runs, the decode checks on the values present that no product, partial sum, requantised value or inverse-GDN term can leave the type its arithmetic is formed in. A stream whose values fail either check is refused (field_range). So a decode with overflow checks and one without compute the same integers, and neither panics.

8. Refusal codes

The list is closed: a sink reports exactly one of these, the first that applies in the order of §2 and §7.

A refusal of the payload or of the parameter set against it, bad_magic through model_mismatch in the table below (and payload_malformed), is final: the sink does not decode the payload by another path, and SPEC §4’s fallback is not tried, since it would decode the same bytes. The other codes say the sink cannot decode through the capability at all, and route as SPEC §4 routes such a capability:

codewhen
bad_magicthe payload does not start with any of the four magics: a tool’s JSON form lands here
format_capability_mismatchthe payload is another of the four formats than the capability’s
keyframe_capability_mismatcha WIV1 keyframe kind that is not the capability’s
versiona version this revision does not define
reserved_nonzeroa reserved field that is not zero
field_rangea field of any section outside its range (a header field, a WIS1 range or table length, a median); or a value the stream decodes that leaves the exact integer range (§7)
too_largea size past the caps of §7
cdf_malformeda table that breaks the rules of §2
stream_lengththe payload ends before a field or section it declares; or a stream length that is not a multiple of 4, or is under 8
nested_lengtha WIV1 keyframe length K that is not the rest of the payload
trailingbytes after the last section
stream_exhausteda stream that ran out before its last symbol
stream_noncanonicala stream that does not end canonically after its last symbol: a word left unread, or a state other than 2^31 (§2)
geometry_mismatcha shape the payload declares that the decode does not give (§7)
model_mismatcha parameter set that does not fit the payload (§7)
set_incompletea parameter set without a member the decode reads (§9); a pinned decode reports SPEC §3.1’s set_incomplete
not_dispatchablea capability SPEC §4 does not dispatch
unsupported_capabilitya capability the sink has no integer decoder for
payload_malformeda wai.neural.int_synth or wai.neural.int_mlicpp payload its decoder refuses
cost_over_limita payload whose declared cost (§7) exceeds the limit the sink applies before decoding

9. Parameter sets

A capability that decodes against a parameter set reads these members, by the names a set’s listing gives them (prior-carriage §7):

capabilityshapemembers
wai.neural.int_hyper, wai.video.int_hyperseveral filesg_s.whm and h_s.whm (networks in WHM1), entropy.json (the tables)
wai.audio.int_codecone filemodel.wam (the synthesis in WAM1)
wai.splat.int_codecone fileg_s.whm (the synthesis in WHM1)
wai.video.int_motionnone—

A set of several files is pinned through its listing (wai.prior.set); a one-file set MAY be pinned form-less, by its file’s SHA-256 (SPEC §3.1). wai.audio.int_codec and wai.splat.int_codec were registered with sets of several files, because no one-file format had been named for them; this revision registers each as one file, the file its decoder reads, which holds every parameter its output depends on. That widens their registrations, as SPEC §8 allows a MINOR revision to; a set pinned through its listing still verifies.

WHM1 (int_hyper_synth::IntModel::from_bytes): magic, then afrac, gf, shift and the layer count (u32 each), then per layer a tag byte: 0 a transposed convolution and 1 a convolution (input and output channels, kernel, stride, padding and, for 0, output padding as u32; weights as i16; per-output-channel multipliers and biases as i64), 2 an inverse GDN (channels as u32; gamma as i32, beta as i64), 3 a ReLU, 4 a leaky ReLU (slope i64, its shift u32). WAM1 (int_audio::AudioModel::from_bytes): the same header, then layers of tag 0, a one-dimensional transposed convolution, and tag 2, an inverse GDN. In both forms nothing follows the last layer: a file with bytes after it is refused (model_mismatch), so one network has one file and one digest. The layers’ integer arithmetic is the reference decoders’ (int_hyper_synth.rs, int_audio.rs, int_transform.rs), which the conformance vectors pin.

10. The tools’ JSON forms

The encode and export tools under tools/ write a JSON bitstream. It is a converter input and a debugging aid, never a payload. An encoder MUST NOT carry it as the payload of any capability of §1, and a sink refuses it (bad_magic). The converters (int_payload::{wiv1,wih1,wia1,wis1}_from_json, and the independent int-payload-conformance/int_payload_verify.py) make explicit what the JSON leaves implicit:

A payload and the JSON it is converted from decode to the same bytes; the conformance vectors check that for every committed JSON bitstream.

11. Envelopes this revision makes non-conforming

This revision completes these strings’ registrations with a payload format they did not have (SPEC §8, Completing a registration). It makes non-conforming, and they still parse:

A sink refuses such a payload when it decodes it (§8), with no fallback attempt. An encoder MUST NOT write one; the reference implementation’s emit paths (wai wrap, the C packers and the browser writers) parse the payload as a sink does and refuse to. An encrypted envelope’s payload (SPEC §6) is ciphertext until a sink decrypts it, so the emit paths do not check it; the sink checks the plaintext when it decodes it (SPEC §6.4). No committed envelope or conformance corpus carried such a payload when this revision was made.

12. Reference implementation and conformance