WAI Extension: Staged-Delivery Measurement
Mirrored from the canonical text at commit 0c6a8897 ().
Status: Draft. How a figure about staged delivery (staged-delivery) is measured and published: which bytes are counted, how quality is computed per prefix, what “first presentable” means, how layering overhead is stated, how decode time and energy are reported, how deduplication is computed, how delivery is simulated, and what a record carries. Reference impl:
wai-rs—measure_framing,quality,output_tree,content_dedup,bench_env,stage_measure::{title, decode, run, sim}; thewai_stage_measuretool; the corpusstaged-measure-conformance/. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.
1. Scope
§1.1 This extension binds any party that publishes a figure about WAI staged or split delivery: bytes, time, quality, overhead, decode time, energy or deduplication. It binds the reference tools. It does not bind a sink’s decode behaviour, which staged-delivery states.
§1.2 A figure that does not meet this extension MUST NOT be presented as a WAI measurement.
§1.3 This extension states no price, and a record MUST NOT carry one. A deployer that derives a cost from transferred bytes does so outside WAI.
2. Terms
| term | meaning |
|---|---|
| title | What a session plays: one or more bases and the layers that refine them (staged-delivery §1). |
| unit | A presentable output, as staged-delivery §2 fixes it per capability: a frame, a picture, an audio clip, an attribute tensor. A title numbers its units globally: a base’s unit i is the title’s unit first_unit + i. |
| stage | A unit’s stage: 0 after its base decodes, k after a layer whose refines.layer is k applies to it. |
| prefix | The stages 0…k of a title: every object whose layer is at most k. Its layers count is k + 1. |
| presentation | A (unit, layers) pair a sink handed to output, with the output it presented. |
| counted bytes | Bytes under §3. Apart from transferred bytes, the only byte figure this extension reports. |
| transferred bytes | Bytes of WAI objects as a sink received them, whole or as byte ranges, without transport or receipts. |
A stage’s position is its layer (0 for the base). A prototype MoQ binding in this
repository carries it as mrl_layer, a stage set’s stage count as mrl_total_layers,
and a subscriber’s upper bound on the layers it accepts as max_mrl_layer; this
extension names that bound layer_cap. Staged-delivery §13 says which of them a sink
reads.
3. Byte accounting
§3.1 Counted bytes of an object. For a WAI1 envelope, 12 + man_len + P; for a
WAI2 envelope, 8 + man_len + 2 + 8·n + Σ Pᵢ. Here man_len is the manifest’s
length as written, n is the component count, and P or Pᵢ is each payload’s
length (SPEC §2, §7.1). These are the envelope’s bytes, exactly: an object with bytes
after its last payload is not the envelope its lengths describe, and is not counted.
Transport bytes (HTTP, QUIC, MoQ object headers, TLS) are not counted bytes; a figure
that includes them MUST state them separately.
§3.2 Only a checked format is counted. An object is countable only where WAI checks
the format of every payload it carries: a capability SPEC §5 pins to a binary format
(WIV1, WIH1, WIA1, WIS1, and WIR1 for a refinement-only capability; Appendix
A) whose payload parses as that format, by the parse a sink runs. Where a WAI1
manifest names a fallback with a pinned format, the payload must parse as the
fallback’s format too, since the fallback decodes the same payload (SPEC §4 step 5).
Otherwise the object is not countable, with a code:
no-format-check: a capability WAI checks no format for, a classical codestream or a string WAI does not register among them. Nothing in WAI tells such a payload’s bytes from a text encoding of them (a JSON text, base64, JSON arrays of byte values), which MUST NOT be counted, as stored or as an estimate of what it would be in binary;- the format parse’s code (
bad_magic,format_capability_mismatch, …: integer-payloads §8, staged-delivery A.6), for a payload that is not its capability’s or its fallback’s format. A JSON form of a pinned format, its base64 text, a JSON string, or the format’s bytes behind a byte-order mark all fail here; sha256-mismatch: aWAI2component whosesha256is not its payload’s (SPEC §7.1);encrypted: an encrypted payload, whose plaintext’s format cannot be checked.
No byte figure is published for an object that is not countable. The JSON that the integer codecs’ encode and export tools write is a converter input (integer-payloads §10); its counted bytes are those of the envelope that carries the converted payload.
§3.3 Framing label. A byte figure carries a framing label. A figure is
registered when every payload it counts is in its capability’s registered format,
which §3.2 requires of every figure this revision defines. A later revision that
counts a payload under another framing names its label. Two figures with different
framing labels MUST NOT be compared, summed or placed on one curve.
§3.4 A record MUST state:
- the counted bytes of each object, split into envelope header, manifest and payload;
- the sums per (base, stage) and per (base, prefix), and over the title per stage and per prefix;
- the framing label;
- the bytes of each base’s parameter set, separately (staged-delivery §11).
§3.5 Receipt and claim bytes are counted separately, as receipt bytes.
4. Quality per prefix
§4.1 Quality is stated per presentation and per prefix against a reference: the source, where the title pins it by SHA-256, and the final prefix’s output, labelled “against the final prefix”. A figure against the final prefix MUST NOT be presented as distortion against the source. A unit whose shape at a stage differs from the reference’s (below an up2 or a rate layer) has no figure at that stage; a figure pooled over a prefix covers the units that have one, and says how many.
§4.2 Pictures and frames (rgb8). For two RGB8 buffers of width W ≥ 1 and
height H ≥ 1:
sse_rgb= Σ (aᵢ − bᵢ)² over the3·W·Hbytes;sse_yis the same sum over the luma planes, withY = (77·R + 150·G + 29·B + 128) >> 8;ssim_y_q32is the mean structural similarity of the luma planes in Q32 fixed point (2³²is 1).
The similarity’s windows measure ww = min(8, W) by wh = min(8, H) pixels. Their
left edges are at every multiple of 4 up to W − ww, and at W − ww itself when it is
not a multiple of 4; their top edges are at every multiple of 4 up to H − wh, and at
H − wh itself when it is not one. A window is at every pair of a left and a top edge,
so every pixel is in at least one window. For a window of n = ww·wh pixels, with Sx = Σx,
Sy = Σy, Sxx = Σx², Syy = Σy² and Sxy = Σxy over the reference’s luma x and
the test’s luma y:
A1 = 20000·Sx·Sy + 65025·n²
A2 = 20000·(n·Sxy − Sx·Sy) + 585225·n²
B1 = 10000·(Sx² + Sy²) + 65025·n²
B2 = 10000·((n·Sxx − Sx²) + (n·Syy − Sy²)) + 585225·n²
s_w = ⌊ A1·A2·2³² / (B1·B2) ⌋ (floor toward −∞)
ssim_y_q32 = ⌊ (Σ_w s_w) / (window count) ⌋ (floor toward −∞)
This is the structural similarity of Wang et al. (2004) with K1 = 0.01, K2 = 0.03, a
peak of 255, uniform windows at stride 4 with an edge-aligned last window on each
axis, and population statistics, multiplied out so
that no floating-point operation occurs. Because n ≤ 64, |A1·A2·2³²| < 2¹¹⁵ and
B1·B2 > 0 (the reference implementation’s quality module proves both), so the
computation fits signed 128-bit integers.
§4.3 Audio (pcm16). For two sequences of i16 samples, channel-interleaved and
in output order: sse = Σ (aᵢ − bᵢ)² and ref_energy = Σ aᵢ² over the reference.
Both are unsigned integers of at least 96 bits, written in records as decimal
strings. Sequences of different lengths are refused.
§4.4 Tensors (tensor-i64), for splat attributes, volumetric attributes and
haptic samples: sse over the capability’s canonical integer output, written as a
decimal string, and max_abs = max |aᵢ − bᵢ|.
§4.5 Pooling. A figure over several units (a clip, or a prefix) sums each sum over
the units; for pictures, ssim_y_q32 is the mean of every window of every unit, one
division over the pooled Σ s_w and window count, and max_abs is the largest.
§4.6 Exact outputs (state-exact), for worlds, feeds, scores, avatars, quantum
and every capability whose output is a state: a prefix is equal or differs, by
SPEC §7 Digests. It is never stated as a distance.
§4.7 PSNR = 10·log10(255²·N / sse) and SNR = 10·log10(ref_energy / sse) are derived
for display only. A record MUST carry the integers. An sse of 0 has no finite
PSNR, and is shown as exact.
§4.8 A record MAY add another full-reference metric. It names the metric’s implementation and version, and the metric is informative only.
5. First presentable
§5.1 bfpf_min is the sum of the counted bytes of the smallest set of objects needed
to present unit 0 at one layer: the base that holds it, and every index or manifest
object a sink reads first. It is computed from the title alone. A base’s parameter
set is stated beside it (§3.4), not in it.
§5.2 bfpf_sched is the counted bytes that had arrived when unit 0 was first
presented, under a stated delivery schedule (§10). It is at least bfpf_min.
§5.3 tfpf is the time from the session’s first request to the first presentation of
unit 0, in microseconds. Under §10 it is virtual and deterministic. Measured, it is a distribution over repeated runs, with the
clock and the presentation event named; in a browser that event is the frame callback
after the canvas write.
§5.4 tffs_k is the time to the first presentation of unit 0 at k + 1 layers, stated
as §5.3 states tfpf.
6. Layering overhead
§6.1 A figure that states a saving from staging MUST be accompanied by its layering overhead.
§6.2 The single-layer reference is the same content coded by the same staged form as one layer at the final layer’s operating point: the title’s bases and, over each, one layer straight from stage 0 to the source, with the final layer’s step and the predictor that reaches the final shape. A reference at further operating points is the same with another step.
§6.3 The overhead is two rate–quality points: (counted bytes of the full title, quality of its final prefix against the source) and (counted bytes of the reference, its quality against the source). A record MUST NOT reduce them to a ratio, and MUST NOT say “at equal quality” unless the two quality integers are equal. It SHOULD add the reference at further operating points as a curve, and state which two reference points bracket the quality of the final prefix.
7. Decode time
§7.1 Decode time is measured per (unit, stage) on a stated machine and build. A record states the runs; the minimum, median, 10th and 90th percentiles; the order in which cases ran; and the discarded warm-up runs.
§7.2 Cases MUST be interleaved, with the order rotated each round.
§7.3 A run whose output digest differs from the reference is not a timing; the bench aborts.
§7.4 A record states the machine (its processor, cores, operating system and architecture, and the build profile), its load before and after each block, its power source, and whether it was otherwise idle. Continuous integration MUST NOT assert any timing.
8. Energy
§8.1 An energy figure MUST follow energy-measurement:
an acquisition class (§2); for a figure obtained under its §3, the bracket a
measurement claim publishes (§6); the build (§5); and a recomputable work unit. For
staged decode, the work unit is the (unit, stage) decodes of one completion.
§8.2 A figure the meter could not obtain is unmetered: it is written as null
with a reason, never as 0.
§8.3 Only a figure whose attribution verdict (energy-measurement §4, §6.2) is clean is published as a figure. A bracket usable for relative comparison MAY be shown, labelled “relative comparison only, not a figure”, and never as an absolute joule figure. Noisy and unusable brackets are kept in the record and not shown.
§8.4 A figure covers decode on the measured machine’s processor rails. A record MUST NOT state, imply or chart energy for displays, radios, networks or other devices. A reading from a whole-system rail that includes a display is not published.
§8.5 A rail that is not live in a bracket (energy-measurement §3.1) gives no figure; the decode is unmetered.
9. Deduplication
§9.1 Over a set of objects:
objectsis their count, a repeated object counted each time;uniqueis the count of distinct content hashes (jwp-receipts §2.1);referenced_bytesis the sum of their counted bytes;stored_bytesis the sum over one object per content hash.
All four are recomputed, never declared. One content hash has one length; a set that gives one hash two lengths has no totals.
§9.2 Deduplication is stated per title, over every object the title and its single-layer references reference (a reference’s bases counted again), and, where content repeats across titles, across the record.
10. Delivery simulation
§10.1 A deterministic delivery figure comes from the staged delivery simulator, a discrete-event model in integer microseconds defined by its model and plan (Appendix B). Its figures are labelled “simulated” with the plan. A simulated figure is never a measurement of a network, and is never presented as one.
§10.2 Delivery timeout. A plan gives each object a carriage, after the
DELIVERY TIMEOUT of MoQ Transport (draft-ietf-moq-transport):
datagram: each object is one datagram, and one not started by its deadline (its availability plus its timeout) is dropped.subgroup: each (base, layer) is one subgroup, whose objects are sent in order. When an object of a subgroup is not fully sent by its deadline, the subgroup is reset: the bytes already sent of it and of every later object of the subgroup are abandoned, and the bytes not yet sent are dropped, whatever the later objects’ own deadlines.
Dropped and abandoned bytes are never delivered, and every offered byte is delivered, dropped or abandoned.
§10.3 Layer cap. An object whose layer exceeds layer_cap is never sent and never
counted as offered.
§10.4 A refinement never stalls its base beyond a bound. A plan’s base-only
model removes every refinement. The stall of a unit is its base presentation time
under the plan minus the same time under the base-only model. The stall bound is
S = S_link + S_decode:
S_link: understrict-priority, the serialisation time of one chunk; underfifowith a refinement timeoutT,Tplus one chunk; otherwise there is no bound.S_decode: underbase-firstdecode, one refinement’s decode time; otherwise there is no bound.
Preconditions. The bound holds for every model and every plan with no base
timeout (delivery_timeout_us.base is null), a bounded link schedule (strict-priority,
or fifo with a refinement timeout) and base-first decode, whatever else the plan
states (carriage, window, live captures, layer cap, workers, decode times, playout
deadlines). A plan outside them has no bound: a base its refinements delay past its own
timeout is not presented at all, and the two unbounded parts above have counter-plans.
Proof. Write c for one chunk’s serialisation time, ⌈chunk_bytes·1000 / bytes_per_ms⌉, the longest any chunk takes; L = c under strict-priority and
L = T + c under fifo; and d_r for one refinement’s decode time. A base below is
one layer-0 object (a base split into several layer-0 objects contributes each).
Number the bases b_1, b_2, … in the order the sender sends them. That order is
greedy under subgroup precedence: whenever the sender starts a base, it takes, of the
bases available then and not held behind an unfinished earlier object of their
subgroup (§10.2; under subgroup, a base’s layer-0 objects form one subgroup, sent in
title order), the least by the order key (Appendix B, Order), whose first part is the
availability under either schedule. It is the same order under the plan and under its
base-only model. By induction on k: when either starts its k-th base, the bases
finished are b_1, …, b_{k−1} in both, so the same bases are held. The base-only
model starts it at the first time after b_{k−1} ends at which a base is available;
the plan starts it no earlier, since b_{k−1} ends no earlier under the plan and it
too needs a base available. A base that becomes available after the base-only model’s
start has a later availability, and so a larger key, than every base eligible then, so
the least is b_k in both. The order can differ from the order of the keys alone:
a base listed after another of its subgroup waits for it, even when it is available
earlier. A base’s service time s_k (the sum of its chunks’ times) is the same in
both.
The link. Let a_k be b_k’s availability, start_k, end_k its first chunk’s start
and last chunk’s end under the plan, and start⁰_k, end⁰_k under the base-only model.
- A base, once started, is sent to its end without a chunk of another object between.
At each boundary while
b_kis sent, the eligible objects are those eligible when it started, less any dropped or reset, and those that became available since, whose availability, and so whose key under either schedule, is larger thanb_k’s. None becomes eligible by its subgroup in between: an earlier object of a subgroup is finished only by being sent, which does not happen whileb_kis sent, and a reset ends the whole tail.b_kwas the least eligible object when it started (understrict-priorityevery base precedes every refinement; underfifoit was chosen), so it stays the least. Soend_k = start_k + s_k. - In the base-only model,
start⁰_k = max(a_k, end⁰_{k−1}): the bases are sent in order, and the link is idle only when no base is available. - Under the plan, at
t_k = max(a_k, end_{k−1})baseb_kis eligible (the bases its subgroup holds it behind are amongb_1, …, b_{k−1}) and, by the order above, the eligible base of least key. Understrict-priority, a base precedes every refinement, sob_kstarts at the first boundary at or aftert_k; a chunk in flight att_kis a refinement’s (step 1), and ends withinc. Sostart_k ≤ t_k + c. Underfifo, objects that precedeb_kare bases (all before it in the order) and refinements available no later thana_k, whose deadlines are therefore at mosta_k + T. At the first boundary aftera_k + T, which comes withincof it, every such refinement not yet finished has been dropped or reset (§10.2); one already finished took no longer. Sostart_k ≤ max(end_{k−1}, a_k + T + c), and in both casesstart_k ≤ max(a_k + L, end_{k−1}), withend_{k−1}exact when it is the larger: a base that ends at a boundary whereb_kis eligible is followed byb_k. - By induction on
k,start_k ≤ start⁰_k + L: with step 3, step 1 and the hypothesis,start_k ≤ max(a_k + L, end⁰_{k−1} + L) = start⁰_k + L. Each base therefore arrives at mostLlater than in the base-only model, and the bases arrive in the same order in both.
The decode. Number the base decode jobs in arrival order, which is the order of step
4. Under base-first the sink takes bases in the order they became ready (Appendix B,
Workers), which is that order, and before any refinement. Let r_k ≤ r⁰_k + L be
b_k’s arrival under the plan and the base-only model, W the workers and d_b a
base’s decode time.
- In the base-only model, with equal decode times and jobs started in order,
u⁰_k = max(r⁰_k, u⁰_{k−1}, u⁰_{k−W} + d_b)isb_k’s start (terms with an index below 1 omitted): at that timeb_{k−1}has started, every base up tob_{k−W}has finished, so at mostW − 1workers are busy. - Under the plan, no refinement starts while a base is ready and waiting. A
refinement that holds a worker after
r_k + d_rtherefore started afterr_kand so afterb_kstarted. Atτ = max(r_k + d_r, u_{k−1}, u_{k−W} + d_b), ifb_khas not started, it is the waiting base of least order, no refinement holds a worker, and at mostW − 1bases do; sou_k ≤ τ. - By induction,
u_k ≤ u⁰_k + L + d_r: each term ofτis at most the matching term of step 5 plusL + d_r.
So every base is presented at most S_link + S_decode = L + d_r after its time in the
base-only model, and a late base is presented all the same (Appendix B,
Deadlines). ∎
The reference implementation checks the bound on generated models and plans drawn
over the whole range of the preconditions (groups and captures in any order, objects
listed in any order, slow links, zero and short timeouts, caps, one to four workers),
and on draws aimed at the orders the bound depends on (captures on and beside chunk
boundaries and out of order, bases split into two layer-0 objects, chunks of one byte),
which fail on their own when the send or the decode order of bases is not the one
above; the corpus checks it on every bounded plan it holds. For each part of the bound a
plan can lose, the corpus holds a counter-plan whose stall passes the bound of the
same plan with that part restored: a fifo link without a refinement timeout against
the same fifo link with one, and fifo decode against the same plan with
base-first decode.
11. Records
§11.1 A published figure MUST cite a dated record in
wai/research/staged-measurements/ by path and commit. The record carries:
- the date and the commit;
- the title’s content identity (Appendix D) and the framing label;
- the machine and build;
- every deterministic figure of §3–§6 and §9, and every measured figure of §5.3–§5.4, §7 and §8 with its method;
- the link hashes of its claims.
§11.2 Records are immutable. A correction is a new record that names the one it corrects.
§11.3 A classical baseline is a reference curve of points (counted bytes, quality per §4) from a named encoder at stated versions and settings, on the same source. A record MUST NOT state a multiplier, percentage or ratio between a WAI figure and a baseline.
12. What continuous integration asserts
§12.1 Only values that are the same on every machine:
- counted bytes,
bfpf_minand simulator outputs; - the integers of §4, output digests and the output tree’s root and leaf count;
- deduplication totals and the overhead points;
- every figure a conformance corpus states.
§12.2 Measured times and energy are stored as artifacts and never asserted.
§12.3 The corpus staged-measure-conformance/ holds
titles (Appendix D) and, for each, every figure of §12.1 (goldens.json) and the
fixture tool’s own reconstruction digests (expect.json). A conforming implementation
reproduces every figure and every digest. The corpus’s README names its second,
independent implementation.
13. Media-class matrix
The last column follows §3.2: an object is countable only under a capability whose
format WAI checks (count_object gives no-format-check for any other), so a class
whose staged form or base has no checked format is not countable in this revision,
whatever else it supports.
| class | unit | staged form | edge operation | quality kind | in this revision |
|---|---|---|---|---|---|
| image | a picture | pinned integer base (wai.neural.int_hyper) + WIR1 layers (wai.image.int_refine: identity, up2) | serve a prefix of layers | rgb8 | measured (corpus); a classical base (PNG and the like) is not countable in this revision |
| video, WAI-native | a frame | wai.video.int_motion or wai.video.int_hyper base + WIR1 layers, per frame or per span (wai.video.int_refine) | drop layers above a cap | rgb8 per frame, pooled per clip | measured (corpus) |
| audio | the clip | wai.audio.int_codec base + WIR1 layers (wai.audio.int_refine: rate, identity) | drop layers above a cap | pcm16 | measured (corpus); a wai.audio.flac base is not countable in this revision |
| splat | the attribute tensor | wai.splat.int_codec base + WIR1 layers (wai.splat.int_refine) | drop layers above a cap | tensor-i64 | measured (corpus) |
| video, enhancement layer | a frame | registered base + wai.video.lcevc companion (SPEC §7.1) | drop the companion | rgb8 after the sink’s decode | not countable in this revision (the companion’s format is not checked); quality needs a sink-supplied decoder |
| image, progressive passes | a picture | byte prefixes of one progressive codestream | serve a byte prefix | rgb8 after the sink’s decode | a staged form not yet registered (staged-delivery §6.2); not countable in this revision |
| audio, codebook prefix | the clip | the further codebooks of a residual-vector-quantised code | drop codebooks above a cap | pcm16 | a staged form not yet registered; not countable in this revision |
| volumetric (4D splat) | a keyframe span | keyframe refinement | serve a prefix | tensor-i64 | a staged form not yet registered; not countable in this revision |
| world / feed / film / avatar | a segment | snapshot and operation deltas | snapshot and tail | state-exact | a staged form not yet registered; not countable in this revision |
| score | a section | none | none | state-exact (mixdown digest) | not countable in this revision |
| haptics | an envelope | keyframe refinement | serve a prefix | tensor-i64 | a staged form not yet registered; not countable in this revision |
| quantum | a circuit | none | none | state-exact | not countable in this revision |
| text and records | an object | none | none | none | not countable in this revision |
Appendix A. Counted formats
The formats §3.2 counts, and how the reference implementation
(measure_framing::count_object) checks each before it counts it. A payload that fails
its check is not countable, with the check’s code.
| capabilities | format | check |
|---|---|---|
wai.video.int_motion (keyframe 0), wai.video.int_hyper (keyframe 1) | WIV1 | integer-payloads §2–§3, in its order |
wai.neural.int_hyper | WIH1 | integer-payloads §2, §4 |
wai.audio.int_codec | WIA1 | integer-payloads §2, §5 |
wai.splat.int_codec | WIS1 | integer-payloads §2, §6 |
wai.image.int_refine, wai.video.int_refine, wai.audio.int_refine, wai.splat.int_refine | WIR1 | staged-delivery Appendix A.1–A.3, for the capability’s unit kind |
| every other capability | — | none: not countable (no-format-check) |
A WAI1 fallback with a pinned format is checked as its capability is. Parameter sets
(SPEC §3.1) are not objects and are never counted bytes; §3.4 states them separately.
The corpus’s counting.json holds an object for each outcome.
Appendix B. Delivery simulator wai-sim/1
A model is a title’s objects in title order, the object index, as delivery sees
them (wai-sim-model/1):
{ "format": "wai-sim-model/1", "unit_us": 40000,
"objects": [ { "id": "base.wai", "base": 0, "layer": 0, "units": [0, 5], "bytes": 2519, "capture_us": null }, … ] }
units is [from, count] over the title’s units and bytes the object’s counted
bytes (§3). A model is well formed when every object has at least one byte and one
unit, no two objects cover one unit at one layer, and every unit an object of layer
k ≥ 1 covers is covered at layer k − 1 by an object of the same base. A measured
title’s model is its objects with their counted bytes, and its bases’ capture times.
A plan (every member required, none other allowed):
{ "plan": "wai-sim/1", "mode": "vod" | "live",
"link": { "rtt_us": 40000, "bytes_per_ms": 2500, "chunk_bytes": 1200 },
"sender": { "schedule": "strict-priority" | "fifo", "carriage": "datagram" | "subgroup",
"layer_cap": null, "delivery_timeout_us": { "base": null, "refinement": 30000 },
"vod_window_us": null },
"sink": { "workers": 1, "decode_order": "base-first" | "fifo",
"decode_us": { "base": 4000, "refinement": 2000 },
"playout_start_us": null, "late_refinement": "discard" | "present" } }
bytes_per_ms, chunk_bytes, workers and both decode times are at least 1. Every
integer of a model or plan is at most 2⁵³ − 1. All arithmetic is in integer
microseconds, and h = ⌊rtt_us / 2⌋; a time or a byte sum past 64 bits is refused
(overflow), never wrapped.
The sender.
- Availability. The sink requests at 0 and the sender holds the request at
h. An object is available ath + max(0, from·unit_us − vod_window_us)on demand (hwith no window), and atmax(h, capture_us)live; a live plan over an object with no capture time is refused. - Deadline. An object’s availability plus its layer’s timeout (
basefor layer 0,refinementotherwise); none when that timeout isnull. - Cap. An object above
layer_capis capped: never sent, never offered. - Datagrams. Under
datagram, an object of more thanchunk_bytesthat is not capped is refused (datagram-too-large). - Link. The link carries one chunk at a time. At each boundary — time 0, and the
end of each chunk — the sender first applies the timeouts, then sends the next chunk:
- Timeouts. In object order, each object still pending whose deadline is before the
boundary’s time: under
datagramit is dropped; undersubgroupits subgroup is reset from it (§10.2). - Eligibility. An object is eligible when it is pending, available, and, under
subgroup, no earlier object of its subgroup is pending. - Order. Of the eligible objects, the least by (layer, then for a base its
availability,
from, object index) understrict-priority, and by (availability, object index) underfifo: under either, bases go in the order they became available. - Chunk.
c = min(chunk_bytes, bytes not yet sent), occupying the link⌈c·1000 / bytes_per_ms⌉µs from the boundary, and arrivinghafter it ends. An object’ssent_first_usis its first chunk’s start. An object whose last chunk ends ateis delivered, arriving ate + h, underdatagram, with no deadline, or wheneis at most its deadline; otherwise it stays pending and the next boundary resets its subgroup. - With nothing eligible, the time moves to the next availability of a pending object, which is a boundary too; the run ends when nothing is pending.
- Timeouts. In object order, each object still pending whose deadline is before the
boundary’s time: under
The sink.
- Ready. A delivered object is ready when it has arrived and, for layer
k ≥ 1, every unit it covers is at stagek − 1; it became ready at the later of its arrival and the time the last of those units reached that stage. - Workers. The sink has
workersnon-preemptive workers. At each event time it first completes the jobs ending then, then gives each free worker the least ready job: by (layer, then for a base the time it became ready,from, object index) underbase-first, so bases are decoded in the order they arrived, and by (the time it became ready, object index) underfifo. A job lastsdecode_us.baseordecode_us.refinement, once per object whatever its units. The next event is the next job end or the next time a job becomes ready. - Presentation. When a job ends at
t, every unit it covers is presented atlayers = k + 1att, and reaches stagek. - Deadlines. With
playout_start_usset, unitu’s deadline isplayout_start_us + u·unit_us. Each unit a job presents after its deadline counts one missed: a base is presented all the same; a refinement is presented underpresent, and underdiscardis not presented and the unit stays at its stage.
Outputs.
- per object:
sent_first_us,arrived_us, and its state:delivered,dropped,abandonedorcapped; - the presentations
(unit, layers, t_us), ordered by unit, then layers; bfpf_sched, the bytes of the chunks arrived by unit 0’s first presentation;tfpf_us, that presentation’s time; andtffs_us, unit 0’s first presentation at each layer count,nullwhere none;offered,delivered,dropped,abandonedandnever_sent_capped, in bytes, withoffered = delivered + dropped + abandoned;missed;- each unit’s stall (§10.4),
nullwhere the plan or the base-only model never presents its base.
A case of the corpus names a model, a plan, and its claim: within-bound (every
stall at most the plan’s bound), exceeds:<plan> (the plan has no bound, and its
largest stall passes the named plan’s), or unbounded (the plan has no bound, and only
its figures are checked).
Appendix C. Output tree
A commitment to every presentation of a session or a title: one leaf per (unit,
layers), so that a record or a claim commits to its outputs by the pair (root,
leaf_count) and any one output is provable by an inclusion path.
out_digest = SHA-256(canonical output) SPEC §7 Digests: RGB8, s16le PCM, the attribute tensor, a 32-byte state hash
leaf = SHA-256("wai:out-leaf\x01" ‖ u64_be(unit) ‖ u16_be(layers) ‖ u16_be(decoder) ‖ out_digest)
node = SHA-256("wai:out-node\x01" ‖ left ‖ right)
decoderis the index of the decoder entry that produced the output: entries are (capability, pin) pairs in first-use order, presentations taken by (layers, unit). For a title, the base’s capability with the pin it verified, then each refinement capability.- Leaves are ordered by (unit, layers), and a (unit, layers) given twice is refused.
- The root over
n ≥ 1leaves is RFC 6962’s Merkle tree hash: one leaf is itself; otherwisenode(MTH(D[0:k]), MTH(D[k:n])), withkthe largest power of two belown. No leaf is ever duplicated, so two leaf lists never give one root. - A commitment is the root with its leaf count; a root alone is not one. A record or a claim that commits to an output tree states both.
- An inclusion proof is RFC 6962’s audit path, checked as RFC 9162 §2.1.3.2 checks one under the leaf count the commitment states, never a count the proof’s holder supplies. An audit path does not fix the tree’s size: leaf 0’s path in a tree of 7 leaves also walks to that root when 8 is taken as its size.
layersis 16 bits, so it holds any stage count.
The JWP group root (jwp-receipts) is a separate, fixed wire format and is not this tree.
Appendix D. Measured titles
A measured title is a fixture directory with a title.json:
{ "format": "wai-measure-title/1", "title": "still", "about": "…",
"unit_us": 40000,
"bases": [ { "case": "staged-conformance/cases/image_int_hyper_2l", "base": 0,
"first_unit": 0, "capture_us": null } ],
"layers": [ "L01.wai", "L02.wai", "L03.wai" ],
"source": { "file": "source.bin", "sha256": "<64 hex>" },
"references": [ { "label": "one layer, q 1", "layers": [ "R0_00.wai" ] } ],
"expect": "expect.json" }
- The stage structure is the staged set’s own. Each base is a base of a
staged-conformance case, named by the case directory and
its index there: the case states its envelope, the parameter set it decodes against
(by member and path) and its stage-0 output. Each layer’s
refinesnames its base, producers, layer and span (staged-delivery §3).title.jsonadds only what that model does not carry. unit_usis one unit’s duration, andcapture_usa base’s capture time for a live delivery plan (nullfor one delivered on demand).sourceholds each unit’s canonical buffer at its final shape, in unit order, with its SHA-256.referencesare the single-layer references of §6.2: each is the title’s bases with the layers it lists.expect.jsonholds, from the fixture tool’s own encoder, the SHA-256 of every unit’s output at every stage, and of every reference’s final units. An evaluation MUST reproduce them.- A base MAY instead be named by its envelope alone,
{ "envelope": <path>, "set": { <member>: <path> }, "first_unit", "capture_us" }: it has no stage-0 output or case to check its decode against. limits(optional) lowers the limits the title is measured under:cost, four figures (output bytes, symbols, working bytes, multiply-accumulates), andchain_budget, eachnullfor the tool’s limit below. A title lowers a limit and never raises one (below).refused(optional) names the code the title’s evaluation MUST be refused with: a title that exists to be refused has no figures.- Every path is below
wai/; a layer, source and expectation file is a file name in the fixture directory.
A title’s content identity is BLAKE3 over "wai:measure-title\x01" and then, in
path order, each measured file’s path and bytes, each prefixed by its length as a
64-bit big-endian integer. The measured files are its objects (bases, layers and the
references’ layers), its bases’ parameter-set files and its source; labels,
title.json, stage-0 outputs and expect.json are not part of it.
Evaluation. Each base is decoded as a sink decodes one (SPEC §4: its capability,
then its fallback, a refinement-only or companion-only capability making the envelope
inert), against the parameter set its pin names (SPEC §3.1), and under the cost
limit, which its payload’s declared cost (integer-payloads §7) is compared with
before any symbol is decoded or any output allocated. It must give the case’s stage 0,
and the capability, fallback, tier and pin the case states. Its layers are applied as a
sink applies them (staged-delivery §4) under the chain budget, with the same
limit on each layer’s declared cost (Appendix A.5), one layer index at a time and,
within one, in the order the title lists them; every layer MUST apply to every unit
it covers, as in a set delivered whole and in order. The figures of §3–§6 and §9 and
the output tree follow from the outputs at every stage. A refusal on the way names its
code: the base decode’s (cost_over_limit, missing_capability, a pin refusal) or the
sink’s (chain-over-limit, a layer’s §4.2 reason). This composition of library
functions is for measurement, not a sink’s dispatch path.
The tools’ limits. The reference tools (wai stage, wai_stage_measure,
wai_stage_bench) decode every base under these cost figures, hold every layer to
their output, symbols and working bytes (staged-delivery Appendix A.5), and hold every
chain to the chain budget. A tool’s --limit OUT,SYMBOLS,WORKING,MACS replaces the four
cost figures. A title is measured under the lower, figure by figure, of its own
limits and the tool’s: a title lowers a limit and MUST NOT raise one, because
whoever writes a title is not whoever runs the tool, and raising a limit takes the
tool’s own --limit.
| figure | default | why |
|---|---|---|
| output bytes | 2²⁵ (32 MiB) | five 1920 × 1080 RGB8 frames, or minutes of 48 kHz stereo PCM: more than one stage set’s base carries, and below the declared clips a tool must refuse |
| symbols | 2²⁵ | one entropy-decode step per output byte at the output limit |
| working bytes | 2²⁸ (256 MiB) | eight bytes per network activation, for networks a few times the output’s size |
| multiply-accumulates | 2³⁸ | a learned decode of several megapixels |
| chain budget | 2²⁸ (256 MiB) | eight times the output limit, for a base’s units and the layers retained over them |