WAI Extension: Object-Based Spatial Audio
Mirrored from the canonical text at commit 117bad22 ().
Status: Draft. Phase 1 of the cargo roadmap. The audio analog of interactive-worlds: ship sound objects + a scene, render to the listener’s own layout. Conformance is mixdown-equivalence (bit-exact PCM for a declared layout), the audio counterpart of replay-equivalence. Reference impl: the
audio_scenefeature ofwai-rs(mixer + receipt); conformance corpus:audio-conformance/; live sink:/audio. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.
1. Scope and model
A wai.audio.scene object describes a set of audio objects — each a
content-addressed source plus a position trajectory and gain — that a
sink renders to its own speaker layout. Today’s audio ships a fixed
channel mix (stereo, 5.1); this ships the scene and lets the sink make
the mix it can actually play. It is lever 3 (instructions-at-the-sink)
for hearing.
In scope: the scene container, the deterministic mixdown render contract, the mixdown-equivalence conformance criterion, the fallback floor, and the receipt profile.
Out of scope: a transport; the source codecs (sources are ordinary
wai.audio.* WAI objects, decoded by the existing path); HRTF design
(an HRTF is named as a requirement-of-existence, never shipped).
Relationship to the core spec: this mirrors interactive-worlds exactly.
A scene is a content-addressed artifact; sources are content-addressed
WAI assets (§4.1 discipline, domain wai:asset\x01); the render is
deterministic under a declared numeric profile; receipts are a JWP
group profile. media takes the informational value "audio-scene";
the capability wai.audio.scene is authoritative.
2. Container (WSCN)
"WSCN" | u16 n_sections | section table | sections
| n × (u8 kind, u32 off, u32 len), 8-byte-aligned bodies
| kind | section | required |
|---|---|---|
0x01 | scene contract (canonical JSON) | REQUIRED |
0x02 | object table | REQUIRED |
0x03 | asset table (§4.1, the source wai.audio.* objects) | REQUIRED |
0x04 | seal: duration + final mixdown hash + Ed25519 sig | OPTIONAL (a sealed scene pins its reference mix) |
Unknown section kinds MUST be ignored (forward compat within MINOR).
2.1 Scene contract (canonical JSON)
{
"sample_rate": 48000,
"numeric": "wai.det.fixed64",
"pan_law": "wai.audio.pan.sqrt_constpower",
"distance": "wai.audio.dist.inverse_clamped",
"duration": 240000,
"floor_layout":"stereo"
}
numeric is the same Q32.32 fixed-point floor as worlds — the mixdown
is bit-exact on every machine. floor_layout (stereo) is the
mandatory render target; richer layouts (5.1, binaural@<hrtf-id>)
are OPTIONAL and, for binaural, perceptual (the HRTF is a
requirement-of-existence, not bit-exact across HRTFs).
2.2 Object table
n_objects u32, then per object:
source_hash (32 B) // content hash of a wai.audio.* WAI envelope (asset table)
gain (i64 Fx)
n_keys (u16)
keys: n × ( pos_sample u64 | x i64 | y i64 | z i64 ) // listener-relative position, Fx metres
Position between keyframes is linear-interpolated in Fx (so it is
reproducible). One keyframe ⇒ a static source.
3. Render contract (mixdown)
For output channel c and sample t, in Fx:
out[c][t] = Σ_objects src_o[t] · gain_o · dist_atten(|p_o(t)|) · pan_gain(p_o(t), c)
src_o[t]: objecto’s source as mono PCM atsample_rate. v0 requires sources already atsample_rate(a deterministic resampler is a registered sub-detail, not yet in the reference impl).dist_atten(d) = 1 / (1 + d)clamped to[0,1](inverse_clamped), whered = sqrt(x² + y² + z²).pan_gainforsqrt_constpowerstereo — chosen to be exactly computable in theFxfloor (sqrt only, no transcendental tables):pan = clamp(x / (|x| + |z|), -1, 1)(0 when|x|+|z|=0), thengL = sqrt((1 - pan)/2),gR = sqrt((1 + pan)/2)— constant power (gL² + gR² = 1).- Accumulate in
Fx, then quantize to the canonical output (interleaved i16 LE) by round-half-away-from-zero in integer space.
4. Conformance — mixdown-equivalence
Given the same scene and the same declared output layout, a conforming sink MUST produce the identical canonical PCM, hence the identical
BLAKE3over it — on every machine, with no tolerance parameter.
mixdown_hash = BLAKE3("wai:audio-mixdown\x01" || layout_id || pcm_i16le)
The mixdown is a function of the sources’ decoded PCM, so the hash is portable
only as far as that PCM is. A source whose capability promises the same bytes on
every sink (wai.audio.flac, an integer codec) keeps it portable; a source decoded
through a codec whose standard conforms decoders to a tolerance (wai.audio.opus,
SPEC §7, StandardTolerance) may decode differently on two conforming sinks, and
then so does the mixdown. The reference registry reports this:
codecs::is_cross_decoder_exact("wai.audio.scene") is false, because a scene’s
payload names its sources by hash and not by capability, and
codecs::is_cross_decoder_exact_with_sources is true once the capability of
every source envelope is known and each promises the same bytes on every sink. It
answers false when it cannot tell: when no source is given (a scene with no
objects included), or when a source is not a wai.audio.* capability (§2.2).
The stereo floor is mandatory and bit-exact. binaural@<hrtf> is
perceptual (conformance is similarity, not bytes — a sink advertising
it states the metric, like the wai.semantic.* path). The conformance
corpus ships scenes + the expected stereo mixdown hash; pass = hash
equality.
5. Fallback floor
model_requirement.fallback SHOULD name wai.audio.opus (or flac),
and the encoder SHOULD ship the WAI2 multi-rendition form: the scene
plus a pre-rendered stereo recording. A sink without the scene runtime
plays the recording. Scene-capable sinks render to their real layout;
below-floor sinks hear a fixed mix — both opened one envelope.
6. Session receipts (JWP profile)
A render session receipts as a JWP group (jwp-receipts.md) — the worlds
SessionReceipt transposed to audio (SceneReceipt in the reference
impl):
- identity
scene_hash=BLAKE3("wai:audio-scene\x01" || contract_bytes || object_table_bytes)(theworld_hashanalog). - the group’s objects = the rendered mix’s PCM blocks (fixed span of
frames), each
BLAKE3("wai:audio-block\x01" || block_index || i16le), Merkle-bound with JWP’s exact leaf/node domains. - one Ed25519 signature over a profile payload carrying
scene_hash, the Merkle root, the block list,work_mac(= output samples × objects mixed — exact, signed) andjoules_micro(carried; the nativewai_world_meterpattern supplies a measured figure,0if unmetered), with the figure’s acquisition class as an optional signed label (energy-measurement §2.8) — absent, the receipt keeps its legacy bytes and the figure is unlabelled. parent_receipt_hashchains successive renders into an append-only signed timeline, exactly as worlds’ receipts do.
Appendix — why this is the first cargo class
Mixdown-equivalence is replay-equivalence with the tick replaced by the
audio sample; the Fx floor, the content-addressed assets, the seal,
the receipt profile, and the fallback floor are all reused verbatim. A
spatial-audio scene is, structurally, a world whose “sim” is a mixer,
so the interactive-worlds rules carry over to a second sense.