Skip to main content

WAI Extension: Object-Based Spatial Audio

Mirrored from the canonical text at commit 117bad22 ().

Status: Draft. Phase 1 of the cargo roadmap. The audio analog of interactive-worlds: ship sound objects + a scene, render to the listener’s own layout. Conformance is mixdown-equivalence (bit-exact PCM for a declared layout), the audio counterpart of replay-equivalence. Reference impl: the audio_scene feature of wai-rs (mixer + receipt); conformance corpus: audio-conformance/; live sink: /audio. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.

1. Scope and model

A wai.audio.scene object describes a set of audio objects — each a content-addressed source plus a position trajectory and gain — that a sink renders to its own speaker layout. Today’s audio ships a fixed channel mix (stereo, 5.1); this ships the scene and lets the sink make the mix it can actually play. It is lever 3 (instructions-at-the-sink) for hearing.

In scope: the scene container, the deterministic mixdown render contract, the mixdown-equivalence conformance criterion, the fallback floor, and the receipt profile.

Out of scope: a transport; the source codecs (sources are ordinary wai.audio.* WAI objects, decoded by the existing path); HRTF design (an HRTF is named as a requirement-of-existence, never shipped).

Relationship to the core spec: this mirrors interactive-worlds exactly. A scene is a content-addressed artifact; sources are content-addressed WAI assets (§4.1 discipline, domain wai:asset\x01); the render is deterministic under a declared numeric profile; receipts are a JWP group profile. media takes the informational value "audio-scene"; the capability wai.audio.scene is authoritative.

2. Container (WSCN)

"WSCN" | u16 n_sections | section table | sections
        | n × (u8 kind, u32 off, u32 len), 8-byte-aligned bodies
kindsectionrequired
0x01scene contract (canonical JSON)REQUIRED
0x02object tableREQUIRED
0x03asset table (§4.1, the source wai.audio.* objects)REQUIRED
0x04seal: duration + final mixdown hash + Ed25519 sigOPTIONAL (a sealed scene pins its reference mix)

Unknown section kinds MUST be ignored (forward compat within MINOR).

2.1 Scene contract (canonical JSON)

{
  "sample_rate": 48000,
  "numeric":     "wai.det.fixed64",
  "pan_law":     "wai.audio.pan.sqrt_constpower",
  "distance":    "wai.audio.dist.inverse_clamped",
  "duration":    240000,
  "floor_layout":"stereo"
}

numeric is the same Q32.32 fixed-point floor as worlds — the mixdown is bit-exact on every machine. floor_layout (stereo) is the mandatory render target; richer layouts (5.1, binaural@<hrtf-id>) are OPTIONAL and, for binaural, perceptual (the HRTF is a requirement-of-existence, not bit-exact across HRTFs).

2.2 Object table

n_objects u32, then per object:

source_hash (32 B)        // content hash of a wai.audio.* WAI envelope (asset table)
gain        (i64 Fx)
n_keys      (u16)
keys: n × ( pos_sample u64 | x i64 | y i64 | z i64 )   // listener-relative position, Fx metres

Position between keyframes is linear-interpolated in Fx (so it is reproducible). One keyframe ⇒ a static source.

3. Render contract (mixdown)

For output channel c and sample t, in Fx:

out[c][t] = Σ_objects  src_o[t] · gain_o · dist_atten(|p_o(t)|) · pan_gain(p_o(t), c)

4. Conformance — mixdown-equivalence

Given the same scene and the same declared output layout, a conforming sink MUST produce the identical canonical PCM, hence the identical BLAKE3 over it — on every machine, with no tolerance parameter.

mixdown_hash = BLAKE3("wai:audio-mixdown\x01" || layout_id || pcm_i16le)

The mixdown is a function of the sources’ decoded PCM, so the hash is portable only as far as that PCM is. A source whose capability promises the same bytes on every sink (wai.audio.flac, an integer codec) keeps it portable; a source decoded through a codec whose standard conforms decoders to a tolerance (wai.audio.opus, SPEC §7, StandardTolerance) may decode differently on two conforming sinks, and then so does the mixdown. The reference registry reports this: codecs::is_cross_decoder_exact("wai.audio.scene") is false, because a scene’s payload names its sources by hash and not by capability, and codecs::is_cross_decoder_exact_with_sources is true once the capability of every source envelope is known and each promises the same bytes on every sink. It answers false when it cannot tell: when no source is given (a scene with no objects included), or when a source is not a wai.audio.* capability (§2.2).

The stereo floor is mandatory and bit-exact. binaural@<hrtf> is perceptual (conformance is similarity, not bytes — a sink advertising it states the metric, like the wai.semantic.* path). The conformance corpus ships scenes + the expected stereo mixdown hash; pass = hash equality.

5. Fallback floor

model_requirement.fallback SHOULD name wai.audio.opus (or flac), and the encoder SHOULD ship the WAI2 multi-rendition form: the scene plus a pre-rendered stereo recording. A sink without the scene runtime plays the recording. Scene-capable sinks render to their real layout; below-floor sinks hear a fixed mix — both opened one envelope.

6. Session receipts (JWP profile)

A render session receipts as a JWP group (jwp-receipts.md) — the worlds SessionReceipt transposed to audio (SceneReceipt in the reference impl):

Appendix — why this is the first cargo class

Mixdown-equivalence is replay-equivalence with the tick replaced by the audio sample; the Fx floor, the content-addressed assets, the seal, the receipt profile, and the fallback floor are all reused verbatim. A spatial-audio scene is, structurally, a world whose “sim” is a mixer, so the interactive-worlds rules carry over to a second sense.