WAI Extension: Score (Music as Instructions)
Mirrored from the canonical text at commit 117bad22 ().
Status: Draft. Phase 6 — the deepest audio lever. Ship the composition, not the recording: a score of note events the sink synthesizes. Conformance is mixdown-equivalence, the music counterpart of spatial audio. Reference impl: the
scorefeature ofwai-rs; corpus:score-conformance/; live sink: wai.transaction.science/score. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.
1. Scope and model
A wai.audio.score object carries no recorded audio — only note
events. Where wai.audio.scene mixes recorded sound objects, a score
is reconstructed from instructions: a deterministic reference synth
turns the notes into a canonical PCM mixdown, and a sink with a richer
instrument (soundfont, neural) renders the same score for presentation.
It is lever 3 (instructions-at-the-sink) for music — a minute of music
is kilobytes of notes, not megabytes of samples.
In scope: the container, the note model, the deterministic reference synthesis, the mixdown-equivalence conformance criterion, the receipt profile. Out of scope: a transport; the sink’s instrument (soundfonts, neural synths — presentation, not conformance); MIDI-file import (an authoring concern).
Relationship to the core spec: mirrors spatial audio. The reconstruct
runs in the wai.det.fixed64 floor; media takes the value "audio";
wai.audio.score is authoritative.
2. Container (WSCR)
"WSCR" | u16 n_sections | section table (u8 kind, u32 off, u32 len) | sections
| kind | section | required |
|---|---|---|
0x01 | contract (canonical JSON) | REQUIRED |
0x02 | note events | REQUIRED |
2.1 Contract
{ "channels": 1, "duration": 34000, "numeric": "wai.det.fixed64",
"sample_rate": 22050 }
duration is the rendered mixdown length in samples.
2.2 Notes
n_notes u32, then per note (65 bytes):
start u64 | dur u64 | phase_inc i64 Fx | gain i64 Fx | waveform u8 | attack u64 | decay u64 | sustain i64 Fx | release u64.
phase_inc is the per-sample phase advance freq / sample_rate,
authored off-line from the pitch — so the floor never evaluates an
exponential to turn a note into a frequency (the same authoring
discipline as the avatar quaternion). waveform selects a reference
timbre: 0 saw, 1 square, 2 triangle, 3 25%-pulse.
3. Reference synthesis (trig-free)
For each note, a phase accumulator advances by phase_inc each sample
(wrapping at 1.0); a piecewise-linear oscillator maps the phase to
[-1, 1] (no trig); a linear ADSR envelope (attack ramp → decay to
sustain → held until dur → release to zero over release) scales
it; the result times gain accumulates into the mixdown buffer. Each
output sample is quantized to i16 (round(v · 32767), half away from
zero, clamped) in integer space. The whole path is + − × ÷ in the
Fx floor — bit-identical on every machine.
4. Conformance — mixdown-equivalence
Given the same
wai.audio.score, a conforming sink MUST reconstruct the identical canonicali16PCM via the reference synth, hence the identical BLAKE3 — on every machine, no tolerance parameter.
mixdown_hash = BLAKE3("wai:score-mixdown\x01" || i16le PCM)
score_hash = BLAKE3("wai:score\x01" || contract || notes) // identity
A sink MAY play a soundfont/neural instrument instead; that is presentation and does not affect conformance, exactly as the audio-scene HRTF render or the worlds renderer does not.
5. Receipt (JWP profile)
score_hash is the identity; the group’s objects are the mix PCM blocks
(per a fixed sample span), Merkle-bound; one Ed25519 signature over
score_hash + root + exact work (= duration × notes) + carried
joules_micro; parent_receipt_hash chains — the worlds/audio receipt
shape reused. The figure’s acquisition class rides beside it as an optional signed label
(energy-measurement §2.8): absent, the receipt keeps
its legacy bytes and the figure is unlabelled; present, it is inside the
signature. The reference meter labels every figure it measures
OnChipCounter, with its declared uncertainty.
Appendix
Mixdown-equivalence is spatial audio’s criterion with the recorded
objects replaced by note events: both render a deterministic i16 mix
in the Fx floor and hash it. The score is the most extreme
instructions-at-the-sink case in audio — there is no recording anywhere
in the payload, only the composition. Sight, hearing, touch, motion,
state, captured-content provenance — and now the composition itself —
are one floor, one verification.