Skip to main content

WAI Extension: Score (Music as Instructions)

Mirrored from the canonical text at commit 117bad22 ().

Status: Draft. Phase 6 — the deepest audio lever. Ship the composition, not the recording: a score of note events the sink synthesizes. Conformance is mixdown-equivalence, the music counterpart of spatial audio. Reference impl: the score feature of wai-rs; corpus: score-conformance/; live sink: wai.transaction.science/score. Keywords MUST, MUST NOT, SHOULD, MAY are RFC 2119/8174.

1. Scope and model

A wai.audio.score object carries no recorded audio — only note events. Where wai.audio.scene mixes recorded sound objects, a score is reconstructed from instructions: a deterministic reference synth turns the notes into a canonical PCM mixdown, and a sink with a richer instrument (soundfont, neural) renders the same score for presentation. It is lever 3 (instructions-at-the-sink) for music — a minute of music is kilobytes of notes, not megabytes of samples.

In scope: the container, the note model, the deterministic reference synthesis, the mixdown-equivalence conformance criterion, the receipt profile. Out of scope: a transport; the sink’s instrument (soundfonts, neural synths — presentation, not conformance); MIDI-file import (an authoring concern).

Relationship to the core spec: mirrors spatial audio. The reconstruct runs in the wai.det.fixed64 floor; media takes the value "audio"; wai.audio.score is authoritative.

2. Container (WSCR)

"WSCR" | u16 n_sections | section table (u8 kind, u32 off, u32 len) | sections
kindsectionrequired
0x01contract (canonical JSON)REQUIRED
0x02note eventsREQUIRED

2.1 Contract

{ "channels": 1, "duration": 34000, "numeric": "wai.det.fixed64",
  "sample_rate": 22050 }

duration is the rendered mixdown length in samples.

2.2 Notes

n_notes u32, then per note (65 bytes): start u64 | dur u64 | phase_inc i64 Fx | gain i64 Fx | waveform u8 | attack u64 | decay u64 | sustain i64 Fx | release u64.

phase_inc is the per-sample phase advance freq / sample_rate, authored off-line from the pitch — so the floor never evaluates an exponential to turn a note into a frequency (the same authoring discipline as the avatar quaternion). waveform selects a reference timbre: 0 saw, 1 square, 2 triangle, 3 25%-pulse.

3. Reference synthesis (trig-free)

For each note, a phase accumulator advances by phase_inc each sample (wrapping at 1.0); a piecewise-linear oscillator maps the phase to [-1, 1] (no trig); a linear ADSR envelope (attack ramp → decay to sustain → held until dur → release to zero over release) scales it; the result times gain accumulates into the mixdown buffer. Each output sample is quantized to i16 (round(v · 32767), half away from zero, clamped) in integer space. The whole path is + − × ÷ in the Fx floor — bit-identical on every machine.

4. Conformance — mixdown-equivalence

Given the same wai.audio.score, a conforming sink MUST reconstruct the identical canonical i16 PCM via the reference synth, hence the identical BLAKE3 — on every machine, no tolerance parameter.

mixdown_hash = BLAKE3("wai:score-mixdown\x01" || i16le PCM)
score_hash   = BLAKE3("wai:score\x01" || contract || notes)   // identity

A sink MAY play a soundfont/neural instrument instead; that is presentation and does not affect conformance, exactly as the audio-scene HRTF render or the worlds renderer does not.

5. Receipt (JWP profile)

score_hash is the identity; the group’s objects are the mix PCM blocks (per a fixed sample span), Merkle-bound; one Ed25519 signature over score_hash + root + exact work (= duration × notes) + carried joules_micro; parent_receipt_hash chains — the worlds/audio receipt shape reused. The figure’s acquisition class rides beside it as an optional signed label (energy-measurement §2.8): absent, the receipt keeps its legacy bytes and the figure is unlabelled; present, it is inside the signature. The reference meter labels every figure it measures OnChipCounter, with its declared uncertainty.

Appendix

Mixdown-equivalence is spatial audio’s criterion with the recorded objects replaced by note events: both render a deterministic i16 mix in the Fx floor and hash it. The score is the most extreme instructions-at-the-sink case in audio — there is no recording anywhere in the payload, only the composition. Sight, hearing, touch, motion, state, captured-content provenance — and now the composition itself — are one floor, one verification.