Backend loading…
Model loading…
1. Audio
● LIVE
Microphone (noise handling)
Noise floor not measured yet
2. Transcribe
Decoder settings
Defaults are tuned on held-out real recordings (onset 0.7, frame 0.4, min note 58 ms; Spotify's Python defaults are 0.5 / 0.3 / 128 ms), the same values the iOS app uses. Changing them re-decodes without re-running the model.
3. Take
The tempo decides note lengths: at double the tempo the same notes read twice as long (eighths become quarters).
Note list
| # | Pitch | MIDI | Start (s) | Dur (s) | Amp |
|---|
How it works
- Mic audio is captured by an
AudioWorkletin a 22,050 HzAudioContext, cleaned up like the iOS app (70 Hz rumble filter, adaptive noise floor, noise gate 10 dB above it) and streamed to a Web Worker; uploads are resampled with anOfflineAudioContext. - The worker runs each 2-second window (43,844 samples) as soon as it fills, every ~0.5 s live (~0.8 s for files), keeping only the newest frames of each window, through Spotify's Basic Pitch (ICASSP 2022, ~0.9 MB TF.js graph) on WebGPU, WebGL or WASM.
- The note and onset heads are decoded into notes with Basic Pitch's own JS port of
note_creation.py; live, the latest ~15 s are re-decoded after every window and older notes are frozen. - Piano roll first: live and after a take you see the piano roll. Only when you open Page (or export MIDI) does a JS port of EarSheet's quantizer estimate the whole take's tempo, meter and key and snap to 16ths; the result is cached, and the adjustable ♩ = N · Assumed tempo mark re-quantizes it. The grand staff (split at middle C) is rendered with abcjs.
Model and decoder: © Spotify AB, Apache-2.0. TensorFlow.js: Apache-2.0. abcjs and @tonejs/midi: MIT. Source: web/ in the EarSheet repo.