EarSheet web demo

Play into the mic and watch the notes appear as you play, or upload a recording. The same Basic Pitch model the EarSheet iOS app uses runs here on your device through TensorFlow.js. Your audio never leaves the browser.

Backend loading…
Model loading…

1. Audio

Microphone (noise handling)
Noise floor not measured yet

2. Transcribe

Decoder settings

Defaults are tuned on held-out real recordings (onset 0.7, frame 0.4, min note 58 ms; Spotify's Python defaults are 0.5 / 0.3 / 128 ms), the same values the iOS app uses. Changing them re-decodes without re-running the model.

How it works

  1. Mic audio is captured by an AudioWorklet in a 22,050 Hz AudioContext, cleaned up like the iOS app (70 Hz rumble filter, adaptive noise floor, noise gate 10 dB above it) and streamed to a Web Worker; uploads are resampled with an OfflineAudioContext.
  2. The worker runs each 2-second window (43,844 samples) as soon as it fills, every ~0.5 s live (~0.8 s for files), keeping only the newest frames of each window, through Spotify's Basic Pitch (ICASSP 2022, ~0.9 MB TF.js graph) on WebGPU, WebGL or WASM.
  3. The note and onset heads are decoded into notes with Basic Pitch's own JS port of note_creation.py; live, the latest ~15 s are re-decoded after every window and older notes are frozen.
  4. Piano roll first: live and after a take you see the piano roll. Only when you open Page (or export MIDI) does a JS port of EarSheet's quantizer estimate the whole take's tempo, meter and key and snap to 16ths; the result is cached, and the adjustable ♩ = N · Assumed tempo mark re-quantizes it. The grand staff (split at middle C) is rendered with abcjs.

Model and decoder: © Spotify AB, Apache-2.0. TensorFlow.js: Apache-2.0. abcjs and @tonejs/midi: MIT. Source: web/ in the EarSheet repo.