← All of Studio

Audio annotator

Transcription and diarization on a multi-track waveform.

Region labels with millisecond precision. Transcribe, tag speakers, and mark events on the same waveform, no jumping between a player and a spreadsheet.

Transcribe, diarize, and tag in one pass

Annotators work directly on the waveform: draw a region, assign a speaker, type the transcript, and move on. No context-switching between audio player and text editor.

  • Multi-track view for overlapping speakers
  • Speaker labels that persist across a session
  • Millisecond-precision region boundaries
  • Non-speech event tagging (music, noise, silence)
00:00 SPEAKER_1 · 00:42–01:18 02:14
Transcript
"So the yield curve inverted back in—" [SPEAKER_2 overlap]

Built for speech datasets

Speaker diarization

Tag and re-identify speakers across a long recording, with overlap handling for cross-talk.

Playback-synced editing

Transcript text stays locked to its waveform region as you scrub, insert, or trim.

Timestamped export

Export to SRT, VTT, or JSONL with speaker labels and millisecond timestamps intact.

Turn raw audio into training data.