Runs locally No signup No watermark
Audio Tool
Speaker Diarization

This tool listens for voice activity and voice identity at the same time: a segmentation model marks where anyone is speaking in overlapping ten-second windows, and a speaker-embedding model turns each continuous stretch into a compact voiceprint. Stretches whose voiceprints agree are grouped into one speaker, so the output is a timeline, Speaker A from here to here, Speaker B next, plus each speaker's total talk time. It works best on clear conversations such as interviews, meetings and podcasts. Processing stays on your device and the timeline exports as CSV.

✓ Core Access ✓ No signup 🔒 Supported file contents are not uploaded for processing
1

Upload

2

Adjust & Preview

Pixlane
Preparing the tool on your device

Pixlane is loading on your device. Supported file contents are not uploaded for processing. Site analytics and the assistant are separate; tool-specific limits may apply.

Speaker Diarization: Who Spoke When in Audio

Split a recording into who-spoke-when: a speaker timeline with per-speaker talk time, built in your browser.

This tool listens for voice activity and voice identity at the same time: a segmentation model marks where anyone is speaking in overlapping ten-second windows, and a speaker-embedding model turns each continuous stretch into a compact voiceprint. Stretches whose voiceprints agree are grouped into one speaker, so the output is a timeline, Speaker A from here to here, Speaker B next, plus each speaker's total talk time. It works best on clear conversations such as interviews, meetings and podcasts. Processing stays on your device and the timeline exports as CSV.

How to Use It

  1. Upload your audio file and wait for the local preview to load.
  2. Adjust the settings: Adjust the available settings and run the tool locally in your browser.
  3. Download the result: Download the processed file directly to your device.

Frequently Asked Questions

What exactly does this produce?

A speaker timeline: an on-screen list of turns like "00:12.3 – 00:31.7 Speaker B", a per-speaker talk-time summary, and a CSV download with start, end and speaker columns that opens directly in a spreadsheet or feeds a subtitling workflow.

Does it transcribe what was said?

No. Diarization answers who spoke when, not what was said. The two are complementary: many transcription workflows run diarization first and attach the words to each speaker afterwards.

How many speakers can it tell apart?

The segmentation stage tracks up to three people speaking at the same moment, and the grouping stage can label more speakers than that across the whole recording. Very similar voices, heavy noise or strong echo make separation harder. The Speaker separation setting exists for exactly those cases.

Why did one voice get split into two speakers?

Usually the voice changed character mid-recording: distance from the microphone, emotion, a phone speaker versus a headset. Switch Speaker separation to Relaxed, which merges voiceprints more readily, and re-run.

Does my recording leave my device?

Supported file contents are processed locally in your browser through the site's neural runtime. The model files are fetched once and cached; the audio itself is not part of any upload.

What audio works best?

Clear speech with each person mostly taking turns: interviews, meetings, podcasts, calls. Heavy crosstalk, music beds and field recordings still produce a timeline, but boundaries get less precise as conditions degrade.

Back to audio tools