Runs locally No signup No watermark
Audio Tool
Audio to Text

The model here is Whisper base, the small multilingual member of the family, quantized to run inside a browser tab. It reads the recording as a mel spectrogram in thirty-second windows and writes the text out token by token, so punctuation and casing come from the model rather than a post-processing pass. Language is detected from the audio itself unless you name one, which is worth doing when a recording opens with music or silence. Being the base model, it is fast and small but not infallible: clear single-speaker speech transcribes well, while heavy accents, crosstalk, and strings of digits are where it slips. Read the transcript before trusting it, and use the language selector when you already know what was spoken.

✓ Core Access ✓ No signup 🔒 Supported file contents are not uploaded for processing
1

Upload

2

Adjust & Preview

Pixlane
Preparing the tool on your device

Pixlane is loading on your device. Supported file contents are not uploaded for processing. Site analytics and the assistant are separate; tool-specific limits may apply.

Audio to Text: Free Speech to Text Transcription

Turn speech into text with a multilingual on-device model: transcripts for interviews, notes and lectures, built in your browser.

The model here is Whisper base, the small multilingual member of the family, quantized to run inside a browser tab. It reads the recording as a mel spectrogram in thirty-second windows and writes the text out token by token, so punctuation and casing come from the model rather than a post-processing pass. Language is detected from the audio itself unless you name one, which is worth doing when a recording opens with music or silence. Being the base model, it is fast and small but not infallible: clear single-speaker speech transcribes well, while heavy accents, crosstalk, and strings of digits are where it slips. Read the transcript before trusting it, and use the language selector when you already know what was spoken.

How to Use It

  1. Upload your audio file and wait for the local preview to load.
  2. Adjust the settings: Leave the language on auto-detect or pick the spoken language, then start the transcription; the model downloads once before the first run.
  3. Download the result: Download the processed file directly to your device.

Controls and Options

Auto-detect suits most recordings. Name the language when a clip opens with music, noise or silence, or when a bilingual speaker switches mid-sentence. Read the transcript over before using it. The base model is fast and small, so digits and crosstalk are where it slips.

Formats and Output

Reads common audio formats your browser can decode, including MP3, WAV, M4A, OGG, and FLAC, and exports the transcript as a plain .txt file.

Frequently Asked Questions

How accurate is it?

It is the base-size Whisper model, chosen so the download stays reasonable for a browser tool. On clear, single-speaker recordings it is dependable; accents, overlapping voices, background music and long digit strings are where errors show up. Treat the output as a first draft to read over, not a finished record.

Which languages does it handle?

The model is multilingual and detects the language on its own. The selector lists the languages we tested explicitly, including Turkish; auto-detect covers others the model knows, and naming the language helps when a recording starts with music or silence.

Why does the first run take longer?

The model downloads once per session before the first transcription. After that, later runs skip straight to the audio.

Does it add timestamps or separate speakers?

No. This card returns continuous text. For a who-spoke-when timeline, run the recording through Speaker Diarization instead, which produces a labeled speaker timeline with talk time.

How long a recording can I transcribe?

Up to thirty minutes in one pass; longer files are cut at that point and the result says so. The model works in thirty-second windows internally, so length mostly costs time rather than accuracy.

Is my audio uploaded anywhere?

Supported media processing runs on your device and file contents are not uploaded for processing. Site analytics and the assistant are separate.

Back to audio tools