By Every Language Matters

ELM Voice Studio

The Complete Workspace for Multilingual Speech Data Collection. Built for creating high-quality ASR and TTS datasets in the languages the world actually speaks.

The Everylanguagematters Voice Studio workspace, showing an audio waveform, prompts and tools.
Everything you need to annotate — waveform, prompts, and tools in one view. Swipe the image to see the full window.
What is ELM Voice Studio

Professional speech annotation, built for annotators — not engineers.

The tool

ELM Voice Studio is a desktop annotation application designed specifically for building high-quality speech datasets used in Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems. It gives annotators a clean, focused workspace to record, review, segment, and label audio — without needing a technical background to get started.

Who it's for

Built for native speakers, community contributors, and professional annotation teams working across Every Language Matters projects and beyond. Whether you are annotating your first audio file or managing hundreds of hours of speech data, ELM Voice Studio scales with your workflow.

Why it matters

The quality of a voice AI system is only as good as the data it was trained on. ELM Voice Studio puts the tools for producing that data directly in the hands of the communities whose languages matter most — removing the technical barriers that have historically kept underrepresented languages out of the AI pipeline.

Features

Everything an annotator needs, nothing they don't.

Audio recording & playback

Record straight into the app, or open audio you already have. Play it back, slow it down, loop a section, and move through the waveform a fraction of a second at a time.

Filters & noise detection

The app points out background noise, clipping and long silences, and gives you filters to clean them up — so weak audio does not end up in your dataset.

Prompted recording

Load a list of sentences and record them one at a time. The app shows each prompt, saves the take, and moves on to the next — the quickest way to build a dataset from a script.

Export to training format

Export as WAV at the sample rate your training needs, with transcripts and labels saved alongside in a plain, readable format. Ready to hand straight to a training script.

Works offline, anywhere

Everything runs on your own machine. No internet needed to record, label or export, and your recordings never leave your computer unless you send them somewhere yourself.

Use Cases

Built for both ASR and TTS annotation workflows.

ASR

Automatic Speech Recognition

Type out what was said and line the text up with the audio. That pairing — the sound and the words — is what a speech recognition model learns from.

  • Listen to the audio
  • Type what you hear
  • Review the transcript
  • Validate the recording
TTS

Text-to-Speech

Read from a script and record clean, steady speech, then add the small details that make a synthetic voice sound natural — including tone languages and clicks.

  • Read the prompt
  • Record your voice
  • Review your takes
  • Validate recordings
How It Works

From your first recording to a production-ready dataset.

Step one

Set up your project

Start a project, pick ASR or TTS, choose the language, and set the labels you want to use. It takes a couple of minutes and the app walks you through it.

Step two

Record or import audio

Record with your microphone, or open WAV files you already have. WAV is what we recommend and what the app exports, so your audio keeps its full quality from the first take to the finished dataset.

Step three

Annotate with precision

Run the filters to clear background noise, hum and clipping, cut the dead air at the ends, and flag anything that still sounds wrong. Cleaning as you go beats fixing a whole set at the end.

Step four

Export & deliver

Export the project and you get the WAV files, the transcripts and labels, and a short summary of what is in the set — ready to train on, or to hand to whoever is doing the training.

Available In

9 languages. One tool for the world.

ELM Voice Studio's interface is fully available in the following languages, so annotators can work comfortably in their own.

English

English

Français

French

Español

Spanish

Português

Portuguese

中文

Chinese

Kiswahili

Swahili

हिन्दी

Hindi

Русский

Russian

العربية

Arabic

More coming soon

Feedback

Tell us what to build next.

ELM Voice Studio is shaped by the people annotating in it. If something is slowing you down, a language needs support we have not added, or you can think of a feature that would make the work easier — send it to us. Every submission is read.

Share feedback or request a feature

Takes about 5-7 minutes. No account needed.

Start annotating today.

ELM Voice Studio is free to download. Available for Windows, macOS, and Linux.