WhatIsTTS
WhatIsTTS

Home / Speech Translation

Speech Translation

Translate spoken audio between 50+ languages while preserving the original speaker's voice characteristics.

Upload Audio to Translate

Upload spoken audio in any supported language. We will translate the speech and regenerate it in the target language.

Drag & drop your file here, or browse

Supports MP3, WAV, FLAC, OGG, M4A. Max 50MB.

file.mp3

0 MB
Translating speech... Transcribing, translating, and regenerating audio.

Translated Audio

0:00 0:00

Translated Text

Credit Cost

5 credits per minute of audio

Billed based on source audio duration, rounded up to the nearest minute.

How It Works

  1. Speech is transcribed using AI (STT)
  2. Text is translated to the target language
  3. Translated text is spoken using voice cloning to match the original speaker
  4. You get both the translated audio and text

Supported Languages

We support translation between 50+ language pairs. Voice preservation works best with major languages supported by our providers: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, and more.

Best For

Content localization, multilingual publishing, dubbing, international communications, and language learning.

Frequently Asked Questions

Speech translation converts spoken audio in one language into spoken audio in another language, preserving the original speaker's voice characteristics. It combines speech recognition, text translation, and voice cloning.

We support translation between 100+ languages using our speech-to-text providers, and voice output in 140+ languages via our TTS providers. The most popular pairs are English ↔ Spanish, English ↔ Chinese, and English ↔ French.

Speech translation costs 5 credits per minute of audio processed. This covers speech-to-text transcription, text translation, and TTS re-synthesis in the target language. The cost is based on the duration of the input audio.

Yes. The pipeline uses voice cloning to maintain your voice characteristics in the translated output. Your voice is captured from the input audio and applied to the TTS generation in the target language, so the translated speech sounds like you speaking the new language.

Upload audio in MP3, WAV, FLAC, OGG, M4A, or WEBM format, up to 50MB. The translated output is delivered as MP3 by default. For other formats, use our Audio Converter tool after translation.

Translation accuracy depends on the language pair, audio clarity, and source content. Common language pairs like English to Spanish or English to French achieve high accuracy. Less common pairs may have lower quality. Clear, well-paced speech produces the best results.

Yes. Upload a video file (MP4, MOV, WEBM) and the audio track will be extracted, translated, and re-synthesized in the target language. The translated audio is returned as a separate file that you can then sync with your video using a video editor.

Free users can translate up to 30 seconds of audio. Paid plans support files up to 10 minutes per request. For longer content like full lectures or podcast episodes, split the audio into segments and translate them individually.

Yes. Our API supports programmatic speech translation. Upload the source audio, specify the target language, and receive the translated audio file. This enables integration into dubbing workflows, e-learning platforms, and content localization pipelines.

Yes. Use our Speech to Text tool first to transcribe the audio, review and correct any errors, then translate the corrected text using text-to-speech in the target language. This two-step approach gives you full control over the translation accuracy.

Popular uses include dubbing videos and films into other languages, translating interviews and meetings for international teams, localizing e-learning courses and training materials, making podcasts accessible to global audiences, and real-time translation for multilingual events.

Audio files are transmitted securely over HTTPS and processed through our speech-to-text and TTS provider pipeline. Files are not stored permanently on our servers and are deleted after processing. We do not use your recordings for training. Each upstream provider has its own data retention policy.

Need a specific TTS workflow?

Compare providers, test voices, then run it through one brokered API.