Official BytePlus partnerBuilt on BytePlus, ByteDance's official AI platform

Multilingual text to speech

Multilingual text to speech in 24 languages

KreatorFlow speaks 24 languages and accents. Audio Generation voices any of them from a written prompt, and Text to Speech has 45 preset voices across 14 languages for a fixed, reusable voice.

  • 24 languages in Audio Generation
  • 14 languages with preset voices
  • 45 preset voices

Two ways to speak many languages

Audio Generation: 24 languages

Pick a language and the studio asks the model to speak it. The voice is described in your prompt, not chosen from a list, so the result can vary between runs. Prompts take up to 3,000 characters and 2 credits per 1,000 characters.

Text to Speech: 14 languages

45 preset voices, each listed under the language it speaks, so the same voice sounds the same every time. Scripts take up to 5,000 characters and 1 credit per 1,000 characters.

Supported languages

Audio Generation (24, prompt-steered)

  • English
  • American English
  • Indian English
  • Mandarin
  • Cantonese
  • Spanish
  • Mexican Spanish
  • French
  • German
  • Italian
  • Portuguese
  • Brazilian Portuguese
  • Japanese
  • Korean
  • Russian
  • Arabic
  • Indonesian
  • Malay
  • Thai
  • Vietnamese
  • Hindi
  • Dutch
  • Turkish
  • Polish

Text to Speech (14, preset voices)

  • English
  • Mandarin
  • Japanese
  • Spanish
  • Portuguese
  • Indonesian
  • French
  • German
  • Korean
  • Arabic
  • Russian
  • Italian
  • Vietnamese
  • Thai

Preset voices by language

One preset voice per Text to Speech language, under its real studio name.

Which one should I use?

Use Text to Speech when you need the same voice across many clips, such as a brand voice or a series. Use Audio Generation for languages and accents without a preset voice, such as Cantonese, Hindi, Dutch, Turkish or Polish, and for expressive scenes you want to describe in words.

Multilingual text to speech questions

How many languages are supported?

24 in Audio Generation and 14 in Text to Speech, where 45 preset voices are each listed under the language they speak.

Does every preset voice speak every language?

No. Each preset voice is listed under one language. For a language without a preset voice, use Audio Generation and describe the voice in your prompt.

What does prompt-steered mean?

In Audio Generation you choose the language and describe the voice and delivery in words. There is no fixed voice per language, so two runs of the same prompt can sound different.

How much does it cost?

Text to Speech costs 1 credit and Audio Generation 2 credits per started 1,000 characters, so the shortest clip costs 1 credit in Text to Speech and 2 credits in Audio Generation. The cost is shown before you generate.

Can I translate a recording into another language?

Yes, with Live Interpretation, which translates a voice clip into one of 11 languages for 3 credits per clip.

Which platform does the audio studio run on?

Built on BytePlus, ByteDance's official AI platform. KreatorFlow is an official BytePlus partner; the audio studio and its credit prices are KreatorFlow's own.

Speak to every audience

Generate speech in 24 languages from one prompt, or pick one of 45 preset voices.

Keep exploring