useTextToSpeech
Text to speech is a task that allows to transform written text into spoken language. It is commonly used to implement features such as voice assistants, accessibility tools, or audiobooks.
It is recommended to use models provided by us, which are available at our Hugging Face repositories: Kokoro and Supertonic 3. You can also use constants shipped with our library.
API Reference
- For detailed API Reference for
useTextToSpeechsee:useTextToSpeechAPI Reference. - For all text to speech models available out-of-the-box in React Native ExecuTorch see: TTS Models.
- For all supported voices in
useTextToSpeechplease refer to: Supported Voices
High Level Overview
You can play the generated waveform in any way most suitable to you; however, in the snippet below we utilize the react-native-audio-api library to play synthesized speech.
Supertonic 3
import { models, useTextToSpeech } from 'react-native-executorch';
import { AudioContext } from 'react-native-audio-api';
const model = useTextToSpeech(models.text_to_speech.supertonic.m1());
const audioContext = new AudioContext({ sampleRate: 44100 });
const handleSpeech = async (text: string) => {
const waveform = await model.forward({
text,
totalSteps: 8,
lang: 'en',
});
const audioBuffer = audioContext.createBuffer(1, waveform.length, 44100);
audioBuffer.getChannelData(0).set(waveform);
const source = audioContext.createBufferSource();
source.buffer = audioBuffer;
source.connect(audioContext.destination);
source.start();
};
Kokoro
import { models, useTextToSpeech } from 'react-native-executorch';
import { AudioContext } from 'react-native-audio-api';
const model = useTextToSpeech(models.text_to_speech.kokoro.en_us.heart());
const audioContext = new AudioContext({ sampleRate: 24000 });
const handleSpeech = async (text: string) => {
const speed = 1.0;
const waveform = await model.forward({ text, speed });
const audioBuffer = audioContext.createBuffer(1, waveform.length, 24000);
audioBuffer.getChannelData(0).set(waveform);
const source = audioContext.createBufferSource();
source.buffer = audioBuffer;
source.connect(audioContext.destination);
source.start();
};
Arguments
useTextToSpeech takes TextToSpeechModelConfig that consists of:
modelof typeTextToSpeechModelSources— model configuration.voiceSourceof typeResourceSource— the voice tensor used for synthesis.phonemizerConfigof typeTextToSpeechPhonemizerConfig— Kokoro only: phonemizer configuration. Unused by Supertonic.langof typeTextToSpeechSupertonicLanguage— Supertonic only: default language token (e.g.'en','na'). Unused by Kokoro.
useTextToSpeech's second optional argument is an object with:
preventLoadwhich prevents auto-loading of the model.
You need more details? Check the following resources:
- For detailed information about
useTextToSpeecharguments check this section:useTextToSpeecharguments. - For all text to speech models available out-of-the-box in React Native ExecuTorch see: Text to Speech Models.
- For all supported voices in
useTextToSpeechplease refer to: Supported Voices - For more information on loading resources, take a look at loading models page.
Returns
useTextToSpeech returns an object called TextToSpeechType containing bunch of functions to interact with TTS. To get more details please read: TextToSpeechType API Reference.
Running the model
The module provides two ways to generate speech. The available parameters differ by model family:
| Parameter | Kokoro | Supertonic |
|---|---|---|
speed | ✓ | |
phonemize | ✓ | |
totalSteps | ✓ | |
lang | ✓ |
Using Text
forward({ text, speed, phonemize, totalSteps, lang }): Generates the complete audio waveform at once. Returns a promise resolving to aFloat32Array.stream({ speed, phonemize, totalSteps, lang, stopAutomatically, onNext, ... }): An async generator-like functionality (managed via callbacks likeonNext) that yields chunks of audio as they are computed. This is ideal for reducing the "time to first audio" for long sentences. You can also dynamically insert text during the generation process usingstreamInsert(text), force-partition trailing content without an end-of-sentence character viastreamFlush(), and stop the stream withstreamStop(instant).
In most cases, the stream() method is recommended over forward(). It significantly reduces latency by allowing audio playback to begin as soon as the first chunk is synthesized, rather than waiting for the entire text to be processed.
Both methods accept a phonemize parameter (defaults to true). This applies to Kokoro only — Supertonic maps text directly through a unicode indexer and does not use phonemization. When set to true, the input text is treated as raw text and converted to phonemes internally. When set to false, the input is expected to be a string of IPA phonemes.
Using Phonemes (Kokoro only)
If you have pre-computed phonemes (e.g., from an external dictionary or a custom G2P model), you can skip the internal phoneme generation step:
forward({ text, phonemize: false, speed }): Generates the complete audio waveform from a phoneme string.stream({ text, phonemize: false, speed, onNext, ... }): Streams audio chunks generated from a phoneme string.
Since forward and stream process the input, they might take a significant amount of time to produce audio for long inputs.
Example
Raw Synthesis (forward)
import React from 'react';
import { Button, View } from 'react-native';
import { models, useTextToSpeech } from 'react-native-executorch';
import { AudioContext } from 'react-native-audio-api';
export default function App() {
// Supertonic 3 — multilingual, any voice works for any language
const tts = useTextToSpeech(models.text_to_speech.supertonic.m1());
// Kokoro — language-specific voice bundle:
// const tts = useTextToSpeech(models.text_to_speech.kokoro.en_us.heart());
const generateAudio = async () => {
const audioData = await tts.forward({
text: 'Hello world! This is a sample text.',
totalSteps: 8,
lang: 'en',
// Kokoro: text only (no totalSteps/lang):
// text: 'Hello world! This is a sample text.',
});
// Playback — sample rate depends on the model
const ctx = new AudioContext({ sampleRate: 44100 });
// Kokoro: sampleRate: 24000
const buffer = ctx.createBuffer(1, audioData.length, ctx.sampleRate);
buffer.getChannelData(0).set(audioData);
const source = ctx.createBufferSource();
source.buffer = buffer;
source.connect(ctx.destination);
source.start();
};
return (
<View style={{ flex: 1, justifyContent: 'center', alignItems: 'center' }}>
<Button title="Speak" onPress={generateAudio} disabled={!tts.isReady} />
</View>
);
}
Streaming Synthesis
import React, { useRef } from 'react';
import { Button, View } from 'react-native';
import { models, useTextToSpeech } from 'react-native-executorch';
import { AudioContext } from 'react-native-audio-api';
export default function App() {
// Supertonic 3
const tts = useTextToSpeech(models.text_to_speech.supertonic.m1());
const contextRef = useRef(new AudioContext({ sampleRate: 44100 }));
const generateStream = async () => {
const ctx = contextRef.current;
await tts.stream({
text: "This is a longer text, which is being streamed chunk by chunk. Let's see how it works!",
totalSteps: 8,
lang: 'en',
onNext: async (chunk) => {
return new Promise((resolve) => {
const buffer = ctx.createBuffer(1, chunk.length, ctx.sampleRate);
buffer.getChannelData(0).set(chunk);
const source = ctx.createBufferSource();
source.buffer = buffer;
source.connect(ctx.destination);
source.onEnded = () => resolve();
source.start();
});
},
});
};
return (
<View style={{ flex: 1, justifyContent: 'center', alignItems: 'center' }}>
<Button title="Stream" onPress={generateStream} disabled={!tts.isReady} />
</View>
);
}
Supported models
| Model | Language |
|---|---|
| Supertonic 3 | 31 languages + na (unknown) — Arabic, Bulgarian, Czech, Danish, Dutch, English, Finnish, French, German, Greek, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Malay, Norwegian, Polish, Portuguese, Romanian, Russian, Slovak, Spanish, Swahili, Swedish, Tagalog, Tamil, Thai, Turkish, Vietnamese |
| Kokoro | English, French, German, Spanish, Portuguese, Italian, Polish, Hindi |