Skip to main content
Version: 0.10.0

Speech-to-Text (STT)

The Speech-to-Text extension transcribes spoken audio into text directly on-device using OpenAI's Whisper model paired with an integrated FSMN Voice Activity Detector (VAD).

The pipeline supports two primary workflows:

  • Live Microphone Streaming: Streams real-time audio straight from the microphone. As the user speaks, Whisper continuously returns draft transcripts and automatically commits finalized sentences upon pauses.
  • Pre-recorded Audio Transcription: Transcribes pre-recorded audio buffers or audio files in a single pass, with optional token-by-token streaming callbacks.
iOSAndroid

Quick Start

The useSpeechToText hook manages downloading model weights, the tokenizer, and the bundled VAD model. To capture live audio from the microphone, feed PCM chunks into streamInsert() using a microphone recorder such as react-native-audio-api:

import { useState } from 'react';
import { models, useSpeechToText, WHISPER_SAMPLE_RATE_HZ } from 'react-native-executorch';
import { useAudioRecorder } from './hooks/useAudioRecorder'; // Custom helper built with react-native-audio-api

function TranscriptionComponent() {
const [committedText, setCommittedText] = useState('');
const [nonCommittedText, setNonCommittedText] = useState('');

const stt = useSpeechToText(models.speechToText.WHISPER.EN.TINY.DEFAULT);
const recorder = useAudioRecorder();

// Hook state:
// stt.isReady — true once Whisper model, tokenizer, and VAD are loaded
// stt.downloadProgress — 0 to 100 download progress across all files
// stt.error — Error instance if download or load failed
// stt.resource — resolved config with all URLs replaced by local file paths

const handleToggleRecording = async () => {
if (recorder.isRecording) {
await recorder.stopRecording();
stt.streamStop?.(); // Signal stream to finalize and close
return;
}

if (!stt.isReady || !stt.stream || !stt.streamInsert) return;

setCommittedText('');
setNonCommittedText('');

// 1. Consume the live transcription stream in the background
(async () => {
const textStream = stt.stream!({ language: 'en' });
for await (const update of textStream) {
setCommittedText(update.committed);
setNonCommittedText(update.nonCommitted);
}
})();

// 2. Start microphone recording (16 kHz mono Float32 PCM)
await recorder.startRecording(WHISPER_SAMPLE_RATE_HZ, (samples) => {
stt.streamInsert?.(samples);
});
};
}
Full Interactive Example in Gallery App

See src/app/(screens)/speech-to-text.tsx in the React Native ExecuTorch Gallery for a complete, runnable screen featuring microphone controls, live audio streaming, and animated transcription UI.

Output Format & Live Streaming

When streaming live microphone audio with stream(), the generator yields transcription updates on every voice activity event:

type WhisperStreamUpdate = {
/** Finalized transcript of completed sentences and clauses */
readonly committed: string;
/** Live in-progress transcript of the active speech segment that may still update */
readonly nonCommitted: string;
};

How Live Streaming Works

  • Committed vs Non-Committed Text: As the user speaks, Whisper continuously transcribes the active speech window into nonCommitted text. Once the speaker pauses or completes a clause (detected by the integrated Voice Activity Detector), that segment is finalized and appended to committed text.
  • Background Audio Buffer: Audio chunks fed via streamInsert(pcmSamples) are accumulated in an internal audio ring buffer on a background thread without blocking the JavaScript UI.
  • Graceful Termination: Calling streamStop() signals the stream to process any remaining speech in the buffer, commit the final clause, and close the generator.

Pre-Recorded Audio Transcription

To transcribe an existing audio recording or batch audio buffer all at once, use transcribe():

// audioData: Float32Array PCM samples at 16000 Hz
const transcript = await stt.transcribe(audioData, {
language: 'en',
});

console.log('Full transcript:', transcript);

You can also pass an optional onToken callback to receive decoded word/subword tokens in real time as Whisper generates them:

const transcript = await stt.transcribe(audioData, { language: 'en' }, (token) => {
console.log('Decoded token:', token);
});

To abort an in-flight transcription prematurely, call transcribeStop():

// Cancels active transcribe() execution and rejects the pending promise
stt.transcribeStop?.();

Imperative API

For background services, offline audio processors, or non-React component logic, instantiate the pipeline imperatively using createWhisperSpeechToText:

import { createWhisperSpeechToText, download, models } from 'react-native-executorch';

// Download and cache Whisper weights, tokenizer, and bundled VAD
const model = await download(models.speechToText.WHISPER.EN.BASE.DEFAULT);
const stt = await createWhisperSpeechToText(model);

try {
const transcript = await stt.transcribe(audioData, { language: 'en' });
console.log('Transcript:', transcript);
} finally {
// Always release native resources when finished
stt.dispose();
}

Synchronous Execution

For synchronous worklet execution contexts or frame-by-frame audio processors, createWhisperSpeechToText exposes a synchronous transcribeWorklet function:

// Called synchronously inside a worklet runtime without Promise scheduling overhead
const transcript = stt.transcribeWorklet(audioData, { language: 'en' });

See Worklets & Threading for details on worklet execution contexts and zero-copy host objects.

Available Models

The library provides ready-to-use Whisper models from the Software Mansion HuggingFace Whisper Collection, available in models.speechToText:

Model FamilyVariantsSize RangeSupported BackendsLanguagesNotes
Whisper TinyMultilingual, English57.1 MB – 221.8 MBXNNPACK (CPU), Core ML (Apple), MLX (Apple), Vulkan (Android)English / WHISPER_LANGUAGES (99+ languages)Ultra-fast transcription with minimal RAM usage.
Whisper BaseMultilingual, English97.8 MB – 380.2 MBXNNPACK (CPU), Core ML (Apple), MLX (Apple), Vulkan (Android)English / WHISPER_LANGUAGES (99+ languages)Balanced accuracy and speed for general voice dictation.
Whisper SmallMultilingual, English276.0 MB – 1.05 GBXNNPACK (CPU), Core ML (Apple), MLX (Apple), Vulkan (Android)English / WHISPER_LANGUAGES (99+ languages)High-capacity model for complex, noisy, or multi-speaker audio.

API Reference

Hooks & Pipelines

Types & Options

Constants & Model Presets

Source Code

View the implementation on GitHub: