Skip to main content
Version: 0.10.0

Voice Activity Detection (VAD)

The Voice Activity Detection extension detects speech presence and segments audio into spoken and silent intervals directly on-device using a lightweight Feedforward Sequential Memory Network (FSMN-VAD) model.

The pipeline supports two primary workflows:

  • Live Microphone Event Stream: Processes incoming audio chunks from a microphone recorder in real time, firing 'speechStart' and 'speechEnd' transitions.
  • Batch Audio Segmentation: Analyzes an entire recorded audio buffer and returns an array of timestamped speech segments in seconds.
iOSAndroid

Quick Start

The useVoiceActivityDetector hook manages downloading the model weights and provides live streaming methods. To capture live audio, stream PCM chunks from a microphone recorder such as react-native-audio-api directly into detectVoiceOnStream():

import { useState } from 'react';
import { FSMN_VAD_SAMPLE_RATE_HZ, models, useVoiceActivityDetector } from 'react-native-executorch';
import { useAudioRecorder } from './hooks/useAudioRecorder'; // Custom helper built with react-native-audio-api

function VadComponent() {
const [isSpeaking, setIsSpeaking] = useState(false);

const vad = useVoiceActivityDetector(models.voiceActivityDetection.FSMN_VAD.DEFAULT);
const recorder = useAudioRecorder();

// Hook state:
// vad.isReady — true once FSMN-VAD model is loaded in memory
// vad.downloadProgress — 0 to 100 download progress
// vad.error — Error instance if download or load failed
// vad.resource — resolved config with all URLs replaced by local file paths

const handleToggleStreaming = async () => {
if (recorder.isRecording) {
await recorder.stopRecording();
vad.resetStream?.();
setIsSpeaking(false);
return;
}

if (!vad.isReady || !vad.detectVoiceOnStream || !vad.resetStream) return;

vad.resetStream(); // Clear internal rolling buffer
setIsSpeaking(false);

// Stream live microphone PCM chunks (16 kHz mono Float32)
await recorder.startRecording(
FSMN_VAD_SAMPLE_RATE_HZ,
(samples) => {
const event = vad.detectVoiceOnStream!(samples, { detectionMargin: 300 });
if (event === 'speechStart') {
setIsSpeaking(true);
} else if (event === 'speechEnd') {
setIsSpeaking(false);
}
},
1600 // ~100 ms chunk size
);
};
}
Full Interactive Example in Gallery App

See src/app/(screens)/voice-activity-detection.tsx in the React Native ExecuTorch Gallery for a complete, runnable screen featuring microphone controls, real-time speech indicators, and live audio streaming.

Live Microphone Streaming

detectVoiceOnStream() appends incoming audio samples to an internal 2.5-second bounded rolling window and runs fast inference (taking ~2–5 ms).

Output Event Type

detectVoiceOnStream() returns a VadEvent on transition states, or undefined when the voice activity state hasn't changed:

type VadEvent = 'speechStart' | 'speechEnd' | undefined;
  • 'speechStart': Fired when speech probability stays above speechThreshold for at least minSpeechDurationMs (default: 250 ms).
  • 'speechEnd': Fired when speech ceases and remains silent for at least minSilenceDurationMs (default: 220 ms).
  • undefined: Fired on regular frames when no transition boundary has occurred.

Before starting a new recording stream, call resetStream() to clear past audio history from the rolling buffer.

Batch Audio Segmentation

To process a pre-recorded audio buffer all at once, call detectVoice():

// audioData: Float32Array PCM samples at 16000 Hz
const segments = await vad.detectVoice(audioData);

for (const segment of segments) {
console.log(`Speech detected from ${segment.start.toFixed(2)}s to ${segment.end.toFixed(2)}s`);
}

Each VadSegment contains timestamps in seconds:

type VadSegment = {
/** Start time of the speech segment in seconds */
readonly start: number;
/** End time of the speech segment in seconds */
readonly end: number;
};

Detection Tuning & Options

You can customize threshold parameters per call by passing VadOptions:

const customSegments = await vad.detectVoice(audioData, {
speechThreshold: 0.5, // Minimum probability threshold (0.0 to 1.0, default: 0.5)
minSpeechDurationMs: 250, // Min continuous speech duration to open a segment (default: 250 ms)
minSilenceDurationMs: 220, // Min silence duration to close a segment (default: 220 ms)
speechPadMs: 300, // Padding added before/after detected speech (default: 300 ms)
mergeGapMs: 400, // Gap below which adjacent segments are merged (default: 400 ms)
});

Imperative API

For background tasks, offline audio preprocessing, or non-React component logic, instantiate the pipeline imperatively using createFsmnVoiceActivityDetector:

import { createFsmnVoiceActivityDetector, download, models } from 'react-native-executorch';

// Download and cache FSMN-VAD weights
const model = await download(models.voiceActivityDetection.FSMN_VAD.DEFAULT);
const detector = await createFsmnVoiceActivityDetector(model);

try {
const segments = await detector.detectVoice(audioData);
console.log('Detected segments:', segments);
} finally {
// Always release native model memory when done
detector.dispose();
}

Synchronous Execution

For synchronous worklet execution contexts or frame-by-frame audio processors, createFsmnVoiceActivityDetector exposes a synchronous detectVoiceWorklet function:

// Called synchronously inside a worklet runtime without Promise scheduling overhead
const segments = detector.detectVoiceWorklet(audioData);

See Worklets & Threading for details on worklet execution contexts and zero-copy host objects.

Available Models

The library provides the optimized FSMN-VAD model from the Software Mansion HuggingFace Voice Activity Detection Collection, available in models.voiceActivityDetection:

ModelVariantsSize RangeSample RateSupported BackendsNotes
FSMN-VADSee1.8 MB16000 HzXNNPACK (CPU)Compact, low-latency Feedforward Sequential Memory Network for continuous speech detection.

API Reference

Hooks & Pipelines

Types & Options

Constants & Model Presets

Source Code