Pose & Keypoints
Pose estimation and keypoint detection locate specific anatomical landmarks on detected subjects — such as human skeletal joints (eyes, shoulders, elbows, wrists, hips, knees, ankles) or facial landmarks (eyes, nose tip, mouth, ears). Each prediction outputs a subject bounding box, detection confidence, and landmark coordinates scaled to the input image with individual landmark confidence scores.
Unlike basic object detection (which only returns box boundaries), keypoint detection tracks body posture, movement, and facial alignment. Common use cases include fitness/workout tracking, gesture controls, motion analysis, face alignment, and AR filters.
| iOS | Android |
|---|---|
Quick Start
The useKeypointDetector hook manages model downloading, initialization, and lifecycle:
import { models, useKeypointDetector } from 'react-native-executorch';
import type { ImageBuffer } from 'react-native-executorch/cv';
function MyComponent() {
const detector = useKeypointDetector(models.keypointDetection.YOLO26_POSE.DEFAULT);
// Hook state:
// detector.isReady — true once model is downloaded and loaded in memory
// detector.downloadProgress — 0 to 100 download progress
// detector.error — Error instance if download or load failed
// detector.resource — resolved config with all URLs replaced by local file paths
const handleDetect = async (imageBuffer: ImageBuffer) => {
if (!detector.isReady || !detector.detectKeypoints) return;
// Run inference on background thread
const detections = await detector.detectKeypoints(imageBuffer, {
confidenceThreshold: 0.25,
iouThreshold: 0.7,
});
console.log('Detected poses:', detections);
};
// Trigger handleDetect from an image picker, button press, or camera frame
}
See src/app/(screens)/keypoint-detection.tsx in the React Native ExecuTorch Gallery for a complete, runnable screen with photo picker, skeleton keypoint overlays, and latency tracking.
Output Format
detectKeypoints() returns an array of KeypointDetection objects:
type KeypointDetection<F extends BoxFormat = 'xyxy', L extends PropertyKey = string> = {
/** Scaled bounding box coordinates matching the input image resolution */
readonly box: BoundingBox<F>;
/** Overall detection confidence score (between 0.0 and 1.0) */
readonly confidence: number;
/** Map of landmark names to their pixel coordinates and confidence scores */
readonly landmarks: Record<L, { x: number; y: number; confidence: number }>;
};
For human pose models (YOLO26_POSE), landmarks includes 17 COCO_LANDMARKS body points:
[
{
box: { format: 'xyxy', xmin: 45.2, ymin: 12.0, xmax: 310.5, ymax: 580.0 },
confidence: 0.93,
landmarks: {
nose: { x: 178.4, y: 85.2, confidence: 0.97 },
leftEye: { x: 190.1, y: 75.4, confidence: 0.95 },
rightEye: { x: 165.8, y: 76.0, confidence: 0.94 },
leftEar: { x: 205.3, y: 80.1, confidence: 0.91 },
rightEar: { x: 150.2, y: 81.0, confidence: 0.9 },
// ... 12 more COCO landmarks (shoulders → ankles)
},
},
];
For face models (BLAZEFACE), landmarks includes 6 facial points from BLAZEFACE_LANDMARKS: leftEye, rightEye, noseTip, mouthCenter, leftEar, rightEar.
Models that regress depth add a depth to each landmark, on the same scale as x and negative towards the camera. It is undefined for the flat models.
Face Mesh
FACEMESH is the dense counterpart to BlazeFace: 468 3-D vertices over a single face, keyed by their index in MediaPipe's canonical face model (see FACEMESH_LANDMARKS), so any lip/eye/oval index list published for MediaPipe Face Mesh applies unchanged.
It differs from the other models on this page in two ways. It does not search an image for faces — it expects one already-cropped, face-filling image and always returns exactly one result, whose box is the hull of the mesh rather than a detection. And its confidence is a face-presence score for the whole crop, while each landmark's own confidence is a constant 1. Feeding it a whole photo rather than a crop drops that score below the default 0.5 threshold on most faces, even ones filling the frame, and detectKeypoints then returns an empty array.
import { createKeypointDetector, download, models } from 'react-native-executorch';
const detector = await createKeypointDetector(
await download(models.keypointDetection.BLAZEFACE.DEFAULT)
);
const mesh = await createKeypointDetector(
await download(models.keypointDetection.FACEMESH.DEFAULT)
);
const [face] = await detector.detectKeypoints(imageBuffer);
if (face) {
const crop = cropToFace(imageBuffer, face); // see below
const [result] = await mesh.detectKeypoints(crop.buffer);
const point = result?.landmarks[33]; // { x, y, confidence: 1, depth }, in crop pixels
if (point) console.log(crop.toSource(point.x, point.y)); // in imageBuffer pixels
}
Like MediaPipe, crop a square 1.5x the detector's box and rotate it so the eyes are level: the mesh expects an upright face, and an axis-aligned crop of a tilted head drifts ~5 px at 30 degrees and loses the face near 90. An ImageBuffer is plain HWC bytes, so this is a resample loop:
import type { BlazeFaceLandmark, KeypointDetection } from 'react-native-executorch';
import type { ImageBuffer } from 'react-native-executorch/cv';
const SIZE = 192; // the mesh's input side
function cropToFace(src: ImageBuffer, face: KeypointDetection<'xyxy', BlazeFaceLandmark>) {
const { box, landmarks } = face;
const { leftEye, rightEye } = landmarks;
const angle = Math.atan2(rightEye.y - leftEye.y, rightEye.x - leftEye.x);
const [cos, sin] = [Math.cos(angle), Math.sin(angle)];
const cx = (box.xmin + box.xmax) / 2;
const cy = (box.ymin + box.ymax) / 2;
const scale = (Math.max(box.xmax - box.xmin, box.ymax - box.ymin) * 1.5) / SIZE;
// Crop pixel -> source pixel. Also maps the mesh's landmarks back.
const toSource = (x: number, y: number) => {
const dx = (x - SIZE / 2) * scale;
const dy = (y - SIZE / 2) * scale;
return { x: cx + dx * cos - dy * sin, y: cy + dx * sin + dy * cos };
};
const channels = src.data.length / (src.width * src.height);
const data = new Uint8Array(SIZE * SIZE * channels);
for (let row = 0; row < SIZE; row++) {
for (let col = 0; col < SIZE; col++) {
const { x, y } = toSource(col + 0.5, row + 0.5);
const sx = Math.floor(x);
const sy = Math.floor(y);
if (sx < 0 || sy < 0 || sx >= src.width || sy >= src.height) continue;
const from = (sy * src.width + sx) * channels;
data.set(src.data.subarray(from, from + channels), (row * SIZE + col) * channels);
}
}
return { buffer: { ...src, data, width: SIZE, height: SIZE }, toSource, scale };
}
In the example above we use nearest-pixel sampling to keep it concise.
Configuration & Options
Pass a DetectKeypointsOptions object to detectKeypoints() to override model defaults:
| Option | Type | Default | Description |
|---|---|---|---|
confidenceThreshold | number | Model default (e.g. 0.25) | Minimum confidence score for a detected subject to be retained. |
iouThreshold | number | Model default (e.g. 0.7) | Non-Maximum Suppression (NMS) IoU overlap threshold. |
Imperative API
For background tasks, headless services, or manual lifecycle management outside React components, create the detector using createKeypointDetector:
import { createKeypointDetector, download, models } from 'react-native-executorch';
// Download and cache model assets before creating the pipeline
const model = await download(models.keypointDetection.YOLO26_POSE.DEFAULT);
const detector = await createKeypointDetector(model);
try {
const poses = await detector.detectKeypoints(imageBuffer, {
confidenceThreshold: 0.3,
});
console.log('Detected poses:', poses);
} finally {
// Always release native resources when finished
detector.dispose();
}
Synchronous Execution
For real-time camera tracking and live fitness apps, createKeypointDetector exposes a synchronous detectKeypointsWorklet function. This runs directly on the worklet thread with zero Promise scheduling overhead:
// Called synchronously inside a VisionCamera frame processor on the UI worklet thread
const poses = detector.detectKeypointsWorklet(frameBuffer, {
confidenceThreshold: 0.3,
});
See Worklets & Threading for details on worklet execution contexts and zero-copy host objects.
Available Models
The library provides ready-to-use pose and landmark detectors from the Software Mansion HuggingFace Pose Estimation Collection, available in models.keypointDetection:
| Model Family | Variants | Keypoints Detected | Size Range | Supported Backends | Notes |
|---|---|---|---|---|---|
| MediaPipe BlazeFace | See | BLAZEFACE_LANDMARKS (6 facial landmarks + box) | 0.6 MB | XNNPACK (CPU) | Ultra-lightweight face bounding box & eye/ear/nose/mouth keypoint tracking (sub-millisecond). |
| MediaPipe Face Mesh | See | FACEMESH_LANDMARKS (468 3-D mesh vertices) | 1.6 MB – 2.5 MB | XNNPACK (CPU), Core ML (Apple) | Dense single-face mesh from a 192x192 crop; needs a face detector in front of it. |
| YOLO26 Pose | See | COCO_LANDMARKS (17 body keypoints) | 11.4 MB | XNNPACK (CPU), Core ML (Apple) | Real-time multi-person full-body skeletal tracking across multiple input resolutions. |
| RF-DETR Keypoint | See | COCO_LANDMARKS (17 body keypoints) | 138.6 MB – 140.9 MB | XNNPACK (CPU), Core ML (Apple) | High-accuracy body keypoint detection transformer for complex, occluded poses. |
To use your own fine-tuned pose or landmark detection .pte model, pass a KeypointDetectorModel configuration object to useKeypointDetector or createKeypointDetector:
const customDetector = await createKeypointDetector({
modelPath: 'https://example.com/my-pose-model.pte',
modelOpts: {
landmarks: ['head', 'leftHand', 'rightHand'],
boxFormat: 'xyxy',
resizeMode: 'letterbox',
interpolation: 'linear',
normalizeOpts: { alpha: 1 / 255.0, beta: 0.0 },
defaultConfidenceThreshold: 0.3,
defaultIouThreshold: 0.6,
},
});
The pipeline automatically verifies that the model's exported input and output shapes match its requirements. To prepare and export your own .pte model to match this pipeline, see Exporting Custom Models.
API Reference
Hooks & Pipelines
useKeypointDetector()— React hook for keypoint detector downloading, state, and lifecycle.createKeypointDetector()— Imperative factory for keypoint and pose detection pipelines.
Types & Options
KeypointDetector— Keypoint detector runner interface (detectKeypoints,detectKeypointsWorklet).KeypointDetection— Detection result structure containingbox,confidence, andlandmarks.DetectKeypointsOptions— Detection options (confidenceThreshold,iouThreshold).KeypointDetectorModel— Model configuration spec for pose and landmark models.KeypointDetectorOptions— Options defining landmark names, box format, and normalization.Landmarks— Record of landmark names mapped to aLandmark.Landmark— One landmark:x,y,confidence, andzfor models that regress depth.BoundingBox— Bounding box structure.ImageBuffer— Input image buffer structure.
Model Presets & Constants
models.keypointDetection— Pre-configured keypoint and pose models registry.COCO_LANDMARKS— List of 17 standard COCO skeletal body keypoints.BLAZEFACE_LANDMARKS— List of 6 standard BlazeFace facial landmarks.FACEMESH_LANDMARKS— The 468 Face Mesh vertex indices.
View the implementation on GitHub: