Skip to main content
Version: Next

Pose & Keypoints

Pose estimation and keypoint detection locate specific anatomical landmarks on detected subjects — such as human skeletal joints (eyes, shoulders, elbows, wrists, hips, knees, ankles) or facial landmarks (eyes, nose tip, mouth, ears). Each prediction outputs a subject bounding box, detection confidence, and landmark coordinates scaled to the input image with individual landmark confidence scores.

Unlike basic object detection (which only returns box boundaries), keypoint detection tracks body posture, movement, and facial alignment. Common use cases include fitness/workout tracking, gesture controls, motion analysis, face alignment, and AR filters.

iOSAndroid

Quick Start​

The useKeypointDetector hook manages model downloading, initialization, and lifecycle:

import { models, useKeypointDetector } from 'react-native-executorch';
import type { ImageBuffer } from 'react-native-executorch/cv';

function MyComponent() {
const detector = useKeypointDetector(models.keypointDetection.YOLO26_POSE.DEFAULT);

// Hook state:
// detector.isReady — true once model is downloaded and loaded in memory
// detector.downloadProgress — 0 to 100 download progress
// detector.error — Error instance if download or load failed
// detector.resource — resolved config with all URLs replaced by local file paths

const handleDetect = async (imageBuffer: ImageBuffer) => {
if (!detector.isReady || !detector.detectKeypoints) return;

// Run inference on background thread
const detections = await detector.detectKeypoints(imageBuffer, {
confidenceThreshold: 0.25,
iouThreshold: 0.7,
});
console.log('Detected poses:', detections);
};

// Trigger handleDetect from an image picker, button press, or camera frame
}
Full Interactive Example in Gallery App

See src/app/(screens)/keypoint-detection.tsx in the React Native ExecuTorch Gallery for a complete, runnable screen with photo picker, skeleton keypoint overlays, and latency tracking.

Output Format​

detectKeypoints() returns an array of KeypointDetection objects:

type KeypointDetection<F extends BoxFormat = 'xyxy', L extends PropertyKey = string> = {
/** Scaled bounding box coordinates matching the input image resolution */
readonly box: BoundingBox<F>;
/** Overall detection confidence score (between 0.0 and 1.0) */
readonly confidence: number;
/** Map of landmark names to their pixel coordinates and confidence scores */
readonly landmarks: Record<L, { x: number; y: number; confidence: number }>;
};

For human pose models (YOLO26_POSE), landmarks includes 17 COCO_LANDMARKS body points:

[
{
box: { format: 'xyxy', xmin: 45.2, ymin: 12.0, xmax: 310.5, ymax: 580.0 },
confidence: 0.93,
landmarks: {
nose: { x: 178.4, y: 85.2, confidence: 0.97 },
leftEye: { x: 190.1, y: 75.4, confidence: 0.95 },
rightEye: { x: 165.8, y: 76.0, confidence: 0.94 },
leftEar: { x: 205.3, y: 80.1, confidence: 0.91 },
rightEar: { x: 150.2, y: 81.0, confidence: 0.9 },
// ... 12 more COCO landmarks (shoulders → ankles)
},
},
];

For face models (BLAZEFACE), landmarks includes 6 facial points from BLAZEFACE_LANDMARKS: leftEye, rightEye, noseTip, mouthCenter, leftEar, rightEar.

Models that regress depth add a depth to each landmark, on the same scale as x and negative towards the camera. It is undefined for the flat models.

Face Mesh​

FACEMESH is the dense counterpart to BlazeFace: 468 3-D vertices over a single face, keyed by their index in MediaPipe's canonical face model (see FACEMESH_LANDMARKS), so any lip/eye/oval index list published for MediaPipe Face Mesh applies unchanged.

It differs from the other models on this page in two ways. It does not search an image for faces — it expects one already-cropped, face-filling image and always returns exactly one result, whose box is the hull of the mesh rather than a detection. And its confidence is a face-presence score for the whole crop, while each landmark's own confidence is a constant 1. Feeding it a whole photo rather than a crop drops that score below the default 0.5 threshold on most faces, even ones filling the frame, and detectKeypoints then returns an empty array.

import { createKeypointDetector, download, models } from 'react-native-executorch';

const detector = await createKeypointDetector(
await download(models.keypointDetection.BLAZEFACE.DEFAULT)
);
const mesh = await createKeypointDetector(
await download(models.keypointDetection.FACEMESH.DEFAULT)
);

const [face] = await detector.detectKeypoints(imageBuffer);
if (face) {
const crop = cropToFace(imageBuffer, face); // see below
const [result] = await mesh.detectKeypoints(crop.buffer);
const point = result?.landmarks[33]; // { x, y, confidence: 1, depth }, in crop pixels
if (point) console.log(crop.toSource(point.x, point.y)); // in imageBuffer pixels
}

Like MediaPipe, crop a square 1.5x the detector's box and rotate it so the eyes are level: the mesh expects an upright face, and an axis-aligned crop of a tilted head drifts ~5 px at 30 degrees and loses the face near 90. An ImageBuffer is plain HWC bytes, so this is a resample loop:

import type { BlazeFaceLandmark, KeypointDetection } from 'react-native-executorch';
import type { ImageBuffer } from 'react-native-executorch/cv';

const SIZE = 192; // the mesh's input side

function cropToFace(src: ImageBuffer, face: KeypointDetection<'xyxy', BlazeFaceLandmark>) {
const { box, landmarks } = face;
const { leftEye, rightEye } = landmarks;
const angle = Math.atan2(rightEye.y - leftEye.y, rightEye.x - leftEye.x);
const [cos, sin] = [Math.cos(angle), Math.sin(angle)];
const cx = (box.xmin + box.xmax) / 2;
const cy = (box.ymin + box.ymax) / 2;
const scale = (Math.max(box.xmax - box.xmin, box.ymax - box.ymin) * 1.5) / SIZE;

// Crop pixel -> source pixel. Also maps the mesh's landmarks back.
const toSource = (x: number, y: number) => {
const dx = (x - SIZE / 2) * scale;
const dy = (y - SIZE / 2) * scale;
return { x: cx + dx * cos - dy * sin, y: cy + dx * sin + dy * cos };
};

const channels = src.data.length / (src.width * src.height);
const data = new Uint8Array(SIZE * SIZE * channels);
for (let row = 0; row < SIZE; row++) {
for (let col = 0; col < SIZE; col++) {
const { x, y } = toSource(col + 0.5, row + 0.5);
const sx = Math.floor(x);
const sy = Math.floor(y);
if (sx < 0 || sy < 0 || sx >= src.width || sy >= src.height) continue;
const from = (sy * src.width + sx) * channels;
data.set(src.data.subarray(from, from + channels), (row * SIZE + col) * channels);
}
}
return { buffer: { ...src, data, width: SIZE, height: SIZE }, toSource, scale };
}

In the example above we use nearest-pixel sampling to keep it concise.

Configuration & Options​

Pass a DetectKeypointsOptions object to detectKeypoints() to override model defaults:

OptionTypeDefaultDescription
confidenceThresholdnumberModel default (e.g. 0.25)Minimum confidence score for a detected subject to be retained.
iouThresholdnumberModel default (e.g. 0.7)Non-Maximum Suppression (NMS) IoU overlap threshold.

Imperative API​

For background tasks, headless services, or manual lifecycle management outside React components, create the detector using createKeypointDetector:

import { createKeypointDetector, download, models } from 'react-native-executorch';

// Download and cache model assets before creating the pipeline
const model = await download(models.keypointDetection.YOLO26_POSE.DEFAULT);
const detector = await createKeypointDetector(model);

try {
const poses = await detector.detectKeypoints(imageBuffer, {
confidenceThreshold: 0.3,
});
console.log('Detected poses:', poses);
} finally {
// Always release native resources when finished
detector.dispose();
}

Synchronous Execution​

For real-time camera tracking and live fitness apps, createKeypointDetector exposes a synchronous detectKeypointsWorklet function. This runs directly on the worklet thread with zero Promise scheduling overhead:

// Called synchronously inside a VisionCamera frame processor on the UI worklet thread
const poses = detector.detectKeypointsWorklet(frameBuffer, {
confidenceThreshold: 0.3,
});

See Worklets & Threading for details on worklet execution contexts and zero-copy host objects.

Available Models​

The library provides ready-to-use pose and landmark detectors from the Software Mansion HuggingFace Pose Estimation Collection, available in models.keypointDetection:

Model FamilyVariantsKeypoints DetectedSize RangeSupported BackendsNotes
MediaPipe BlazeFaceSeeBLAZEFACE_LANDMARKS (6 facial landmarks + box)0.6 MBXNNPACK (CPU)Ultra-lightweight face bounding box & eye/ear/nose/mouth keypoint tracking (sub-millisecond).
MediaPipe Face MeshSeeFACEMESH_LANDMARKS (468 3-D mesh vertices)1.6 MB – 2.5 MBXNNPACK (CPU), Core ML (Apple)Dense single-face mesh from a 192x192 crop; needs a face detector in front of it.
YOLO26 PoseSeeCOCO_LANDMARKS (17 body keypoints)11.4 MBXNNPACK (CPU), Core ML (Apple)Real-time multi-person full-body skeletal tracking across multiple input resolutions.
RF-DETR KeypointSeeCOCO_LANDMARKS (17 body keypoints)138.6 MB – 140.9 MBXNNPACK (CPU), Core ML (Apple)High-accuracy body keypoint detection transformer for complex, occluded poses.
Using Custom Models

To use your own fine-tuned pose or landmark detection .pte model, pass a KeypointDetectorModel configuration object to useKeypointDetector or createKeypointDetector:

const customDetector = await createKeypointDetector({
modelPath: 'https://example.com/my-pose-model.pte',
modelOpts: {
landmarks: ['head', 'leftHand', 'rightHand'],
boxFormat: 'xyxy',
resizeMode: 'letterbox',
interpolation: 'linear',
normalizeOpts: { alpha: 1 / 255.0, beta: 0.0 },
defaultConfidenceThreshold: 0.3,
defaultIouThreshold: 0.6,
},
});

The pipeline automatically verifies that the model's exported input and output shapes match its requirements. To prepare and export your own .pte model to match this pipeline, see Exporting Custom Models.

API Reference​

Hooks & Pipelines​

Types & Options​

Model Presets & Constants​

Source Code

View the implementation on GitHub: