How On-Device Pose Estimation Works (And Why It Stays On Your Phone)
本機辨識
A camera pointed at your body is about the most sensitive sensor there is. Here is what the app does with those frames, and where they go.
A camera pointed at your body in your own home is about the most sensitive sensor there is. Here is exactly what the app does with those frames, where they go, and what the architecture does and does not guarantee.
What pose estimation is
A pose estimation model takes an image and returns the estimated positions of a set of body landmarks — typically around thirty-three points covering head, shoulders, elbows, wrists, hips, knees, ankles and feet. Each comes with x and y coordinates and a confidence value.
The critical property for privacy purposes: the output is a list of numbers. Once a frame has been processed, what remains is a set of coordinates. The picture is not part of the result.
- 形xíng
- Form, shapeWhat the model can see — the external geometry of a posture.
- 意yì
- Intent, attentionWhat the model cannot see, and what most of the practice consists of.
The pipeline
- 01Camera permission
Nothing starts until you explicitly grant camera access, and only on the screens that use it. Browsers show their own indicator whenever a camera is live — an independent signal not controlled by the page.
- 02Frame capture
Frames are drawn to an in-memory canvas in the page. They are not written to disk and are not placed in any upload queue.
- 03Inference
MediaPipe Pose, running through TensorFlow.js with a WebGL backend, processes the frame using your device's GPU and returns landmark coordinates.
- 04Comparison
Those coordinates are compared against reference values for the current movement — angles and relative positions rather than raw pixels. See Reading a Posture Score.
- 05Feedback
A score and a written cue are rendered to the screen.
- 06Discard
The frame is replaced by the next one. Nothing accumulates.
How to verify it yourself
You should not take a privacy claim on trust from the party making it. This one is checkable in about a minute:
- Open your browser's developer tools and select the Network tab.
- Start a practice session with the camera on.
- Watch the outbound requests. You will see the model weights download once, at the start.
- Look for anything being uploaded during practice — repeated POST requests, large outbound payloads, WebSocket traffic carrying frame data.
- For a stronger test: put the browser in offline mode after the model has loaded. Pose guidance keeps working.
That last test is the decisive one. Something that continues to function with the network disconnected is not sending your frames anywhere.
What we are careful not to claim
It would be easy to write "we never collect any data". We do not say that, because it would not be true of the product as a whole.
| Component | Where it runs | What leaves your device |
|---|---|---|
| Pose estimation | Your browser | Nothing — coordinates stay local |
| Practice sessions offline | Your browser | Nothing |
| Account and sync | Server | Your email and practice history, if you create an account |
| Payments | Payment processor | Handled by the processor; card details never reach us |
| Model weights | Downloaded once | A standard file request |
The precise claim is narrower and defensible: camera frames are analysed on your device and do not need to be uploaded for posture guidance to function. Practise offline and nothing at all leaves.
Why it was built this way
Partly principle and partly practicality. Streaming video to a server for analysis would cost real money per user per minute, add latency that makes real-time feedback unusable, and require operating a system holding video of people exercising at home — a liability nobody sensible wants.
Local inference removes all three problems at once. It also happens to be the right answer, which is a pleasant coincidence rather than a moral achievement.