Capability briefMarkerless motion capture

Markerless motion capture.
Captured, not generated.

mimem turns multi-camera video into production-ready 3D character animation, with quality comparable to high-end optical systems — but markerless, no calibration, two minutes to set up.

See the evidence
  • Up to 12 cameras
  • No suit, no markers
  • No calibration
  • Body + hands

EvidenceIndependent reviews

Don't take our word for it.

“It's astounding how good the animations are, I barely have to clean up the animations — I'd say they look as if they are from optical mocap.

Victor Tan · Independent creator

Same footageTwo solvers

“Same footage… processed using Move Pro, the $7,000 per year system. And on the right we have Mimem, which is using just three of the six GoPros. They're both really good.”

Charlie Driscoll · YouTube creator, UE5 filmmaker

Stress testFeet & occlusion

“Most systems do have a major problem when you leave the floor… In this example, he sits down, which is already a difficult task, and lifts his feet. This blew me away.”

Nils Gallist · YouTube reviewer

“In terms of mocap quality I prefer mimem to Rokoko… I don't have time to manually fix up 1 hour+ of footage. For footlocking, mimem wins.”

Dylan · Filmmaker

Independent reviews — not sponsored

Markerless vs traditional mocap

Measured against the systems it displaces — and honest about where they still win

Feature
mimem.aiRecommended
Optical studio
Inertial suit
Single-camera AI
Cameras per take
Up to 1212–50+None (sensors)1
No suit, no markers
YesNoNoYes
Calibration
Solved from the takeWand calibrationPose calibrationNone
Capture volume
Any room, outdoorsDedicated stageAnywhereAny room
Wardrobe
Street clothes or costumeMarker suitSuit under clothingStreet clothes
Hands & fingers
YesAdd-on ($)Gloves ($)No
Performers per take
Up to 3ManyOne suit eachUsually 1
Foot contact
Firmly plantedReference-gradeDrifts on long takesSome sliding
Occlusion handling
Covered by other anglesCovered by other anglesUnaffectedEstimated
Setup time
2 minutesCalibrated install15–30 min suit-up2 minutes
Starting price
Free
hardware €0 – ~€1,800
$50,000+$2,500+Freemium

Pricing out the alternatives? The motion capture cost hub tracks what every system on this table actually costs — and if you're weighing an inertial suit, the mocap suit guide covers when one is worth it.

When would you still pick a traditional system?

Vicon / OptiTrack

Sub-millimeter precision for fast, complex body deformation that AI pose estimation still struggles with — acrobatics, contortion, or biomechanics research.

Xsens

Works where you can't put cameras. Inertial suits don't need line of sight or a capture volume — the sensors are on the body.

For everything else, markerless sets you free

Real human performance — any location, any camera, any wardrobe.

AI-Powered Pose Estimation

Deep learning models track the whole performance from video alone — full body, hands, fingers, and feet.

Multi-View Triangulation

Combine footage from 2-12 cameras. Our AI synchronizes views and triangulates true 3D positions.

No Calibration Required

Skip the hours of setup. Our AI handles camera synchronization and spatial alignment automatically.

Full Body + Hands + Feet

Track the entire body including individual finger joints and toe positions.

How It Works

Any cameras you own.

Phones, webcams, GoPros, DSLRs — mixed freely, on tripods around the action. No markers, no calibration, no genlock. Shooting on iPhones? There's a native app →

Upload every angle.

Every video goes straight into one session. Sync is solved from the footage itself — up to three seconds of start offset absorbed automatically.

One solve, all views.

Camera geometry and the performance are reconstructed together from the take. Full body, hands and fingers — feet planted, no cleanup pass.

Your character, your engine.

IK retargeting to any rig you upload — stylized proportions included — then standard FBX into every major DCC. See the full video-to-FBX pipeline →

MethodTriangulation, not inference

Triangulated, never guessed.

Every additional viewpoint does two jobs. It covers what another camera can't see — and it tightens the triangulation on everything the cameras share, so the fine detail of the performance survives into the solve: weight, hesitation, intention.

Some single-camera tools fill in what the lens never saw with a generative model. mimem never invents motion — if it's in your FBX, a camera saw it happen.

Where each camera stands matters as much as how many you use. Plan the rig in the camera placement guide — it ships with an interactive 3D simulator.

unseen
1 camera — the far half is inference
12 cameras — every angle measured

Footage in. FBX out.

A full session — record, solve, export — shot on iPhones running the companion app. The same pipeline runs on any footage you upload.

UnrealUnityBlenderMayaMotionBuilderCascadeur+ more

PricingPublic — no quotes

Pilot it this week.

No quote, no sales call, no seat licensing — every price is public. Start free with three cameras you already own, export FBX the same day, and put the result in front of your team.

Free

Pilot starts here

€0

50 tokens a day, up to 3 cameras, 3 FBX exports a month — full commercial rights

Essential

€25/mo

500 tokens a day, priority processing, 40 FBX exports a month

Plus

€60/mo

1,500 tokens a day, up to 6 cameras, unlimited FBX exports

Pro

€199/mo

5,000 tokens a day, up to 12 cameras, unlimited FBX exports

Allowances renew daily · Cancel anytimeHardware: €0 with your own cameras — about €1,800 for a 12-camera rig of refurbished iPhone 11 units
Full pricing

Frequently Asked Questions

Markerless motion capture uses computer vision and AI to track human movement from video, without requiring the performer to wear special markers, suits, or sensors.

Traditional motion capture systems require either:

  • Optical systems: Reflective markers on a suit, tracked by infrared cameras
  • Inertial systems: Sensor suits with accelerometers and gyroscopes

Markerless systems like mimem.ai eliminate this requirement entirely, using AI to detect body pose directly from regular video footage.

Published validations of multi-camera markerless systems report joint-angle errors within a few degrees of marker-based reference for everyday movement, and the gap narrows every year. Geometry does most of the work: multiple views triangulate true 3D positions instead of inferring them from a single image.

The honest limits: transverse-plane rotations and extreme poses — contortion, acrobatics, heavy self-occlusion — remain the hardest cases for any vision-based system, and benefit from more cameras.

For production animation, independent reviewers comparing the output to optical mocap footage put it plainly: “they look as if they are from optical mocap.”

It depends on your accuracy requirements:

  • 1 camera: Good for simple movements, walking, gestures
  • 2-3 cameras: Great for most use cases, handles some occlusion
  • 4-6 cameras: Professional quality, handles complex movements
  • 6-12 cameras: Maximum accuracy for challenging performances

Placement matters as much as count — our camera placement guide shows where to put them, with an interactive 3D simulator. Our free plan supports up to 3 cameras. Pro plan supports up to 12. See our pricing page for details.

Any camera that records video works with mimem.ai: smartphones (iPhone, Android), webcams, GoPros, DSLRs, mirrorless cameras, cinema cameras, or security cameras — mixed freely in the same take. Higher resolution and frame rate produce better results, but even basic webcams can capture usable motion data.

Shooting on iPhones? The native companion app records 4K 60fps on device and synchronizes up to 12 phones automatically.

No. Start each camera whenever you like — mimem aligns the views from the motion in the footage itself, absorbing up to about three seconds of start-time offset automatically. No timecode, no clapperboard, no sync cables.

Yes — up to 3 performers per take. Everyone is solved from the same footage, so actors can play a scene together instead of being captured separately and assembled later. We recommend at least 6 cameras for multi-performer takes: more bodies on stage means more occlusion to cover.

Performers can wear almost anything. For best results:

  • Fitted clothing works better than very loose or flowing garments
  • Contrasting colors against the background help visibility
  • Avoid very dark clothing in dark environments
  • Costumes are fine—capture in character wardrobe when needed

Unlike marker-based systems, there's no need for tight lycra suits or marker placement.

Markerless mocap does need good lighting and clear visibility of the subject—but unlike optical systems you don't need an infrared rig or dedicated studio. A well-lit living room, garage, or outdoor area with natural daylight works well. Inertial suits (Rokoko, Xsens) can work in the dark, but they accumulate drift on longer takes.

mimem.ai currently processes recorded video rather than streaming real-time capture. Processing typically takes 2-5 minutes for a 30-second clip. For real-time applications, we recommend capturing video and processing in batches. Real-time streaming is on our roadmap for future releases.

Real motion.
Really captured.

From the cameras you already own.

View pricing
A 3D character with arms raised, animated from a captured performance