Work / Simulation & Synthetic Data
Porch Watch: testing a person detector with synthetic data
Isaac Sim, OpenUSD and seven rigged MetaHumans. Same 2,000 scenes, only the person changed: daytime recall went from 84% to 99.9%.

Short versionThe test was wrong, not the detector
I built a simulated front porch, filmed it with a virtual doorbell camera, and generated 2,000 labelled images of a person walking up to the door under randomized cameras, lighting and props. Then I used it to test an off-the-shelf person detector.
The first test said the detector missed 16% of people in broad daylight. The real problem was my test: the "person" was a game-engine mannequin that doesn't look human. I replaced it with seven realistic MetaHuman characters and re-ran the exact same 2,000 scenes, with only the person changed. Daytime recall went from 84% to 99.9%.
In simulation-based testing, the quality of the 3D assets is part of the measurement. A bad asset doesn't just look bad; it produces a wrong answer.
BuiltFive parts
1. A SimReady asset set. Four porch props (package, doormat, planter, porch light) built in Python as OpenUSD and packaged as self-contained USDZ. Each has real-world scale, mass and colliders, physics materials, semantic labels, and two materials: a portable USD material plus an RTX material. All four pass NVIDIA's SimReady Prop-Robotics-Physx validation (56 automated requirements), run headless as a pipeline step, plus a physics drop test.
2. A layered scene. The porch scene is split into USD layers (environment, props, actor), so each can be swapped or randomized without touching the others. The walker is a USD variant set with eight characters.
3. Characters from Unreal to Isaac Sim. Seven MetaHuman-based characters baked from Unreal's Mutable customization system and exported to USD. Getting them to simulation quality was the rigging part of the job:
- Body and head export as separate skinned meshes on the same 270-joint skeleton. Both are driven from one walk clip, so the neck seam never separates.
- The hair is attached to the head joint with a rigid skin binding. I measured each character's scalp to fit it (about 1 cm of lift), because Mutable reshapes every head slightly.
- I fixed the baked normal maps (range remap plus the DirectX-to-OpenGL green flip).
- Walk speed is measured from the clip's planted-foot velocity, so the feet lock to the ground: planted feet move 5 cm/s while the body travels 3 m/s.

4. Randomized synthetic data at scale. Every walk randomizes the camera (mount height, tilt, yaw, field of view), the walker (which character, and the path), the lighting (time of day, with a 25% chance of night under porch light only) and the props (package present or absent, and where). Each frame saves the image plus pixel-perfect labels: person bounding box, semantic segmentation, distance to the camera and camera parameters. A manifest records every random setting, so any result traces back to its exact cause.

5. Evaluation. An off-the-shelf YOLO11n person detector (the small size class that runs on edge devices like doorbells) is scored against the ground truth, frame by frame.
ResultsA paired before-and-after
The same seeds reproduce the same camera, lighting, path and props, so the two runs are a paired comparison: 2,000 matched frames where only the person differs. The detector itself never changed.
| Mannequin | 7 MetaHumans | |
|---|---|---|
| Recall, daytime | 83.9% | 99.9% |
| Recall, night | 71.9% | 74.7% |
| Recall, overall | 81.3% | 94.5% |
| Frames that changed | 297 miss → hit | 34 hit → miss |

- The daytime misses were the mannequin, not the detector. With realistic people, daytime misses drop to near zero at every distance out to 10 m.
- Night barely moved, which points to lighting: the porch light barely reaches the walkway. Real doorbells switch to infrared at night, so the next step is an infrared camera model, not a better detector.
- At night, dark clothing looks harder to detect (44% misses for a black T-shirt versus 7% for a bright top). That's a hypothesis to test, not a result: each character only had 4–10 night walks.


QAThe bugs the pipeline caught
Most of the engineering went into making the data trustworthy. A few of the problems my automated checks caught:
- Capture lag. The first ~20 frames of a run lagged the simulation by a frame (duplicates, with the walker not yet reset). A check that requires the person's box to grow as they approach caught it; now two throwaway warm-up walks run first.
- Partial labels. In 21 of 200 walks, only the character's hair carried the "person" label after a character swap, so the box was head-sized. The run looked fine; the QA flagged it. Every mesh is now labelled directly, and a second check compares box size to distance.
- Buried hair. The beard shells sat behind the face on every character, because they were fitted to a different MetaHuman. I measured it, dropped them, and logged it.
- Validator false confidence. The SimReady validator reports PASS with zero rules if its rule modules aren't imported when it runs headless. My runner fails loudly instead.
Every issue goes in a log (33 entries so far) with symptom, cause, fix and how it was verified.
LimitsWhat I'd do next
- Demographics: all seven walkers are men with light-to-medium skin. Before claiming coverage, I'd add darker skin tones and women.
- Camera: an infrared/low-light model for night, and a fisheye lens (doorbells use 150°+ lenses).
- Motion: one walk clip today; next is mocap variety (running, carrying packages, groups of people).
- Metrics: recall only so far. Next: precision/AP curves, and a check against a small set of real doorbell images to measure the sim-to-real gap.
- Production rig: merge body and head into a single skeleton, so retargeting and interaction rigs have one thing to drive.
WhyThe same craft, a new question
I've spent 16+ years making characters behave correctly inside game engines (Bungie, Microsoft, Counterplay, Rascal Games). This project points that craft at a new question: not "does it look right?" but "does the robot see it right, and how do we know?" The rigging, pipeline and validation habits carry over directly. The new part is treating every asset as part of a measurement.
Tools: Isaac Sim 6.1, OpenUSD, Omniverse Replicator, Unreal Engine 5.6/5.7 (Mutable, MetaHuman), Python, NVIDIA SimReady validator, ONNX Runtime. Characters: Epic Games Mutable Sample / MetaHuman. Detector: Ultralytics YOLO11n.