Work / Simulation & Synthetic Data

Porch Watch: testing a person detector with synthetic data

Isaac Sim, OpenUSD and seven rigged MetaHumans. Same 2,000 scenes, only the person changed: daytime recall went from 84% to 99.9%.

Two identical porch-camera frames side by side; a yellow mannequin with no detection box on the left, a man in a red and black jacket with a blue person 0.90 detection box on the right.
Same scene, same camera, same seed. Left: the game-engine mannequin, missed. Right: a MetaHuman, detected (0.90).

Short versionThe test was wrong, not the detector

I built a simulated front porch, filmed it with a virtual doorbell camera, and generated 2,000 labelled images of a person walking up to the door under randomized cameras, lighting and props. Then I used it to test an off-the-shelf person detector.

The first test said the detector missed 16% of people in broad daylight. The real problem was my test: the "person" was a game-engine mannequin that doesn't look human. I replaced it with seven realistic MetaHuman characters and re-ran the exact same 2,000 scenes, with only the person changed. Daytime recall went from 84% to 99.9%.

In simulation-based testing, the quality of the 3D assets is part of the measurement. A bad asset doesn't just look bad; it produces a wrong answer.
84% → 99.9%daytime recall, same 2,000 scenes, only the person changed
7MetaHumans rigged from Unreal into Isaac Sim
12 minper 2,000 labelled frames on one RTX 4090

BuiltFive parts

1. A SimReady asset set. Four porch props (package, doormat, planter, porch light) built in Python as OpenUSD and packaged as self-contained USDZ. Each has real-world scale, mass and colliders, physics materials, semantic labels, and two materials: a portable USD material plus an RTX material. All four pass NVIDIA's SimReady Prop-Robotics-Physx validation (56 automated requirements), run headless as a pipeline step, plus a physics drop test.

2. A layered scene. The porch scene is split into USD layers (environment, props, actor), so each can be swapped or randomized without touching the others. The walker is a USD variant set with eight characters.

3. Characters from Unreal to Isaac Sim. Seven MetaHuman-based characters baked from Unreal's Mutable customization system and exported to USD. Getting them to simulation quality was the rigging part of the job:

  • Body and head export as separate skinned meshes on the same 270-joint skeleton. Both are driven from one walk clip, so the neck seam never separates.
  • The hair is attached to the head joint with a rigid skin binding. I measured each character's scalp to fit it (about 1 cm of lift), because Mutable reshapes every head slightly.
  • I fixed the baked normal maps (range remap plus the DirectX-to-OpenGL green flip).
  • Walk speed is measured from the clip's planted-foot velocity, so the feet lock to the ground: planted feet move 5 cm/s while the body travels 3 m/s.
Seven men in varied clothes and builds walking toward the camera.
Seven MetaHuman walkers baked from Unreal's Mutable sample, rigged and exported to USD for Isaac Sim.

4. Randomized synthetic data at scale. Every walk randomizes the camera (mount height, tilt, yaw, field of view), the walker (which character, and the path), the lighting (time of day, with a 25% chance of night under porch light only) and the props (package present or absent, and where). Each frame saves the image plus pixel-perfect labels: person bounding box, semantic segmentation, distance to the camera and camera parameters. A manifest records every random setting, so any result traces back to its exact cause.

Three panels: a rendered porch scene, a colour-coded segmentation mask with the person in pink, and a greyscale depth image.
One frame's free ground truth: RGB, per-pixel class labels, and distance to camera.

5. Evaluation. An off-the-shelf YOLO11n person detector (the small size class that runs on edge devices like doorbells) is scored against the ground truth, frame by frame.

ResultsA paired before-and-after

The same seeds reproduce the same camera, lighting, path and props, so the two runs are a paired comparison: 2,000 matched frames where only the person differs. The detector itself never changed.

Mannequin7 MetaHumans
Recall, daytime83.9%99.9%
Recall, night71.9%74.7%
Recall, overall81.3%94.5%
Frames that changed297 miss → hit34 hit → miss
Line charts: daytime miss rate falls from up to 27% to near 0% at all distances; night lines nearly overlap; bar chart of night miss rate per character.
Miss rate by distance, day and night, before (mannequin) and after (MetaHumans).
  • The daytime misses were the mannequin, not the detector. With realistic people, daytime misses drop to near zero at every distance out to 10 m.
  • Night barely moved, which points to lighting: the porch light barely reaches the walkway. Real doorbells switch to infrared at night, so the next step is an infrared camera model, not a better detector.
  • At night, dark clothing looks harder to detect (44% misses for a black T-shirt versus 7% for a bright top). That's a hypothesis to test, not a result: each character only had 4–10 night walks.
Dark porch scene with a dimly lit person and a detection box.
Night: porch light only.
Matched frames: mannequin missed, MetaHuman detected.
Another matched pair at 3.3 m: mannequin missed, MetaHuman detected.

QAThe bugs the pipeline caught

Most of the engineering went into making the data trustworthy. A few of the problems my automated checks caught:

  • Capture lag. The first ~20 frames of a run lagged the simulation by a frame (duplicates, with the walker not yet reset). A check that requires the person's box to grow as they approach caught it; now two throwaway warm-up walks run first.
  • Partial labels. In 21 of 200 walks, only the character's hair carried the "person" label after a character swap, so the box was head-sized. The run looked fine; the QA flagged it. Every mesh is now labelled directly, and a second check compares box size to distance.
  • Buried hair. The beard shells sat behind the face on every character, because they were fitted to a different MetaHuman. I measured it, dropped them, and logged it.
  • Validator false confidence. The SimReady validator reports PASS with zero rules if its rule modules aren't imported when it runs headless. My runner fails loudly instead.

Every issue goes in a log (33 entries so far) with symptom, cause, fix and how it was verified.

LimitsWhat I'd do next

  • Demographics: all seven walkers are men with light-to-medium skin. Before claiming coverage, I'd add darker skin tones and women.
  • Camera: an infrared/low-light model for night, and a fisheye lens (doorbells use 150°+ lenses).
  • Motion: one walk clip today; next is mocap variety (running, carrying packages, groups of people).
  • Metrics: recall only so far. Next: precision/AP curves, and a check against a small set of real doorbell images to measure the sim-to-real gap.
  • Production rig: merge body and head into a single skeleton, so retargeting and interaction rigs have one thing to drive.

WhyThe same craft, a new question

I've spent 16+ years making characters behave correctly inside game engines (Bungie, Microsoft, Counterplay, Rascal Games). This project points that craft at a new question: not "does it look right?" but "does the robot see it right, and how do we know?" The rigging, pipeline and validation habits carry over directly. The new part is treating every asset as part of a measurement.

Tools: Isaac Sim 6.1, OpenUSD, Omniverse Replicator, Unreal Engine 5.6/5.7 (Mutable, MetaHuman), Python, NVIDIA SimReady validator, ONNX Runtime. Characters: Epic Games Mutable Sample / MetaHuman. Detector: Ultralytics YOLO11n.