Axon VR TrainingComfyUI

AI Motion Capture Pipeline

Five CG-style tutorial videos teaching Axon customers how to handle the VR headset and Taser, produced in two weeks by one designer using live-action footage to drive ComfyUI generation.

Role
Direction, filming, pipeline, edit
Timeline
2 weeks
Tools
ComfyUI, Sony A7R V, After Effects
5
CG-style tutorial videos shipped
2 wks
from pilot to five finished videos
0
dedicated modelers or animators needed

The problem

Axon's VR training platform needed tutorial videos showing customers how to handle both the headset and the Taser. Police and security teams aren't especially technical, so the videos had to be clear, step-by-step demonstrations of real device handling.

An earlier tutorial video set the baseline: one person producing everything, with no modelers or animators scoped. On that timeline a 2D illustrated approach was the right call, but it limited how ambitious the videos could be.

Before: 2D illustrated tutorial, produced solo.

After: the same job with the ComfyUI pipeline.

The bet

By the time the next round of videos came up, I had been learning ComfyUI on my own time and sharing progress with the design team. I'd also spent years around game motion capture, and I had a hunch that real footage could give us the control AI video usually lacks.

I filmed a teammate against a backdrop using basic photography techniques, then used that live-action reference to drive the ComfyUI generation.

The live-action reference: a teammate filmed against a backdrop.

The pipeline

The workflow was simple: shoot the live-action footage, feed it to ComfyUI with a target image defining the character (a clean gray mannequin-style figure), then use point-prompt segmentation to separate character from background so the performer can be replaced with the generated character while keeping the original motion.

A scroll through the ComfyUI workflow: the reference footage in, the segmentation points, and the finished output.

Results

The results were consistent. Across all five videos the character, style, and proportions stayed locked shot to shot, without the drift and morphing common in generated footage. The headset also stayed on-model: an accurate Vive Focus 3, the device customers actually use, which matters when the video is teaching them to handle it.

Reflection

This was a first run-through, and it showed where the pipeline gets more efficient: better lighting on the reference footage, so the model reads the performance cleanly instead of filling gaps with visuals we didn't want, and the speed that comes with repetition.

These videos were generated with the WAN model; now that LTX supports audio, I'd love to try a workflow built on that. The video generation space moves fast, so who knows what's coming next.