The problem
Axon's VR training platform needed tutorial videos showing customers how to handle both the headset and the Taser. Police and security teams aren't especially technical, so the videos had to be clear, step-by-step demonstrations of real device handling.
An earlier tutorial video set the baseline: one person producing everything, with no modelers or animators scoped. On that timeline a 2D illustrated approach was the right call, but it limited how ambitious the videos could be.
Before: 2D illustrated tutorial, produced solo.
After: the same job with the ComfyUI pipeline.
The bet
By the time the next round of videos came up, I had been learning ComfyUI on my own time and sharing progress with the design team. I'd also spent years around game motion capture, and I had a hunch that real footage could give us the control AI video usually lacks.
I filmed a teammate against a backdrop using basic photography techniques, then used that live-action reference to drive the ComfyUI generation.
The live-action reference: a teammate filmed against a backdrop.
The pipeline
The workflow was simple: shoot the live-action footage, feed it to ComfyUI with a target image defining the character (a clean gray mannequin-style figure), then use point-prompt segmentation to separate character from background so the performer can be replaced with the generated character while keeping the original motion.
A scroll through the ComfyUI workflow: the reference footage in, the segmentation points, and the finished output.
Results
The results were consistent. Across all five videos the character, style, and proportions stayed locked shot to shot, without the drift and morphing common in generated footage. The headset also stayed on-model: an accurate Vive Focus 3, the device customers actually use, which matters when the video is teaching them to handle it.
Reflection
This was a first run-through, and it showed where the pipeline gets more efficient: better lighting on the reference footage, so the model reads the performance cleanly instead of filling gaps with visuals we didn't want, and the speed that comes with repetition.
These videos were generated with the WAN model; now that LTX supports audio, I'd love to try a workflow built on that. The video generation space moves fast, so who knows what's coming next.