Active-speaker-driven video cropping: what the inference actually cost
PyTorch Forums
Active-speaker-driven video cropping: what the inference actually cost
Sharing a local video pipeline and, more usefully, the two inference lessons that cost me the most time. The task: crop a 16:9 stream to 9:16 following whoever is speaking. YOLOv8 for people, OpenCV for the crop path, TalkNet-ASD to pick the speaker from the tracked faces. The model was never the expensive part. Feeding it was. On an RTX 3060, for a 60-second clip: tracking at 8fps 62.3s TalkNet forward pass, two people 1.2s frame reads + crops reusing...
0 comments
No comments yet.