TensorX
返回文献探索

Paper · arXiv 2401.04718

Jump Cut Smoothing for Talking Heads

Xiaojuan Wang, Taesung Park, Yang Zhou, Eli Shechtman, Richard Zhang

20 upvotesJanuary 9, 2024arXiv 预印本
AI 摘要

A framework for smoothing jump cuts in talking head videos using DensePose keypoints, face landmarks, and cross-modal attention for seamless transitions.

DensePose keypointsface landmarksimage translation networkcross-modal attentionvideo interpolationtalking head videosjump cutsseamless transitions

Abstract

A jump cut offers an abrupt, sometimes unwanted change in the viewing experience. We present a novel framework for smoothing these jump cuts, in the context of talking head videos. We leverage the appearance of the subject from the other source frames in the video, fusing it with a mid-level representation driven by DensePose keypoints and face landmarks. To achieve motion, we interpolate the keypoints and landmarks between the end frames around the cut. We then use an image translation network from the keypoints and source frames, to synthesize pixels. Because keypoints can contain errors, we propose a cross-modal attention scheme to select and pick the most appropriate source amongst multiple options for each key point. By leveraging this mid-level representation, our method can achieve stronger results than a strong video interpolation baseline. We demonstrate our method on various jump cuts in the talking head videos, such as cutting filler words, pauses, and even random cuts. Our experiments show that we can achieve seamless transitions, even in the challenging cases where the talking head rotates or moves drastically in the jump cut.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Jump Cut Smoothing for Talking Heads | TensorX