TensorX
返回文献探索

Paper · arXiv 2512.05081

Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression

Jung Yi, Wooseok Jang, Paul Hyunbin Cho, Jisu Nam, Heeji Yoon, Seungryong Kim

33 upvotesDecember 4, 2025arXiv 预印本
AI 摘要

Deep Forcing, a training-free method, enhances real-time video diffusion by addressing temporal repetition and motion issues through Deep Sink and Participative Compression, achieving high-quality, long-duration video generation.

autoregressive video diffusionStreamingLLM-style attentionDeep ForcingDeep Sinktemporal RoPE phaseParticipative CompressionKV cache pruningLongLiveRollingForcingdynamic degree

Abstract

Recent advances in autoregressive video diffusion have enabled real-time frame streaming, yet existing solutions still suffer from temporal repetition, drift, and motion deceleration. We find that naively applying StreamingLLM-style attention sinks to video diffusion leads to fidelity degradation and motion stagnation. To overcome this, we introduce Deep Forcing, which consists of two training-free mechanisms that address this without any fine-tuning. Specifically, 1) Deep Sink dedicates half of the sliding window to persistent sink tokens and re-aligns their temporal RoPE phase to the current timeline, stabilizing global context during long rollouts. 2) Participative Compression performs importance-aware KV cache pruning that preserves only tokens actively participating in recent attention while safely discarding redundant and degraded history, minimizing error accumulation under out-of-distribution length generation. Together, these components enable over 12x extrapolation (e.g. 5s-trained to 60s+ generation) with better imaging quality than LongLive, better aesthetic quality than RollingForcing, almost maintaining overall consistency, and substantial gains in dynamic degree, all while maintaining real-time generation. Our results demonstrate that training-free KV-cache management can match or exceed training-based approaches for autoregressively streaming long-video generation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression | TensorX