TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 8 – Jul 14, 2024

50 篇论文 · 按点赞排序

36

GTA: A Benchmark for General Tool Agents

Jize Wang, Zerun Ma, Yining Li +4 authors

GTA benchmark evaluates LLMs' tool-use capabilities with real user queries, deployed tools, and multimodal inputs, revealing significant challenges and identifying bottlenecks.

16General Tool Agentsreal user queriesHF ↗arXiv ↗
37

Controlling Space and Time with Diffusion Models

Daniel Watson, Saurabh Saxena, Lala Li +2 authors

Cascaded diffusion model for 4D novel view synthesis with joint 3D, 4D, and video training, enabling metric scale camera control and state-of-the-art fidelity with temporal dynamics handling.

16cascaded diffusion model4D novel view synthesisHF ↗arXiv ↗
38

VEnhancer: Generative Space-Time Enhancement for Video Generation

Jingwen He, Tianfan Xue, Dongyang Liu +6 authors

VEnhancer enhances video quality by increasing spatial and temporal resolution using a unified video diffusion model and video ControlNet with space-time data augmentation and video-aware conditioning.

16generative space-time enhancementvideo diffusion modelHF ↗arXiv ↗
39

Video-to-Audio Generation with Hidden Alignment

Manjie Xu, Chenxing Li, Yong Ren +4 authors

The study explores vision encoders, auxiliary embeddings, and data augmentation to enhance video-to-audio generation quality and synchronization, demonstrating state-of-the-art capabilities with the VTA-LDM model.

16vision encodersauxiliary embeddingsHF ↗arXiv ↗
42

Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Haoji Zhang, Yiqin Wang, Yansong Tang +4 authors

Flash-VStream is a video-language model that efficiently processes online video streams in real-time and responds to user queries, achieving superior performance compared to existing methods on both online and offline video understanding benchmarks.

15large language modelscross-modal alignmentHF ↗arXiv ↗
44

Compositional Video Generation as Flow Equalization

Xingyi Yang, Xinchao Wang

Vico is a framework that enhances text-to-video diffusion models by ensuring balanced representation of all concepts in the generated video through spatial-temporal attention graphs and max-flow approximations.

13diffusion modelscompositional video generationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号