TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 7 – Oct 13, 2024

50 篇论文 · 按点赞排序

37

UniMuMo: Unified Text, Music and Motion Generation

Han Yang, Kun Su, Yutong Zhang +4 authors

UniMuMo generates outputs across text, music, and motion modalities using a unified encoder-decoder transformer architecture and aligned unpaired data based on rhythmic patterns.

19Unified multimodal modeltoken-based representationHF ↗arXiv ↗
41

NL-Eye: Abductive NLI for Images

Mor Ventura, Michael Toker, Nitay Calderon +3 authors

NL-Eye is a benchmark that assesses VLMs' visual abductive reasoning skills by evaluating their ability to predict and explain the plausibility of hypothesis images based on premise images.

18Visual Language ModelVLMHF ↗arXiv ↗
45

Intriguing Properties of Large Language and Vision Models

Young-Jun Lee, Byungsoo Ko, Han-Gyu Kim +2 authors

The investigation into large language and vision models reveals that while they excel in advanced reasoning, they may lack robust perception, as seen through various benchmarks evaluating different aspects like permutation invariance and cross-modal alignment.

16large language and vision modelsLLVMsHF ↗arXiv ↗
46

Progressive Autoregressive Video Diffusion Models

Desai Xie, Zhan Xu, Yicong Hong +5 authors

Progressive noise assignment in autoregressive video diffusion models enables high-quality generation of long videos without quality degradation or abrupt scene changes.

16video diffusion modelsautoregressive video diffusion modelsHF ↗arXiv ↗
50

Temporal Reasoning Transfer from Text to Video

Lei Li, Yuanxin Liu, Linli Yao +6 authors

Textual Temporal reasoning Transfer (T3) enhances Video LLMs' temporal understanding by leveraging text-based temporal tasks, leading to superior performance on video benchmarks without additional video data.

14Video LLMstemporal changesHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号