TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 22 – Jan 28, 2024
本周最热82

Lumiere: A Space-Time Diffusion Model for Video Generation

Omer Bar-Tal, Hila Chefer, Omer Tov +11 authors

A text-to-video diffusion model using Space-Time U-Net architecture generates realistic, diverse, and coherent videos through a single pass, achieving state-of-the-art results and supporting various content creation and editing tasks.

Space-Time U-Netdiffusion modeltemporal durationspatial down- and up-samplingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Lihe Yang, Bingyi Kang, Zilong Huang +3 authors

Depth Anything is a robust monocular depth estimation model built on a large-scale dataset with data augmentation and auxiliary supervision strategies, achieving state-of-the-art results on various datasets and enhancing depth-conditioned ControlNet.

64monocular depth estimationdata engineHF ↗arXiv ↗
07

MambaByte: Token-free Selective State Space Model

Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan +1 authors

MambaByte, a token-free byte-level autoregressive model, demonstrates computational efficiency and competitive performance compared to subword token-based models, with the added benefit of fast inference due to linear length scaling.

58MambaBytetoken-free language modelsHF ↗arXiv ↗
18

Rethinking Patch Dependence for Masked Autoencoders

Letian Fu, Long Lian, Renhao Wang +6 authors

By using only cross-attention for masked patch reconstruction, CrossMAE achieves comparable and sometimes superior performance to MAE with significantly reduced computational cost.

26masked autoencodersself-attentionHF ↗arXiv ↗
19

Zero Bubble Pipeline Parallelism

Penghui Qi, Xinyi Wan, Guangxing Huang +1 authors

A novel scheduling strategy achieves zero pipeline bubbles in synchronous training by splitting backward computation and bypassing synchronizations during optimization, significantly improving throughput.

25pipeline parallelismpipeline bubblesHF ↗arXiv ↗
22

CheXagent: Towards a Foundation Model for Chest X-Ray Interpretation

Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck +14 authors

A large-scale instruction-tuning dataset and an instruction-tuned foundation model with a clinical large language model and vision encoder are introduced to automate CXR interpretation, outperforming existing models on clinical tasks and evaluated for fairness.

22vision-language foundation modelsCheXinstructHF ↗arXiv ↗
25

WARM: On the Benefits of Weight Averaged Reward Models

Alexandre Ramé, Nino Vieillard, Léonard Hussenot +4 authors

Weight Averaged Reward Models (WARM) improve the alignment and quality of large language models by averaging fine-tuned reward models, addressing distribution shifts and preference inconsistencies.

19large language modelsreinforcement learningHF ↗arXiv ↗
28

Make-A-Shape: a Ten-Million-scale 3D Shape Model

Ka-Hei Hui, Aditya Sanghi, Arianna Rampini +4 authors

Make-A-Shape, a novel 3D generative model, efficiently trains using wavelet-tree representation and diffusion model techniques, generating high-quality shapes from various inputs in under two seconds.

17wavelet-tree representationdiffusion modelHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号