TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 27 – Jun 2, 2024
本周最热91

An Introduction to Vision-Language Modeling

Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay +38 authors

Introduction to vision-language models (VLMs) covering their applications, training, evaluation, and extension to videos, addressing challenges in mapping visual data to language.

Large Language Models (LLMs)vision-language models (VLMs)visual assistantgenerative modelsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Transformers Can Do Arithmetic with the Right Embeddings

Sean McLeish, Arpit Bansal, Alex Stein +8 authors

Transformers achieve state-of-the-art performance on large arithmetic tasks and other reasoning tasks by addressing positional tracking with embeddings and integrating architectural modifications.

55transformerspositional trackingHF ↗arXiv ↗
03

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Byung-Kwan Lee, Chae Won Kim, Beomchan Park +1 authors

Meteor, an efficient large language and vision model, enhances understanding and answering capabilities by embedding multifaceted rationales using the Mamba architecture, leading to improved vision language performance without increasing model size or using additional vision encoders.

54large language and vision modelsvisual instruction tuningHF ↗arXiv ↗
05

Phased Consistency Model

Fu-Yun Wang, Zhaoyang Huang, Alexander William Bergman +9 authors

The Phased Consistency Model (PCM) addresses limitations in Latent Consistency Models (LCM) and outperforms them in high-resolution, text-conditioned image and few-step text-to-video generation.

47consistency modeldiffusion modelsHF ↗arXiv ↗
06

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Chunjiang Ge, Sijie Cheng, Ziming Wang +6 authors

ConvLLaVA addresses excessive visual tokens and quadratic complexity in high-resolution multimodal models by using ConvNeXt as a hierarchical backbone, optimizing pretrained ConvNeXt for high resolution, and adding a successive stage for further compression, achieving competitive performance.

45High-resolution Large Multimodal ModelsLMMsHF ↗arXiv ↗
09

Matryoshka Multimodal Models

Mu Cai, Jianwei Yang, Jianfeng Gao +1 authors

Large Multimodal Models (LMMs) such as LLaVA have shown strong performance in visual-linguistic reasoning. These models first embed images into a fixed large number of visual tokens and then feed them into a Large Language Model (LLM). However, this design causes an excessive number of tokens for dense visual scenarios such as high-resolution images and videos, leading to great inefficiency. While token pruning/merging methods do exist, they produce a single length output for each image and do not afford flexibility in trading off information density v.s. efficiency. Inspired by the concept of Matryoshka Dolls, we propose M3: Matryoshka Multimodal Models, which learns to represent visual content as nested sets of visual tokens that capture information across multiple coarse-to-fine granularities. Our approach offers several unique benefits for LMMs: (1) One can explicitly control the visual granularity per test instance during inference, e.g. , adjusting the number of tokens used to represent an image based on the anticipated complexity or simplicity of the content; (2) M3 provides a framework for analyzing the granularity needed for existing datasets, where we find that COCO-style benchmarks only need around ~9 visual tokens to obtain accuracy similar to that of using all 576 tokens; (3) Our approach provides a foundation to explore the best trade-off between performance and visual token length at sample level, where our investigation reveals that a large gap exists between the oracle upper bound and current fixed-scale representations.

35Multimodal Models (LMMs)LLaVAHF ↗arXiv ↗
13

Zamba: A Compact 7B SSM Hybrid Model

Paolo Glorioso, Quentin Anthony, Yury Tokpanov +4 authors

Zamba, a 7B SSM-transformer hybrid model, achieves competitive performance with minimal parameter cost and efficient inference by combining a Mamba backbone with a shared attention module.

26SSM-transformerMamba backboneHF ↗arXiv ↗
15

2BP: 2-Stage Backpropagation

Christopher Rae, Joseph K. L. Lee, James Richings

Introducing 2-stage backpropagation enhances the throughput of pipeline parallelism in training large DNNs by reducing idle compute time.

25Deep Neural Networks (DNNs)pipeline parallelismHF ↗arXiv ↗
16

The Road Less Scheduled

Aaron Defazio, Xingyu, Yang +4 authors

A new schedule-free method achieves state-of-the-art performance in optimization without requiring a stopping time or additional hyper-parameters, by unifying scheduling and iterate averaging.

25learning rate schedulesoptimization stopping stepHF ↗arXiv ↗
22

Yuan 2.0-M32: Mixture of Experts with Attention Router

Shaohua Wu, Jiangang Luo, Xi Chen +12 authors

Yuan 2.0-M32, using a mixture-of-experts architecture with an Attention Router, surpasses Llama3-70B on MATH and ARC-Challenge benchmarks with lower computational requirements and parameters.

20mixture-of-experts architectureAttention RouterHF ↗arXiv ↗
24

Xwin-LM: Strong and Scalable Alignment Practice for LLMs

Bolin Ni, JingCheng Hu, Yixuan Wei +4 authors

Xwin-LM is a suite of alignment methodologies for large language models that includes supervised finetuning, reward modeling, rejection sampling finetuning, and direct preference optimization, showing consistent improvements on evaluation benchmarks.

17supervised finetuningreward modelingHF ↗arXiv ↗
29

Human4DiT: Free-view Human Video Generation with 4D Diffusion Transformer

Ruizhi Shao, Youxin Pang, Zerong Zheng +2 authors

We present a novel approach for generating high-quality, spatio-temporally coherent human videos from a single image under arbitrary viewpoints. Our framework combines the strengths of U-Nets for accurate condition injection and diffusion transformers for capturing global correlations across viewpoints and time. The core is a cascaded 4D transformer architecture that factorizes attention across views, time, and spatial dimensions, enabling efficient modeling of the 4D space. Precise conditioning is achieved by injecting human identity, camera parameters, and temporal signals into the respective transformers. To train this model, we curate a multi-dimensional dataset spanning images, videos, multi-view data and 3D/4D scans, along with a multi-dimensional training strategy. Our approach overcomes the limitations of previous methods based on GAN or UNet-based diffusion models, which struggle with complex motions and viewpoint changes. Through extensive experiments, we demonstrate our method's ability to synthesize realistic, coherent and free-view human videos, paving the way for advanced multimedia applications in areas such as virtual reality and animation. Our project website is https://human4dit.github.io.

16U-Netsdiffusion transformersHF ↗arXiv ↗
30

iVideoGPT: Interactive VideoGPTs are Scalable World Models

Jialong Wu, Shaofeng Yin, Ningya Feng +4 authors

Interactive VideoGPT is a scalable multimodal autoregressive transformer for world modeling that integrates visual observations, actions, and rewards, enabling competitive video prediction, planning, and reinforcement learning.

16autoregressive transformertokenization techniqueHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号