TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

548

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Enric Corona, Andrei Zanfir, Eduard Gabriel Bazavan +3 authors

VLOGGER generates audio-driven human videos from a single image using a diffusion-based method that includes 3D motion and text-to-image models, outperforming existing methods in quality, identity, and consistency.

36stochastic human-to-3d-motion diffusion modeldiffusion-based architectureHF ↗arXiv ↗
549

Pre-training Small Base LMs with Fewer Tokens

Sunny Sanyal, Sujay Sanghavi, Alexandros G. Dimakis

Inheritune leverages transformer blocks from large language models and minimal data to create smaller, efficient base models with performance competitive to larger models.

36transformer blocksInherituneHF ↗arXiv ↗
554

LLMs + Persona-Plug = Personalized LLMs

Jiongnan Liu, Yutao Zhu, Shuting Wang +6 authors

A novel model constructs user-specific embeddings using a lightweight plug-in module to personalize LLM outputs without fine-tuning, improving performance on various tasks.

35large language models (LLMs)personalized LLMHF ↗arXiv ↗
555

Matryoshka Multimodal Models

Mu Cai, Jianwei Yang, Jianfeng Gao +1 authors

Large Multimodal Models (LMMs) such as LLaVA have shown strong performance in visual-linguistic reasoning. These models first embed images into a fixed large number of visual tokens and then feed them into a Large Language Model (LLM). However, this design causes an excessive number of tokens for dense visual scenarios such as high-resolution images and videos, leading to great inefficiency. While token pruning/merging methods do exist, they produce a single length output for each image and do not afford flexibility in trading off information density v.s. efficiency. Inspired by the concept of Matryoshka Dolls, we propose M3: Matryoshka Multimodal Models, which learns to represent visual content as nested sets of visual tokens that capture information across multiple coarse-to-fine granularities. Our approach offers several unique benefits for LMMs: (1) One can explicitly control the visual granularity per test instance during inference, e.g. , adjusting the number of tokens used to represent an image based on the anticipated complexity or simplicity of the content; (2) M3 provides a framework for analyzing the granularity needed for existing datasets, where we find that COCO-style benchmarks only need around ~9 visual tokens to obtain accuracy similar to that of using all 576 tokens; (3) Our approach provides a foundation to explore the best trade-off between performance and visual token length at sample level, where our investigation reveals that a large gap exists between the oracle upper bound and current fixed-scale representations.

35Multimodal Models (LMMs)LLaVAHF ↗arXiv ↗
557

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97 authors

InternLM2 is an open-source LLM that outperforms predecessors through innovative pre-training and optimization techniques, including Supervised Fine-Tuning and Conditional Online Reinforcement Learning from Human Feedback.

35Large Language ModelsLLMsHF ↗arXiv ↗
563

Extending Llama-3's Context Ten-Fold Overnight

Peitian Zhang, Ninglu Shao, Zheng Liu +4 authors

Llama-3-8B-Instruct's context length is extended from 8K to 80K using QLoRA fine-tuning with minimal additional training samples, demonstrating significant potential for further context extension with increased computational resources.

34QLoRA fine-tuningHF ↗arXiv ↗
564

FlowMind: Automatic Workflow Generation with LLMs

Zhen Zeng, William Watson, Nicole Cho +4 authors

FlowMind uses Large Language Models with a generic prompt recipe to generate automatic workflows, addressing spontaneous tasks and ensuring data integrity, and it is evaluated using a new financial dataset NCEN-QA.

34Large Language ModelsGenerative Pretrained TransformerHF ↗arXiv ↗
565

Pegasus-v1 Technical Report

Raehyuk Jung, Hyojun Go, Jaehyuk Yi +41 authors

Pegasus-1 is a multimodal language model designed for video content comprehension, handling spatiotemporal information and demonstrated in benchmarks for video conversation, zero-shot video question answering, and video summarization.

34multimodal language modelvideo content understandingHF ↗arXiv ↗
19 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号