TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 11 – Mar 17, 2024
本周最热130

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier +28 authors

A study of Multimodal Large Language Models finds a mix of pre-training data and image encoder design crucial for SOTA performance on multimodal benchmarks.

Multimodal Large Language Modelsimage encodervision language connectorpre-training dataHF ↗arXiv ↗

47 篇论文 · 按点赞排序

02

Stealing Part of a Production Language Model

Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +10 authors

A model-stealing attack is introduced that can extract detailed information such as the embedding projection layer from black-box language models with minimal cost.

90model-stealing attackembedding projection layerHF ↗arXiv ↗
09

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105 authors

Gemma, a family of lightweight and high-performing language models, outperforms similarly sized open models across text-based tasks and emphasizes the importance of responsible model development and safety.

52lightweightstate-of-the art open modelsHF ↗arXiv ↗
10

Chronos: Learning the Language of Time Series

Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen +14 authors

Chronos, a framework using pretrained transformer-based models for time series forecasting, outperforms classical methods and achieves comparable zero-shot performance on unseen datasets using tokenized time series data.

51pretrained probabilistic time series modelsChronosHF ↗arXiv ↗
11

DeepSeek-VL: Towards Real-World Vision-Language Understanding

Haoyu Lu, Wen Liu, Bo Zhang +11 authors

DeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.

49DeepSeek-VLVision-Language ModelHF ↗arXiv ↗
12

Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma +8 authors

Branch-Train-MiX (BTX) method enhances Large Language Models by asynchronously training experts in parallel and integrating them using Mixture-of-Expert layers with token-level routing for improved accuracy and efficiency.

45Large Language ModelsBranch-Train-MiXHF ↗arXiv ↗
13

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Xiwei Hu, Rui Wang, Yixiao Fang +3 authors

ELLA, an Efficient Large Language Model Adapter, enhances text-to-image diffusion models by integrating powerful Large Language Models through a Timestep-Aware Semantic Connector, improving dense prompt comprehension and generation quality.

44diffusion modelstext-to-image generationHF ↗arXiv ↗
14

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Enric Corona, Andrei Zanfir, Eduard Gabriel Bazavan +3 authors

VLOGGER generates audio-driven human videos from a single image using a diffusion-based method that includes 3D motion and text-to-image models, outperforming existing methods in quality, identity, and consistency.

36stochastic human-to-3d-motion diffusion modeldiffusion-based architectureHF ↗arXiv ↗
21

Personalized Audiobook Recommendations at Spotify Through Graph Neural Networks

Marco De Nadai, Francesco Fabbri, Paul Gigioli +11 authors

A scalable recommendation system integrating heterogeneous graph neural networks and a two-tower model enhances personalized audiobook recommendations by leveraging user preferences and reducing model complexity, resulting in improved streaming rates and benefits to other content types.

23Heterogeneous Graph Neural Networks (HGNNs)Two Tower (2T) modelHF ↗arXiv ↗
24

Video Editing via Factorized Diffusion Distillation

Uriel Singer, Amit Zohar, Yuval Kirstain +4 authors

Emu Video Edit (EVE) uses unsupervised distillation, specifically Factorized Diffusion Distillation, to perform video editing by aligning image editing and video generation adapters without needing supervised data.

22text-to-image modelFactorized Diffusion DistillationHF ↗arXiv ↗
27

Algorithmic progress in language models

Anson Ho, Tamay Besiroglu, Ege Erdil +6 authors

Analysis of language model evaluations shows that compute requirements halve faster than hardware improvements, with compute making the larger contribution to performance gains.

19pre-training language modelsdeep learningHF ↗arXiv ↗
30

On the Societal Impact of Open Foundation Models

Sayash Kapoor, Rishi Bommasani, Kevin Klyman +22 authors

Foundation models are powerful technologies: how they are released publicly directly shapes their societal impact. In this position paper, we focus on open foundation models, defined here as those with broadly available model weights (e.g. Llama 2, Stable Diffusion XL). We identify five distinctive properties (e.g. greater customizability, poor monitoring) of open foundation models that lead to both their benefits and risks. Open foundation models present significant benefits, with some caveats, that span innovation, competition, the distribution of decision-making power, and transparency. To understand their risks of misuse, we design a risk assessment framework for analyzing their marginal risk. Across several misuse vectors (e.g. cyberattacks, bioweapons), we find that current research is insufficient to effectively characterize the marginal risk of open foundation models relative to pre-existing technologies. The framework helps explain why the marginal risk is low in some cases, clarifies disagreements about misuse risks by revealing that past work has focused on different subsets of the framework with different assumptions, and articulates a way forward for more constructive debate. Overall, our work helps support a more grounded assessment of the societal impact of open foundation models by outlining what research is needed to empirically validate their theoretical benefits and risks.

17HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号