TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热157

Your Transformer is Secretly Linear

Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova +4 authors

Transformer decoders exhibit near-perfect linear relationships between layers, which can be reduced with cosine-similarity-based regularization, leading to improved performance on benchmarks.

transformer decodersProcrustes similarity scoreresidual componentlinear blocksHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Chameleon: Mixed-Modal Early-Fusion Foundation Models

Chameleon Team

Chameleon is a mixed-modal early-fusion token-based model that achieves state-of-the-art performance across various tasks, including image captioning, text generation, and long-form mixed-modal generation, using a unified architecture.

135early-fusiontoken-basedHF ↗arXiv ↗
05

Octopus v4: Graph of language models

Wei Chen, Zhiyuan Li

The Octopus v4 model uses functional tokens to integrate and direct queries to task-specific open-source language models, achieving SOTA performance with models under 10B parameters.

116functional tokensOctopus v4HF ↗arXiv ↗
06

KAN: Kolmogorov-Arnold Networks

Ziming Liu, Yixuan Wang, Sachin Vaidya +5 authors

Kolmogorov-Arnold Networks (KANs) outperform Multi-Layer Perceptrons (MLPs) in accuracy and interpretability by using learnable activation functions and spline-based weights.

115Kolmogorov-Arnold NetworksKANsHF ↗arXiv ↗
07

What matters when building vision-language models?

Hugo Laurençon, Léo Tronchon, Matthieu Cord +1 authors

Idefics2, a vision-language model with 8 billion parameters, achieves state-of-the-art performance on multimodal benchmarks through extensive experimental validation.

104vision-language modelslarge language modelsHF ↗arXiv ↗
08

An Introduction to Vision-Language Modeling

Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay +38 authors

Introduction to vision-language models (VLMs) covering their applications, training, evaluation, and extension to videos, addressing challenges in mapping visual data to language.

91Large Language Models (LLMs)vision-language models (VLMs)HF ↗arXiv ↗
09

LoRA Learns Less and Forgets Less

Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes +9 authors

LoRA, a parameter-efficient finetuning method for large language models, underperforms full finetuning in target domains but provides better regularization and maintains diverse generation compared to other techniques.

91Low-Rank AdaptationLoRAHF ↗arXiv ↗
10

Better & Faster Large Language Models via Multi-token Prediction

Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière +2 authors

Training large language models to predict multiple future tokens simultaneously enhances their sample efficiency and performance, particularly on generative benchmarks and small algorithmic tasks, with decreased inference time.

82next-token prediction lossmulti-token predictionHF ↗arXiv ↗
12

RLHF Workflow: From Reward Modeling to Online RLHF

Hanze Dong, Wei Xiong, Bo Pang +7 authors

Online iterative reinforcement learning from human feedback achieves state-of-the-art performance in large language models using open-source datasets and proxy preference models.

71RLHFOnline Iterative RLHFHF ↗arXiv ↗
13

WildChat: 1M ChatGPT Interaction Logs in the Wild

Wenting Zhao, Xiang Ren, Jack Hessel +3 authors

Chatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bridge this gap, we offered free access to ChatGPT for online users in exchange for their affirmative, consensual opt-in to anonymously collect their chat transcripts and request headers. From this, we compiled WildChat, a corpus of 1 million user-ChatGPT conversations, which consists of over 2.5 million interaction turns. We compare WildChat with other popular user-chatbot interaction datasets, and find that our dataset offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study. In addition to timestamped chat transcripts, we enrich the dataset with demographic data, including state, country, and hashed IP addresses, alongside request headers. This augmentation allows for more detailed analysis of user behaviors across different geographical regions and temporal dimensions. Finally, because it captures a broad range of use cases, we demonstrate the dataset's potential utility in fine-tuning instruction-following models. WildChat is released at https://wildchat.allen.ai under AI2 ImpACT Licenses.

66HF ↗arXiv ↗
15

Transformers Can Do Arithmetic with the Right Embeddings

Sean McLeish, Arpit Bansal, Alex Stein +8 authors

Transformers achieve state-of-the-art performance on large arithmetic tasks and other reasoning tasks by addressing positional tracking with embeddings and integrating architectural modifications.

55transformerspositional trackingHF ↗arXiv ↗
16

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Byung-Kwan Lee, Chae Won Kim, Beomchan Park +1 authors

Meteor, an efficient large language and vision model, enhances understanding and answering capabilities by embedding multifaceted rationales using the Mamba architecture, leading to improved vision language performance without increasing model size or using additional vision encoders.

55large language and vision modelsvisual instruction tuningHF ↗arXiv ↗
18

MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning

Ting Jiang, Shaohan Huang, Shengyue Luo +8 authors

MoRA, a high-rank updating method using square matrices, enhances the ability of large language models to learn and memorize new knowledge, especially in memory-intensive tasks, compared to LoRA.

50low-rank adaptationparameter-efficient fine-tuningHF ↗arXiv ↗
19

Iterative Reasoning Preference Optimization

Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3 authors

An iterative preference optimization method using a modified DPO loss improves reasoning accuracy on various datasets by optimizing winning and losing reasoning steps in Chain-of-Thought candidates.

50iterative preference optimizationChain-of-Thought (CoT)HF ↗arXiv ↗
21

Phased Consistency Model

Fu-Yun Wang, Zhaoyang Huang, Alexander William Bergman +9 authors

The Phased Consistency Model (PCM) addresses limitations in Latent Consistency Models (LCM) and outperforms them in high-resolution, text-conditioned image and few-step text-to-video generation.

48consistency modeldiffusion modelsHF ↗arXiv ↗
24

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Chunjiang Ge, Sijie Cheng, Ziming Wang +6 authors

ConvLLaVA addresses excessive visual tokens and quadratic complexity in high-resolution multimodal models by using ConvNeXt as a hierarchical backbone, optimizing pretrained ConvNeXt for high resolution, and adding a successive stage for further compression, achieving competitive performance.

46High-resolution Large Multimodal ModelsLMMsHF ↗arXiv ↗
27

Not All Language Model Features Are Linear

Joshua Engels, Isaac Liao, Eric J. Michaud +2 authors

Research explores multi-dimensional features in language models, discovering interpretable circular representations in GPT-2, Mistral 7B, and Llama 3 8B, which are used for modular arithmetic tasks.

40linear representation hypothesismulti-dimensional featuresHF ↗arXiv ↗
28

SUTRA: Scalable Multilingual Language Model Architecture

Abhijit Bendale, Michael Sapienza, Steven Ripplinger +3 authors

SUTRA, a multilingual Large Language Model architecture, achieves superior performance on multilingual tasks by decoupling conceptual understanding from language-specific processing using a Mixture of Experts framework.

38Multilingual Large Language ModelMixture of Experts frameworkHF ↗arXiv ↗
30

Matryoshka Multimodal Models

Mu Cai, Jianwei Yang, Jianfeng Gao +1 authors

Large Multimodal Models (LMMs) such as LLaVA have shown strong performance in visual-linguistic reasoning. These models first embed images into a fixed large number of visual tokens and then feed them into a Large Language Model (LLM). However, this design causes an excessive number of tokens for dense visual scenarios such as high-resolution images and videos, leading to great inefficiency. While token pruning/merging methods do exist, they produce a single length output for each image and do not afford flexibility in trading off information density v.s. efficiency. Inspired by the concept of Matryoshka Dolls, we propose M3: Matryoshka Multimodal Models, which learns to represent visual content as nested sets of visual tokens that capture information across multiple coarse-to-fine granularities. Our approach offers several unique benefits for LMMs: (1) One can explicitly control the visual granularity per test instance during inference, e.g. , adjusting the number of tokens used to represent an image based on the anticipated complexity or simplicity of the content; (2) M3 provides a framework for analyzing the granularity needed for existing datasets, where we find that COCO-style benchmarks only need around ~9 visual tokens to obtain accuracy similar to that of using all 576 tokens; (3) Our approach provides a foundation to explore the best trade-off between performance and visual token length at sample level, where our investigation reveals that a large gap exists between the oracle upper bound and current fixed-scale representations.

35Multimodal Models (LMMs)LLaVAHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号