TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

61

Neural Network Diffusion

Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors

Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.

100diffusion modelsautoencoderHF ↗arXiv ↗
62

Emu3: Next-Token Prediction is All You Need

Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors

Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.

99next-token predictionmultimodal modelsHF ↗arXiv ↗
63

GenEx: Generating an Explorable World

Taiming Lu, Tianmin Shu, Junfei Xiao +8 authors

GenEx generates 3D environments from a single image, enabling AI agents to explore and interact with a consistent, expansive space through guided generative imagination.

98Generative imaginationpanoramic video streamsHF ↗arXiv ↗
67

TinyLlama: An Open-Source Small Language Model

Peiyuan Zhang, Guangtao Zeng, Tianduo Wang +1 authors

TinyLlama, a compact 1.1B language model, leverages FlashAttention to achieve high performance in downstream tasks with enhanced computational efficiency.

96FlashAttentionLlama 2HF ↗arXiv ↗
69

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Yuan Yao, Tianyu Yu, Ao Zhang +20 authors

MiniCPM-V presents a series of efficient Multimodal Large Language Models optimized for end-side deployment, offering high performance and practical usability compared to larger models.

96Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
71

Rho-1: Not All Tokens Are What You Need

Zhenghao Lin, Zhibin Gou, Yeyun Gong +8 authors

Rho-1, a novel language model using Selective Language Modeling, improves efficiency and performance by selectively training on useful tokens rather than all tokens in the corpus.

95next-token predictiontoken-level training dynamicsHF ↗arXiv ↗
72

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You +3 authors

Correlation between cross-modal alignment and vision representation improves performance in multimodal large language models, enabling identification and training of optimal vision representation with reduced computational cost.

95cross-modal alignmentvision representationHF ↗arXiv ↗
76

Sapiens: Foundation for Human Vision Models

Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez +5 authors

Sapiens, a family of vision models for human-centric tasks, achieves superior performance with self-supervised pretraining and can be easily fine-tuned for 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction.

93self-supervised pretrainingmodel adaptationHF ↗arXiv ↗
77

SaulLM-7B: A pioneering Large Language Model for Law

Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf +8 authors

SaulLM-7B, a large language model with 7 billion parameters, excels in legal text comprehension and generation using instructional fine-tuning on a legal corpus.

92large language modellegal domainHF ↗arXiv ↗
80

An Introduction to Vision-Language Modeling

Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay +38 authors

Introduction to vision-language models (VLMs) covering their applications, training, evaluation, and extension to videos, addressing challenges in mapping visual data to language.

91Large Language Models (LLMs)vision-language models (VLMs)HF ↗arXiv ↗
81

Learn Your Reference Model for Real Good Alignment

Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov +5 authors

A new method, Trust Region DPO (TR-DPO), is proposed to improve policy alignment in reinforcement learning, outperforming Direct Preference Optimization (DPO) by up to 19% on key datasets by updating the reference policy during training.

91Reinforcement Learning From Human Feedback (RLHF)Kullback-Leibler divergenceHF ↗arXiv ↗
83

Stealing Part of a Production Language Model

Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +10 authors

A model-stealing attack is introduced that can extract detailed information such as the embedding projection layer from black-box language models with minimal cost.

90model-stealing attackembedding projection layerHF ↗arXiv ↗
84

LoRA Learns Less and Forgets Less

Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes +9 authors

LoRA, a parameter-efficient finetuning method for large language models, underperforms full finetuning in target domains but provides better regularization and maintains diverse generation compared to other techniques.

90Low-Rank AdaptationLoRAHF ↗arXiv ↗
85

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Kevin Qinghong Lin, Linjie Li, Difei Gao +6 authors

ShowUI is a vision-language-action model that enhances GUI assistants by using UI-guided token selection and interleaved vision-language-action streaming, achieving high accuracy and efficiency in zero-shot screenshot grounding across different environments.

90vision-language-action modelUI-Guided Visual Token SelectionHF ↗arXiv ↗
89

GPT-4o System Card

OpenAI, Aaron Hurst, Adam Lerer +416 authors

GPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.

88autoregressive modelomnimodalHF ↗arXiv ↗
90

Baichuan-Omni Technical Report

Yadong Li, Haoze Sun, Mingan Lin +24 authors

Baichuan-Omni, a 7B open-source Multimodal Large Language Model, excels in processing image, video, audio, and text, showcasing competitive performance across multimodal benchmarks.

88Multimodal Large Language Modelmultimodal training schemaHF ↗arXiv ↗
3 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号