TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热262

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan +81 authors

Phi-3-mini, a compact 3.8 billion parameter language model, achieves competitive performance with larger models through an enhanced training dataset and alignment.

language modelMMLUMT-benchrobustnessHF ↗arXiv ↗

50 篇论文 · 按点赞排序

06

ReFT: Representation Finetuning for Language Models

Zhengxuan Wu, Aryaman Arora, Zheng Wang +4 authors

Representation Finetuning (ReFT) methods, exemplified by Low-rank Linear Subspace ReFT (LoReFT), achieve high efficiency and performance by adapting representations in frozen base models, outperforming state-of-the-art Parameter-efficient Fine-tuning (PEFT) methods.

101Parameter-efficient fine-tuningRepresentation FinetuningHF ↗arXiv ↗
07

Rho-1: Not All Tokens Are What You Need

Zhenghao Lin, Zhibin Gou, Yeyun Gong +8 authors

Rho-1, a novel language model using Selective Language Modeling, improves efficiency and performance by selectively training on useful tokens rather than all tokens in the corpus.

95next-token predictiontoken-level training dynamicsHF ↗arXiv ↗
08

Learn Your Reference Model for Real Good Alignment

Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov +5 authors

A new method, Trust Region DPO (TR-DPO), is proposed to improve policy alignment in reinforcement learning, outperforming Direct Preference Optimization (DPO) by up to 19% on key datasets by updating the reference policy during training.

91Reinforcement Learning From Human Feedback (RLHF)Kullback-Leibler divergenceHF ↗arXiv ↗
09

Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, Eldon Schoop +5 authors

Ferret-UI, a multimodal large language model tailored for mobile UI screens, enhances understanding and interaction through region annotations and a comprehensive dataset of UI tasks, outperforming existing models including GPT-4V.

83multimodal large language modelsMLLMsHF ↗arXiv ↗
11

OmniFusion Technical Report

Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk +6 authors

The OmniFusion model, leveraging pretrained LLMs and visual adapters, demonstrates superior performance across multiple visual-language benchmarks compared to open-source alternatives.

78multimodal architecturesOmniFusion modelHF ↗arXiv ↗
14

LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach +3 authors

LLM2Vec transforms decoder-only LLMs into effective text encoders using bidirectional attention, masked next token prediction, and unsupervised contrastive learning, achieving state-of-the-art performance on text embedding tasks.

67decoder-only language modelsLLM2VecHF ↗arXiv ↗
17

Multi-Head Mixture-of-Experts

Xun Wu, Shaohan Huang, Wenhui Wang +1 authors

MH-MoE enhances SMoE by splitting tokens into sub-tokens processed by diverse experts in parallel, improving expert activation, context understanding, and overfitting, and demonstrating effectiveness across various modeling tasks.

61Sparse Mixtures of Experts (SMoE)Multi-Head Mixture-of-Experts (MH-MoE)HF ↗arXiv ↗
20

Make Your LLM Fully Utilize the Context

Shengnan An, Zexiong Ma, Zeqi Lin +2 authors

FILM-7B enhances long-context processing by using information-intensive training with synthesized datasets, improving retrieval and performance on real-world tasks without compromising short-context performance.

55information-intensive traininglong-context trainingHF ↗arXiv ↗
28

Dynamic Typography: Bringing Words to Life

Zichen Liu, Yihao Meng, Hao Ouyang +4 authors

Dynamic Typography generates coherent and semantically meaningful text animations using neural displacement fields, end-to-end optimization, and perceptual loss regularization, outperforming baseline methods.

46neural displacement fieldsend-to-end optimizationHF ↗arXiv ↗
29

Advancing LLM Reasoning Generalists with Preference Trees

Lifan Yuan, Ganqu Cui, Hanbin Wang +12 authors

Eurus, a suite of reasoning-optimized large language models, achieves state-of-the-art performance on various benchmarks through UltraInteract, a large-scale, high-quality alignment dataset, and a novel reward modeling objective.

46large language models (LLMs)Mistral-7BHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号