TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 9 – Sep 15, 2024
本周最热72

Towards a Unified View of Preference Learning for Large Language Models: A Survey

Bofei Gao, Feifan Song, Yibo Miao +21 authors

Large Language Models (LLMs) exhibit remarkably powerful capabilities. One of the crucial factors to achieve success is aligning the LLM's output with human preferences. This alignment process often requires only a small amount of data to efficiently enhance the LLM's performance. While effective, research in this area spans multiple domains, and the methods involved are relatively complex to understand. The relationships between different methods have been under-explored, limiting the development of the preference alignment. In light of this, we break down the existing popular alignment strategies into different components and provide a unified framework to study the current alignment strategies, thereby establishing connections among them. In this survey, we decompose all the strategies in preference learning into four components: model, data, feedback, and algorithm. This unified view offers an in-depth understanding of existing alignment algorithms and also opens up possibilities to synergize the strengths of different strategies. Furthermore, we present detailed working examples of prevalent existing algorithms to facilitate a comprehensive understanding for the readers. Finally, based on our unified perspective, we explore the challenges and future research directions for aligning large language models with human preferences.

HF ↗arXiv ↗

49 篇论文 · 按点赞排序

06

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Run Luo, Haonan Zhang, Longze Chen +13 authors

MMEvol, a multimodal instruction data evolution framework, enhances the capabilities of Multimodal Large Language Models by generating diverse and complex image-text instruction datasets, leading to improved performance across various vision-language tasks.

48Multimodal Large Language ModelsMMEvolHF ↗arXiv ↗
10

Agent Workflow Memory

Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried +1 authors

Agent Workflow Memory (AWM) enhances language model-based agents' performance on complex tasks by inducing and using reusable task workflows, resulting in improved success rates and reduced steps compared to baselines.

32Agent Workflow MemoryAWMHF ↗arXiv ↗
18

SongCreator: Lyrics-based Universal Song Generation

Shun Lei, Yixuan Zhou, Boshi Tang +7 authors

SongCreator is a dual-sequence language model with an attention mask strategy that generates songs from lyrics, achieving state-of-the-art performance in lyrics-to-song and lyrics-to-vocals tasks while controlling acoustic conditions independently.

22dual-sequence language modelDSLMHF ↗arXiv ↗
23

Self-Harmonized Chain of Thought

Ziqi Jin, Wei Lu

ECHO is a self-harmonized chain-of-thought prompting method that improves reasoning performance by consolidating diverse solution paths into a uniform pattern.

18Chain-of-Thought (CoT) promptinglarge language modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号