PingPong: A Benchmark for Role-Playing Language Models with User Emulation and Multi-Model Evaluation
Ilya Gusev
A benchmark evaluates language models' role-playing abilities using automated components to simulate and judge dialogues.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Bofei Gao, Feifan Song, Yibo Miao +21 authors
Large Language Models (LLMs) exhibit remarkably powerful capabilities. One of the crucial factors to achieve success is aligning the LLM's output with human preferences. This alignment process often requires only a small amount of data to efficiently enhance the LLM's performance. While effective, research in this area spans multiple domains, and the methods involved are relatively complex to understand. The relationships between different methods have been under-explored, limiting the development of the preference alignment. In light of this, we break down the existing popular alignment strategies into different components and provide a unified framework to study the current alignment strategies, thereby establishing connections among them. In this survey, we decompose all the strategies in preference learning into four components: model, data, feedback, and algorithm. This unified view offers an in-depth understanding of existing alignment algorithms and also opens up possibilities to synergize the strengths of different strategies. Furthermore, we present detailed working examples of prevalent existing algorithms to facilitate a comprehensive understanding for the readers. Finally, based on our unified perspective, we explore the challenges and future research directions for aligning large language models with human preferences.
49 篇论文 · 按点赞排序
Ilya Gusev
A benchmark evaluates language models' role-playing abilities using automated components to simulate and judge dialogues.
Liqiang Jing, Zhehui Huang, Xiaoyang Wang +6 authors
DSBench evaluates large language and vision-language models on realistic data science tasks, revealing significant performance gaps that indicate the need for improved autonomous agents.
Qingkai Fang, Shoutao Guo, Yan Zhou +3 authors
LLaMA-Omni integrates a speech encoder, adaptor, LLM, and decoder for low-latency and high-quality speech interaction with LLMs, outperforming previous models in both content and style.
Praveen K Kanithi, Clément Christophe, Marco AF Pimentel +7 authors
MEDIC framework evaluates Large Language Models across five clinical dimensions to guide model selection in healthcare applications, identifying performance trade-offs and ensuring practical implementation.
Run Luo, Haonan Zhang, Longze Chen +13 authors
MMEvol, a multimodal instruction data evolution framework, enhances the capabilities of Multimodal Large Language Models by generating diverse and complex image-text instruction datasets, leading to improved performance across various vision-language tasks.
Rogerio Bonatti, Dan Zhao, Francesco Bonacci +8 authors
Windows Agent Arena provides a general, reproducible environment for evaluating multi-modal agent performance on Windows OS tasks, with scalability and parallelizability features.
Chenglei Si, Diyi Yang, Tatsunori Hashimoto
LLM-generated research ideas are perceived as more novel than those from human experts but are deemed slightly less feasible, based on blind evaluations by NLP researchers.
Yejie Wang, Keqing He, Dayuan Fu +11 authors
XCoder, a family of models fine-tuned from LLaMA3, achieves state-of-the-art performance on code instruction tasks using a novel data pruning strategy that addresses data leakage and quality in training datasets.
Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried +1 authors
Agent Workflow Memory (AWM) enhances language model-based agents' performance on complex tasks by inducing and using reusable task workflows, resulting in improved success rates and reduced steps compared to baselines.
Jintian Zhang, Cheng Peng, Mengshu Sun +6 authors
OneGen is a framework that integrates retrieval and generation in a single pass within Large Language Models, improving retrieval performance without compromising generative capabilities.
Hongjin Qian, Peitian Zhang, Zheng Liu +2 authors
MemoRAG enhances retrieval-augmented generation by incorporating long-term memory and a dual-system architecture, leading to superior performance across both complex and straightforward tasks.
Chaojun Xiao, Zhengyan Zhang, Chenyang Song +20 authors
The paper explores a modular approach for large language models (LLMs) by decomposing them into functional modules (bricks) to enhance computational efficiency and scalability.
Ji Ha Jang, Hoigi Seo, Se Young Chun
INTRA addresses challenges in weakly supervised affordance grounding by using contrastive learning with exocentric images and vision-language embeddings, enhancing flexibility and robustness.
Guanyu Lin, Tao Feng, Pengrui Han +2 authors
Paper Copilot is a self-evolving LLM system that provides personalized academic support and streamlines the research process by maintaining a real-time updated database and saving researchers time.
Zhuoyan Luo, Fengyuan Shi, Yixiao Ge +3 authors
Open-MAGVIT2 models, ranging from 300M to 1.5B parameters, achieve state-of-the-art image reconstruction and exploration in auto-regressive models with super-large token vocabularies.
Yuan Liu, Zhongyin Zhao, Ziyuan Zhuang +3 authors
We developed a robust vision-language model with a filtered dataset and model soup tuning, achieving competitive performance with a 9B parameter model.
Shun Lei, Yixuan Zhou, Boshi Tang +7 authors
SongCreator is a dual-sequence language model with an attention mask strategy that generates songs from lyrics, achieving state-of-the-art performance in lyrics-to-song and lyrics-to-vocals tasks while controlling acoustic conditions independently.
Yinwei Wu, Xianpan Zhou, Bing Ma +3 authors
The Instance Feature Adapter enhances Text-to-Image diffusion models by accurately positioning and depicting instance features using appearance tokens and semantic maps.
Yu Zhang, Songlin Yang, Ruijie Zhu +9 authors
Gated Slot Attention (GSA) improves memory capacity and efficiency in Transformers through a bounded-memory-control gating mechanism, enhancing recall and facilitating efficient training and inference.
Haibo Yang, Yang Chen, Yingwei Pan +4 authors
Hi3D is a novel video diffusion model that generates high-resolution, multi-view consistent images with detailed textures using 3D-aware priors and refinement techniques.
Alisia Lupidi, Carlos Gemmell, Nicola Cancedda +5 authors
Source2Synth improves LLM performance in structured reasoning and tool usage scenarios by generating high-quality synthetic data points and filtering out low-quality ones.
Ziqi Jin, Wei Lu
ECHO is a self-harmonized chain-of-thought prompting method that improves reasoning performance by consolidating diverse solution paths into a uniform pattern.
Jing Wang, Ao Ma, Jiasong Feng +3 authors
The Proxy Token Diffusion Transformer (PT-DiT) uses sparse representative token attention to efficiently capture global visual information and reduce redundancy in diffusion transformers, achieving competitive performance with reduced computational complexity.
Vickie Ye, Ruilong Li, Justin Kerr +8 authors
Gsplat is an open-source library for Gaussian Splatting with optimized CUDA kernels, offering speed, memory, and convergence improvements over the original implementation.
Qi Yang, Binjie Mao, Zili Wang +6 authors
A controllable video-to-audio synthesis model, Draw an Audio, uses masked instructions and loudness signals for audio-visual synchronization in challenging V2A tasks.
Lorenza Prospero, Abdullah Hamdi, Joao F. Henriques +1 authors
A transformer model predicts adjustments to Gaussian-splatted 3D human models from a single image, improving 3D pose estimation without test-time optimization or 3D point supervision.
NaHyeon Park, Kunhee Kim, Hyunjung Shim
Selective fine-tuning of text encoders, along with augmentation tokens, knowledge-preservation loss, and SNR-weighted sampling, enhances personalized image generation from single reference images.
Sandeep Kumar, Tirthankar Ghosal, Vinayak Goyal +1 authors
LLMs, particularly Claude-2, generate more novel and diverse research ideas compared to other models, as assessed through human evaluation across multiple domains.
Teng Hu, Jiangning Zhang, Ran Yi +3 authors
A method called SaRA reutilizes ineffective parameters in pre-trained diffusion models to enhance fine-tuning, improving generative capabilities and memory efficiency.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号