TensorX
返回文献探索

Paper · arXiv 2501.17195

Atla Selene Mini: A General Purpose Evaluation Model

Andrei Alexandru, Antonia Calvi, Henry Broomfield, Jackson Golden, Kyle Dai, Mathias Leys, Maurice Burger, Max Bartolo, Roman Engeler, Sashank Pisupati, Toby Drane, Young Sun Park

35 upvotesJanuary 27, 2025arXiv 预印本
AI 摘要

Atla Selene Mini, an 8B language model-as-a-judge, excels across various benchmarks using enhanced data curation and a combined training approach, achieving top performance in zero-shot evaluations and real-world scenarios.

data curationsynthetically generated critiquesdirect preference optimizationsupervised fine-tuningpromptable evaluatorzero-shot agreementhuman expert evaluations

Abstract

We introduce Atla Selene Mini, a state-of-the-art small language model-as-a-judge (SLMJ). Selene Mini is a general-purpose evaluator that outperforms the best SLMJs and GPT-4o-mini on overall performance across 11 out-of-distribution benchmarks, spanning absolute scoring, classification, and pairwise preference tasks. It is the highest-scoring 8B generative model on RewardBench, surpassing strong baselines like GPT-4o and specialized judges. To achieve this, we develop a principled data curation strategy that augments public datasets with synthetically generated critiques and ensures high quality through filtering and dataset ablations. We train our model on a combined direct preference optimization (DPO) and supervised fine-tuning (SFT) loss, and produce a highly promptable evaluator that excels in real-world scenarios. Selene Mini shows dramatically improved zero-shot agreement with human expert evaluations on financial and medical industry datasets. It is also robust to variations in prompt format. Preliminary results indicate that Selene Mini is the top-ranking evaluator in a live, community-driven Judge Arena. We release the model weights on HuggingFace (https://hf.co/AtlaAI/Selene-1-Mini-Llama-3.1-8B) and Ollama to encourage widespread community adoption.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Atla Selene Mini: A General Purpose Evaluation Model | TensorX