TensorX
返回文献探索

Paper · arXiv 2505.00358

R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training

Albert Ge, Tzu-Heng Huang, John Cooper, Avi Trost, Ziyi Chu, Satya Sai Srinath Namburi GNVV, Ziyang Cai, Kendall Park, Nicholas Roberts, Frederic Sala

26 upvotesMay 1, 2025arXiv 预印本
AI 摘要

R&B, a framework that repartitions and balances training data based on semantic similarity and domain gradients, enhances language model performance with minimal additional computational cost.

semantic similarityGram matrixdomain gradientsdata mixing strategiesevaluation informationregularity conditionsmultimodal tasks

Abstract

Data mixing strategies have successfully reduced the costs involved in training language models. While promising, such methods suffer from two flaws. First, they rely on predetermined data domains (e.g., data sources, task types), which may fail to capture critical semantic nuances, leaving performance on the table. Second, these methods scale with the number of domains in a computationally prohibitive way. We address these challenges via R&B, a framework that re-partitions training data based on semantic similarity (Regroup) to create finer-grained domains, and efficiently optimizes the data composition (Balance) by leveraging a Gram matrix induced by domain gradients obtained throughout training. Unlike prior works, it removes the need for additional compute to obtain evaluation information such as losses or gradients. We analyze this technique under standard regularity conditions and provide theoretical insights that justify R&B's effectiveness compared to non-adaptive mixing approaches. Empirically, we demonstrate the effectiveness of R&B on five diverse datasets ranging from natural language to reasoning and multimodal tasks. With as little as 0.01% additional compute overhead, R&B matches or exceeds the performance of state-of-the-art data mixing strategies.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training | TensorX