TensorX
返回文献探索

Paper · arXiv 2508.10180

For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMs

Wenlong Deng, Qi Zeng, Jiaming Zhang, Minghui Chen, Zixin Ding, Christos Thrampoulidis, Boying Gong, Xiaoxiao Li

20 upvotesApril 25, 2026arXiv 预印本
AI 摘要

For-Value is a forward-only data valuation framework that efficiently estimates data value using final hidden representations and prediction errors, enabling scalable batch processing without gradient computations.

data valuationlarge language modelsvision-language modelsgradient computationsforward-onlybatch-scalablehidden representationsprediction errorsclosed-form expressionsingle forward passgradient-based baselines

Abstract

Data valuation is essential for enhancing the transparency and accountability of large language models (LLMs) and vision-language models (VLMs). However, existing methods typically rely on gradient computations, making them computationally prohibitive for billion-parameter models and precluding batch parallelization. In this work, we introduce For-Value, a forward-only data valuation framework that enables efficient batch-scalable value estimation while maintaining effectiveness. Leveraging the expressive power of pretrained LLMs/VLMs, we theoretically demonstrate that data valuation can be captured by the alignment between the final hidden representations and prediction errors at the last layer. In light of this insight, For-Value computes data value using a simple closed-form expression with a single forward pass, eliminating the need for costly backpropagation and enabling efficient batch calculating at scale. Extensive experiments show that For-Value matches or outperforms gradient-based baselines in detecting influential data and mislabeled data, while achieving significant efficiency improvements.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号