TensorX
返回文献探索

Paper · arXiv 2504.21117

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

Hanhua Hong, Chenghao Xiao, Yang Wang, Yiqi Liu, Wenge Rong, Chenghua Lin

26 upvotesApril 29, 2025arXiv 预印本
AI 摘要

Inversion learning automates the generation of effective evaluation prompts for language models, improving robustness and efficiency over manual processes.

natural language generationNLGhuman evaluationinconsistenciesstandardisationdemographic biasesLLM-based evaluationprompt designinversion learningreverse mappingsinput instructionssingle evaluation sample

Abstract

Evaluating natural language generation (NLG) systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardisation, and demographic biases, limiting reproducibility. LLM-based evaluation offers a scalable alternative but is highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts | TensorX