When an LLM is apprehensive about its answers -- and when its uncertainty is justified
Petr Sychev, Andrey Goncharov, Daniil Vyazhev +2 authors
Token-wise entropy and model-as-judge (MASJ) are evaluated as uncertainty estimation methods for multiple-choice question-answering tasks, showing that entropy performs better in knowledge-dependent domains but requires reasoning, while MASJ needs refinement.