TensorX
返回文献探索

Paper · arXiv 2608.29974

SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models

Amanuel Gizachew Abebe, Yasmin Moslem

6 upvotesAugust 30, 2026arXiv 预印本
AI 摘要

A hybrid system combining a multimodal sequence tagger and a generative vision-language model improves hallucination span detection and calibration through union-calibrated fusion.

Large Vision-Language Modelshallucination detectionspan localizationsequence taggerXLM-RoBERTa-LargeSigLIPcross-attentiongenerative VLMQwen3.5-4BUnion-Calibrated Fusioncalibration correlationIoU

Abstract

Detecting hallucinations in Large Vision-Language Models (LVLMs) requires both accurate span localization and well-calibrated confidence scores. Fine-tuned generative VLMs excel at identifying hallucinated text spans but suffer from overconfidence and high inference latency. Discriminative sequence taggers offer deterministic speed and superior calibration but exhibit conservative span recall. We present SpanCalib-VLM, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with our fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT). Through a Union-Calibrated Fusion strategy, candidate spans from the generative model are re-scored with calibrated probabilities from the sequence tagger. On the SHROOM-Visions English evaluation split, our ensemble achieves a Pearson calibration correlation of 0.41 and an overall IoU of 0.39, with a clean-response IoU of 0.91} and overall detection accuracy of 70.7%. We make our model weights and code publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号