TensorX
返回文献探索

Paper · arXiv 2503.00955

SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking

Nam V. Nguyen, Dien X. Tran, Thanh T. Tran, Anh T. Hoang, Tai V. Duong, Di T. Le, Phuc-Lu Le

28 upvotesMarch 2, 2025arXiv 预印本
AI 摘要

SemViQA combines Semantic-based Evidence Retrieval and Two-step Verdict Classification to enhance Vietnamese fact-checking, achieving top performance with improved speed.

Large Language ModelsLLMSSemantic-based Evidence RetrievalSERTwo-step Verdict ClassificationTVCISE-DSC01ViWikiFCUIT Data Science Challenge

Abstract

The rise of misinformation, exacerbated by Large Language Models (LLMs) like GPT and Gemini, demands robust fact-checking solutions, especially for low-resource languages like Vietnamese. Existing methods struggle with semantic ambiguity, homonyms, and complex linguistic structures, often trading accuracy for efficiency. We introduce SemViQA, a novel Vietnamese fact-checking framework integrating Semantic-based Evidence Retrieval (SER) and Two-step Verdict Classification (TVC). Our approach balances precision and speed, achieving state-of-the-art results with 78.97\% strict accuracy on ISE-DSC01 and 80.82\% on ViWikiFC, securing 1st place in the UIT Data Science Challenge. Additionally, SemViQA Faster improves inference speed 7x while maintaining competitive accuracy. SemViQA sets a new benchmark for Vietnamese fact verification, advancing the fight against misinformation. The source code is available at: https://github.com/DAVID-NGUYEN-S16/SemViQA.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
SemViQA: A Semantic Question Answering System for Vietnamese Information Fact-Checking | TensorX