TensorX
返回文献探索

Paper · arXiv 2409.01437

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Sushant Gautam, Andrea Storås, Cise Midoglu, Steven A. Hicks, Vajira Thambawita, Pål Halvorsen, Michael A. Riegler

71 upvotesSeptember 2, 2024arXiv 预印本
AI 摘要

Kvasir-VQA is a dataset with question-and-answer annotations for GI diagnostics, supporting image captioning, VQA, synthetic image generation, object detection, and classification.

Visual Question Answering (VQA)image captioningobject detectionclassification

Abstract

We introduce Kvasir-VQA, an extended dataset derived from the HyperKvasir and Kvasir-Instrument datasets, augmented with question-and-answer annotations to facilitate advanced machine learning tasks in Gastrointestinal (GI) diagnostics. This dataset comprises 6,500 annotated images spanning various GI tract conditions and surgical instruments, and it supports multiple question types including yes/no, choice, location, and numerical count. The dataset is intended for applications such as image captioning, Visual Question Answering (VQA), text-based generation of synthetic medical images, object detection, and classification. Our experiments demonstrate the dataset's effectiveness in training models for three selected tasks, showcasing significant applications in medical image analysis and diagnostics. We also present evaluation metrics for each task, highlighting the usability and versatility of our dataset. The dataset and supporting artifacts are available at https://datasets.simula.no/kvasir-vqa.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Kvasir-VQA: A Text-Image Pair GI Tract Dataset | TensorX