TensorX
返回文献探索

Paper · arXiv 2312.12433

Tracking Any Object Amodally

Cheng-Yen Hsieh, Tarasha Khurana, Achal Dave, Deva Ramanan

12 upvotesDecember 19, 2023arXiv 预印本
AI 摘要

A new TAO-Amodal benchmark and an amodal expander module improve amodal detection and tracking of occluded objects in video sequences.

amodal perceptionTAO-Amodal benchmarkamodal bounding boxesmodal bounding boxesobject permanenceamodal expanderfine-tuningdata augmentation

Abstract

Amodal perception, the ability to comprehend complete object structures from partial visibility, is a fundamental skill, even for infants. Its significance extends to applications like autonomous driving, where a clear understanding of heavily occluded objects is essential. However, modern detection and tracking algorithms often overlook this critical capability, perhaps due to the prevalence of modal annotations in most datasets. To address the scarcity of amodal data, we introduce the TAO-Amodal benchmark, featuring 880 diverse categories in thousands of video sequences. Our dataset includes amodal and modal bounding boxes for visible and occluded objects, including objects that are partially out-of-frame. To enhance amodal tracking with object permanence, we leverage a lightweight plug-in module, the amodal expander, to transform standard, modal trackers into amodal ones through fine-tuning on a few hundred video sequences with data augmentation. We achieve a 3.3\% and 1.6\% improvement on the detection and tracking of occluded objects on TAO-Amodal. When evaluated on people, our method produces dramatic improvements of 2x compared to state-of-the-art modal baselines.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Tracking Any Object Amodally | TensorX