TensorX
返回文献探索

Paper · arXiv 2407.20179

Theia: Distilling Diverse Vision Foundation Models for Robot Learning

Jinghuan Shang, Karl Schmeckpeper, Brandon B. May, Maria Vittoria Minniti, Tarik Kelestemur, David Watkins, Laura Herlant

47 upvotesJuly 29, 2024arXiv 预印本
AI 摘要

Theia, a vision foundation model for robot learning, outperforms existing models with less data and smaller sizes by distilling knowledge from multiple vision foundation models.

vision foundation modelrobot policy learningvisual representationsoff-the-shelf modelsfeature norm distributionsentropy

Abstract

Vision-based robot policy learning, which maps visual inputs to actions, necessitates a holistic understanding of diverse visual tasks beyond single-task needs like classification or segmentation. Inspired by this, we introduce Theia, a vision foundation model for robot learning that distills multiple off-the-shelf vision foundation models trained on varied vision tasks. Theia's rich visual representations encode diverse visual knowledge, enhancing downstream robot learning. Extensive experiments demonstrate that Theia outperforms its teacher models and prior robot learning models using less training data and smaller model sizes. Additionally, we quantify the quality of pre-trained visual representations and hypothesize that higher entropy in feature norm distributions leads to improved robot learning performance. Code and models are available at https://github.com/bdaiinstitute/theia.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Theia: Distilling Diverse Vision Foundation Models for Robot Learning | TensorX