TensorX
返回文献探索

Paper · arXiv 2409.07239

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

Yang Liu, Pengxiang Ding, Siteng Huang, Min Zhang, Han Zhao, Donglin Wang

14 upvotesSeptember 11, 2024arXiv 预印本
AI 摘要

A new video-language model, PiTe, uses trajectory-guided pixel-temporal alignment to achieve superior performance across various multimodal video tasks by leveraging a large pre-training dataset with precise object trajectories.

Large Language Models (LLMs)Large Visual-Language Models (LVLMs)videospatial-temporal dataLarge Video-Language Models (LVidLMs)latent spacemulti-modal tasksfine-grained alignmentobject trajectorytrajectory-guided Pixel-Temporal AlignmentPiTepre-training datasetPiTe-143kautomatic annotation pipeline

Abstract

Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due to the complexity of the relationship between language and spatial-temporal data structure. Recent Large Video-Language Models (LVidLMs) align feature of static visual data like image into latent space of language feature, by general multi-modal tasks to leverage abilities of LLMs sufficiently. In this paper, we explore fine-grained alignment approach via object trajectory for different modalities across both spatial and temporal dimensions simultaneously. Thus, we propose a novel LVidLM by trajectory-guided Pixel-Temporal Alignment, dubbed PiTe, that exhibits promising applicable model property. To achieve fine-grained video-language alignment, we curate a multi-modal pre-training dataset PiTe-143k, the dataset provision of moving trajectories in pixel level for all individual objects, that appear and mention in the video and caption both, by our automatic annotation pipeline. Meanwhile, PiTe demonstrates astounding capabilities on myriad video-related multi-modal tasks through beat the state-of-the-art methods by a large margin.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
PiTe: Pixel-Temporal Alignment for Large Video-Language Model | TensorX