TensorX
返回文献探索

Paper · arXiv 2407.03958

Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge

Young-Jun Lee, Dokyong Lee, Junyoung Youn, Kyeongjin Oh, Byungsoo Ko, Jonghwan Hyeon, Ho-Jin Choi

20 upvotesJuly 4, 2024arXiv 预印本
AI 摘要

A large-scale dataset and framework are introduced to capture long-term multi-modal conversations, enhancing a model's visual imagination ability.

multi-modal conversation datasetStarkmulti-modal contextualization frameworkMcuPlan-and-Execute image alignerChatGPTmulti-modal conversation modelUltron 7Bvisual imagination

Abstract

Humans share a wide variety of images related to their personal experiences within conversations via instant messaging tools. However, existing works focus on (1) image-sharing behavior in singular sessions, leading to limited long-term social interaction, and (2) a lack of personalized image-sharing behavior. In this work, we introduce Stark, a large-scale long-term multi-modal conversation dataset that covers a wide range of social personas in a multi-modality format, time intervals, and images. To construct Stark automatically, we propose a novel multi-modal contextualization framework, Mcu, that generates long-term multi-modal dialogue distilled from ChatGPT and our proposed Plan-and-Execute image aligner. Using our Stark, we train a multi-modal conversation model, Ultron 7B, which demonstrates impressive visual imagination ability. Furthermore, we demonstrate the effectiveness of our dataset in human evaluation. We make our source code and dataset publicly available.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Stark: Social Long-Term Multi-Modal Conversation with Persona Commonsense Knowledge | TensorX