Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
Junda Wu, Zachary Novack, Amit Namburi +5 authors
FUTGA enhances music captioning by generating detailed, temporally-aware descriptions using generative augmentation with temporal compositions and large language models.