CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models
Hao-Wen Dong, Xiaoyu Liu, Jordi Pons +5 authors
A conditional diffusion model leverages pretrained CLIP and diffusion prior models to generate audio from text queries using unlabeled videos, offering competitive performance in text-to-audio and image-to-audio synthesis.