TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 8 – Jan 14, 2024

50 篇论文 · 按点赞排序

34

Pheme: Efficient and Conversational Speech Generation

Paweł Budzianowski, Taras Sereda, Tomasz Cichy +1 authors

The Pheme model series achieves compact and high-quality voice generation with parallel processing, efficient training on smaller datasets, and improved voice quality through distillation.

18speech generationhierarchical neural audio codecsHF ↗arXiv ↗
40

Diffusion Priors for Dynamic View Synthesis from Monocular Videos

Chaoyang Wang, Peiye Zhuang, Aliaksandr Siarohin +4 authors

A finetuned RGB-D diffusion model combined with dynamic and static Neural Radiance Fields (NeRF) enhances dynamic novel view synthesis by achieving geometric consistency and robust hallucination of unseen video regions.

12RGB-D diffusion modelNeural Radiance FieldsHF ↗arXiv ↗
41

Efficient LLM inference solution on Intel GPU

Hui Wu, Yi Gan, Feng Yuan +7 authors

The proposed LLM inference solution reduces latency and increases throughput by simplifying decoder layers, using a segment KV cache policy, and optimizing the Scaled-Dot-Product-Attention kernel.

12TransformerLarge Language ModelsHF ↗arXiv ↗
43

Score Distillation Sampling with Learned Manifold Corrective

Thiemo Alldieck, Nikos Kolotouros, Cristian Sminchisescu

The paper addresses issues in the Score Distillation Sampling loss function by training a shallow network to remove noisy gradients introduced by high text guidance, demonstrating improvements in various applications like image synthesis, network training, and 3D synthesis.

12image diffusion modelScore Distillation Sampling (SDS)HF ↗arXiv ↗
45

Object-Centric Diffusion for Efficient Video Editing

Kumara Kahatapitiya, Adil Karjauv, Davide Abati +3 authors

Object-Centric Diffusion reduces memory and computational cost in diffusion-based video editing by focusing on salient regions and merging redundant tokens in the background, achieving up to a 10x latency reduction with comparable quality.

10diffusion-based video editingtextual edit promptsHF ↗arXiv ↗
46

AGG: Amortized Generative 3D Gaussians for Single Image to 3D

Dejia Xu, Ye Yuan, Morteza Mardani +4 authors

An Amortized Generative 3D Gaussian framework (AGG) efficiently generates 3D Gaussians from a single image, offering competitive performance and speed advantages over existing methods.

103D Gaussian splattingAmortized Generative 3D Gaussian frameworkHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号