TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 12 – Jun 18, 2023
本周最热113

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

Shuai Yang, Yifan Zhou, Ziwei Liu +1 authors

A novel framework adapts image diffusion models for video by generating key frames with hierarchical constraints and propagating them using patch matching and blending, achieving high-quality and temporally-coherent videos.

diffusion modelstext-guided video-to-video translationkey frame translationhierarchical cross-frame constraintsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

TryOnDiffusion: A Tale of Two UNets

Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors

A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.

75diffusion-based architectureParallel-UNetHF ↗arXiv ↗
03

Judging LLM-as-a-judge with MT-Bench and Chatbot Arena

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10 authors

Using strong large language models as judges for evaluating other LLM-based chat assistants achieves high agreement with human preferences, offering a scalable and explainable solution compared to traditional benchmarks.

44large language modelLLMHF ↗arXiv ↗
04

Seeing the World through Your Eyes

Hadi Alzayer, Kevin Zhang, Brandon Feng +2 authors

A method for reconstructing 3D scenes beyond a camera's line of sight using eye reflections is proposed, refining cornea poses, radiance fields, and iris textures.

34cornea posesradiance fieldHF ↗arXiv ↗
08

One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning

Arnav Chavan, Zhuang Liu, Deepak Gupta +2 authors

GLoRA, an advanced method for parameter-efficient fine-tuning, enhances LoRA with a generalized prompt module and modular layer-wise structure search, offering superior performance across diverse tasks with fewer parameters and computational costs.

25Generalized LoRAGLoRAHF ↗arXiv ↗
09

Controlling Text-to-Image Diffusion by Orthogonal Finetuning

Zeju Qiu, Weiyang Liu, Haiwen Feng +6 authors

Orthogonal Finetuning and Constrained Orthogonal Finetuning methods enhance text-to-image diffusion models by preserving hyperspherical energy and improving stability, leading to better generation quality and speed.

25diffusion modelsOrthogonal FinetuningHF ↗arXiv ↗
10

Knowledge Distillation of Large Language Models

Yuxian Gu, Li Dong, Furu Wei +1 authors

MiniLLM distills knowledge from large generative language models to smaller models using reverse KLD for better precision, quality, and performance.

24Knowledge Distillation (KD)large language models (LLMs)HF ↗arXiv ↗
11

Benchmarking Neural Network Training Algorithms

George E. Dahl, Frank Schneider, Zachary Nado +22 authors

A new benchmark, AlgoPerf: Training Algorithms, addresses challenges in evaluating training algorithms by providing a competitive, time-to-result benchmark across workloads and optimizers.

22update rulestuning protocolsHF ↗arXiv ↗
14

h2oGPT: Democratizing Large Language Models

Arno Candel, Jon McKinney, Philipp Singer +12 authors

h2oGPT provides open-source, fine-tuned LLMs based on Generative Pretrained Transformers with 100% private document search capabilities.

19Generative Pretrained TransformersLLMsHF ↗arXiv ↗
16

Augmenting Language Models with Long-Term Memory

Weizhi Wang, Li Dong, Hao Cheng +4 authors

A framework called LongMem enables large language models to utilize long-term memory, overcoming input length limitations and improving performance on long-context tasks.

19language modelslong-term memoryHF ↗arXiv ↗
18

DreamHuman: Animatable 3D Avatars from Text

Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir +3 authors

DreamHuman generates realistic animatable 3D human avatars using text by integrating text-to-image synthesis, neural radiance fields, and statistical human body models.

17text-to-3D methodsneural radiance fieldsHF ↗arXiv ↗
19

Scalable 3D Captioning with Pretrained Models

Tiange Luo, Chris Rockwell, Honglak Lee +1 authors

Cap3D generates high-quality descriptive text for 3D objects using pretrained models and datasets, surpassing human performance in quality, cost, and speed.

17image captioningimage-text alignmentHF ↗arXiv ↗
22

Language to Rewards for Robotic Skill Synthesis

Wenhao Yu, Nimrod Gileadi, Chuyuan Fu +17 authors

A new method uses large language models to define reward parameters for control policies, bridging high-level language instructions to low-level robotic actions through an interactive system.

13large language modelsin-context learningHF ↗arXiv ↗
23

High-Fidelity Audio Compression with Improved RVQGAN

Rithesh Kumar, Prem Seetharaman, Alejandro Luebs +2 authors

A universal neural audio compression algorithm achieves high fidelity at 8 kbps bandwidth by combining advancements in audio generation, vector quantization, and improved loss functions.

13neural compression modelhigh-fidelity audio generationHF ↗arXiv ↗
25

Image Captioners Are Scalable Vision Learners Too

Michael Tschannen, Manoj Kumar, Andreas Steiner +3 authors

Captioning alone, when done with a carefully controlled comparison, proves to be as effective and sometimes more powerful than contrastive pretraining in vision and vision-language tasks.

12contrastive pretrainingimage-text pairsHF ↗arXiv ↗
27

ATT3D: Amortized Text-to-3D Object Synthesis

Jonathan Lorraine, Kevin Xie, Xiaohui Zeng +7 authors

A framework called Amortized Text-to-3D (ATT3D) efficiently generates 3D objects from text by sharing computation across multiple text prompts and enabling knowledge sharing for novel setups and smooth animations.

11text-to-3D modellinggenerative text-to-image modelsHF ↗arXiv ↗
30

Diffusion Models for Zero-Shot Open-Vocabulary Segmentation

Laurynas Karazija, Iro Laina, Andrea Vedaldi +1 authors

Zero-shot open-vocabulary segmentation uses text-to-image diffusion models to sample support images, enhancing localization and background segmentation with pre-trained feature extractors.

10zero-shot open-vocabulary segmentationcontrastive trainingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号