TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热170

Simple and Controllable Music Generation

Jade Copet, Felix Kreuk, Itai Gat +5 authors

MusicGen, a single-stage transformer language model, generates high-quality music conditioned on text or melodic features using efficient token interleaving patterns, outperforming existing models on text-to-music benchmarks.

Language Modeltransformer LMtoken interleaving patternsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Textbooks Are All You Need

Suriya Gunasekar, Yi Zhang, Jyoti Aneja +16 authors

A new compact Transformer-based large language model for code, phi-1, achieves high accuracy on coding benchmarks despite having fewer parameters than competing models.

159Transformer-basedHumanEvalHF ↗arXiv ↗
03

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

Shuai Yang, Yifan Zhou, Ziwei Liu +1 authors

A novel framework adapts image diffusion models for video by generating key frames with hierarchical constraints and propagating them using patch matching and blending, achieving high-quality and temporally-coherent videos.

113diffusion modelstext-guided video-to-video translationHF ↗arXiv ↗
04

Long-range Language Modeling with Self-retrieval

Ohad Rubin, Jonathan Berant

The Retrieval-Pretrained Transformer jointly trains a retrieval-augmented language model to model long texts by fusing retrieved information for improved prediction.

97retrieval-augmented language modelsRetrieval-Pretrained TransformerHF ↗arXiv ↗
05

TryOnDiffusion: A Tale of Two UNets

Luyang Zhu, Dawei Yang, Tyler Zhu +5 authors

A diffusion-based architecture unifies garment detail preservation and warping for pose and shape variation in virtual try-on tasks.

75diffusion-based architectureParallel-UNetHF ↗arXiv ↗
08

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar +3 authors

Orca, a 13-billion parameter model, enhances small models by imitating the reasoning process from large foundation models using rich signals and diverse data, surpassing existing models in complex reasoning benchmarks.

51imitation learninglarge foundation modelsHF ↗arXiv ↗
11

Judging LLM-as-a-judge with MT-Bench and Chatbot Arena

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10 authors

Using strong large language models as judges for evaluating other LLM-based chat assistants achieves high agreement with human preferences, offering a scalable and explainable solution compared to traditional benchmarks.

44large language modelLLMHF ↗arXiv ↗
13

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou +4 authors

Generalized Knowledge Distillation addresses distribution mismatch in auto-regressive sequence models by training the student on its own output sequences and integrates with RL fine-tuning for effective distillation in tasks like summarization and instruction-tuning.

38Knowledge distillationauto-regressive sequence modelsHF ↗arXiv ↗
16

Fast Segment Anything

Xu Zhao, Wenchao Ding, Yongqi An +5 authors

A speed-up method using a CNN detector with an instance segmentation branch achieves comparable performance to SAM with 50 times faster runtime by training on a small subset of SAM's dataset.

36segment anything modelSAMHF ↗arXiv ↗
17

Seeing the World through Your Eyes

Hadi Alzayer, Kevin Zhang, Brandon Feng +2 authors

A method for reconstructing 3D scenes beyond a camera's line of sight using eye reflections is proposed, refining cornea poses, radiance fields, and iris textures.

34cornea posesradiance fieldHF ↗arXiv ↗
23

MotionGPT: Human Motion as a Foreign Language

Biao Jiang, Xin Chen, Wen Liu +3 authors

MotionGPT, a unified motion-language model using discrete vector quantization, achieves top performance in various motion-related tasks by treating motion as a language similar to text.

28discrete vector quantizationmotion tokensHF ↗arXiv ↗
26

One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning

Arnav Chavan, Zhuang Liu, Deepak Gupta +2 authors

GLoRA, an advanced method for parameter-efficient fine-tuning, enhances LoRA with a generalized prompt module and modular layer-wise structure search, offering superior performance across diverse tasks with fewer parameters and computational costs.

25Generalized LoRAGLoRAHF ↗arXiv ↗
27

Controlling Text-to-Image Diffusion by Orthogonal Finetuning

Zeju Qiu, Weiyang Liu, Haiwen Feng +6 authors

Orthogonal Finetuning and Constrained Orthogonal Finetuning methods enhance text-to-image diffusion models by preserving hyperspherical energy and improving stability, leading to better generation quality and speed.

25diffusion modelsOrthogonal FinetuningHF ↗arXiv ↗
28

Knowledge Distillation of Large Language Models

Yuxian Gu, Li Dong, Furu Wei +1 authors

MiniLLM distills knowledge from large generative language models to smaller models using reverse KLD for better precision, quality, and performance.

24Knowledge Distillation (KD)large language models (LLMs)HF ↗arXiv ↗
29

Generate Anything Anywhere in Any Scene

Yuheng Li, Haotian Liu, Yangming Wen +1 authors

A text-to-image diffusion model is enhanced with adapter layers and regionally-guided sampling to achieve controlled generation of personalized objects with high fidelity.

23diffusion modelsentanglement issuesHF ↗arXiv ↗
30

Training Transformers with 4-bit Integers

Haocheng Xi, Changhao Li, Jianfei Chen +1 authors

A novel method for training transformers with 4-bit quantization achieves competitive accuracy and accelerates training on current GPUs.

23activation quantizationweight quantizationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号