TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

61

Generative Image Dynamics

Zhengqi Li, Richard Tucker, Noah Snavely +1 authors

A frequency-coordinated diffusion sampling process is used to predict long-term motion representations for still images, enabling dynamic video creation and interactive scene manipulation.

55frequency-coordinated diffusion sampling processneural stochastic motion textureHF ↗arXiv ↗
62

AppAgent: Multimodal Agents as Smartphone Users

Chi Zhang, Zhao Yang, Jiaxuan Liu +5 authors

A novel LLM-based multimodal agent learns to operate smartphone apps through autonomous exploration or imitation, demonstrating proficiency across diverse tasks.

54large language modelsmultimodal agentHF ↗arXiv ↗
68

LRM: Large Reconstruction Model for Single Image to 3D

Yicong Hong, Kai Zhang, Jiuxiang Gu +7 authors

A Large Reconstruction Model using a transformer-based architecture predicts 3D neural radiance fields from single images using massive multi-view training data.

52Large Reconstruction Modeltransformer-based architectureHF ↗arXiv ↗
69

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Shilong Liu, Hao Cheng, Haotian Liu +10 authors

LLaVA-Plus, a general-purpose multimodal assistant, enhances large multimodal models by integrating pre-trained vision and vision-language models, performing tool-assisted tasks and improving interaction through direct image grounding.

52multimodal assistantpre-trained vision and vision-language modelsHF ↗arXiv ↗
71

Challenges and Applications of Large Language Models

Jean Kaddour, Joshua Harris, Maximilian Mozes +3 authors

Large Language Models (LLMs) went from non-existent to ubiquitous in the machine learning discourse within a few years. Due to the fast pace of the field, it is difficult to identify the remaining challenges and already fruitful application areas. In this paper, we aim to establish a systematic set of open problems and application successes so that ML researchers can comprehend the field's current state more quickly and become productive.

51Large Language Models (LLMs)discourseHF ↗arXiv ↗
72

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar +3 authors

Orca, a 13-billion parameter model, enhances small models by imitating the reasoning process from large foundation models using rich signals and diverse data, surpassing existing models in complex reasoning benchmarks.

51imitation learninglarge foundation modelsHF ↗arXiv ↗
73

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +939 authors

Gemini, a family of multimodal models, achieves state-of-the-art performance across various benchmarks, including human-expert performance on MMLU, through advanced cross-modal reasoning and language understanding.

51multimodal modelscross-modal reasoningHF ↗arXiv ↗
76

Diffusion Model Alignment Using Direct Preference Optimization

Bram Wallace, Meihua Dang, Rafael Rafailov +7 authors

A method called Diffusion-DPO aligns text-to-image diffusion models to human preferences using direct optimization on comparison data, improving visual appeal and prompt alignment.

49Reinforcement Learning from Human Feedback (RLHF)human comparison dataHF ↗arXiv ↗
80

StemGen: A music generation model that listens

Julian D. Parker, Janne Spijkervet, Katerina Kosta +6 authors

A transformer-based, non-autoregressive model generates musically coherent audio by responding to context, achieving audio quality comparable to text-conditioned models.

48non-autoregressivetransformer-basedHF ↗arXiv ↗
82

Drivable 3D Gaussian Avatars

Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito +3 authors

A new 3D controllable avatar model uses Gaussian splats for photorealistic rendering in real-time, employing cage deformations driven by joint angles and keypoints, outperforming existing methods.

47Gaussian splats3D Gaussian SplattingHF ↗arXiv ↗
85

Instant3D: Instant Text-to-3D Generation

Ming Li, Pan Zhou, Jia-Wei Liu +4 authors

A framework named Instant3D generates 3D objects from text prompts in under one second using a novel network and adaptive algorithms to enhance efficiency and quality.

47text-to-3D generationneural fieldHF ↗arXiv ↗
87

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Yixin Song, Zeyu Mi, Haotong Xie +1 authors

PowerInfer, a high-speed LLM inference engine for personal computers, enhances efficiency using hotspot neuron analysis, GPU-CPU hybrid computation, adaptive predictors, and neuron-aware sparse operators, achieving performance close to server-grade GPUs.

46Large Language Model (LLM)inference engineHF ↗arXiv ↗
88

Matryoshka Diffusion Models

Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2 authors

Matryoshka Diffusion Models use a NestedUNet architecture for joint denoising at multiple resolutions, enabling efficient high-resolution image and video synthesis.

46diffusion modelshigh-resolution image and video synthesisHF ↗arXiv ↗
3 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号