TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

363

LDM3D: Latent Diffusion Model for 3D

Gabriela Ben Melech Stan, Diana Wofk, Scottie Fox +8 authors

Latent Diffusion Model for 3D (LDM3D) generates RGBD images from text prompts and is used to create immersive 360-degree-view experiences.

13Latent Diffusion Model for 3DLDM3DHF ↗arXiv ↗
364

Voyager: An Open-Ended Embodied Agent with Large Language Models

Guanzhi Wang, Yuqi Xie, Yunfan Jiang +5 authors

Voyager is an LLM-powered agent in Minecraft that autonomously explores, learns, and discovers skills through a curriculum, skill library, and prompting mechanism, demonstrating superior lifelong learning and task-solving capabilities.

13LLM-poweredembodied lifelong learning agentHF ↗arXiv ↗
367

Personalize Segment Anything Model with One Shot

Renrui Zhang, Zhengkai Jiang, Ziyu Guo +5 authors

A training-free and fine-tuning variant of the Segment Anything Model (SAM), PerSAM and PerSAM-F, achieves personalized image and video segmentation using a single reference image and minimal fine-tuning, improving performance on personalized and dreambooth applications.

10Segment Anything ModelSAMHF ↗arXiv ↗
368

MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

Lili Yu, Dániel Simig, Colin Flaherty +3 authors

Megabyte, a multi-scale decoder architecture, enables efficient byte-level modeling of long sequences through sub-quadratic self-attention, larger feedforward layers, and improved parallelism, achieving performance competitive with subword models and state-of-the-art results.

10Megabytemulti-scale decoder architectureHF ↗arXiv ↗
374

Recommender Systems with Generative Retrieval

Shashank Rajput, Nikhil Mehta, Anima Singh +10 authors

A generative retrieval model using Semantic IDs and a Transformer sequence-to-sequence approach improves recommendation system performance and generalization, especially for cold-start items.

9dual-encoder modelApproximate Nearest NeighborHF ↗arXiv ↗
375

PaLM 2 Technical Report

Rohan Anil, Andrew M. Dai, Orhan Firat +125 authors

PaLM 2, a Transformer-based language model, improves multilingual and reasoning capabilities with enhanced efficiency and performance across various tasks compared to its predecessor.

9Transformer-based modelmixture of objectivesHF ↗arXiv ↗
378

Generating Images with Multimodal Language Models

Jing Yu Koh, Daniel Fried, Ruslan Salakhutdinov

A method to integrate text-only large language models with image encoders and decoders through embedding space mapping achieves multimodal capabilities including image retrieval, generation, and dialogue.

8large language modelsLLMsHF ↗arXiv ↗
383

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Yao Zhao, Rishabh Joshi, Tianqi Liu +3 authors

Sequence Likelihood Calibration (SLiC) is shown to be an effective and simpler alternative to Reinforcement Learning from Human Feedback (RLHF) for learning from human preferences in language models.

7Reinforcement Learning from Human Feedback (RLHF)Sequence Likelihood Calibration (SLiC)HF ↗arXiv ↗
386

Model Dementia: Generated Data Makes Models Forget

Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao +3 authors

Model dementia, caused by the use of model-generated content in training, affects the quality of generative models like LLMs, VAEs, and GMMs, highlighting the need for genuine human interaction data.

7Variational AutoencodersGaussian Mixture ModelsHF ↗arXiv ↗
13 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号