TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热64

QLoRA: Efficient Finetuning of Quantized LLMs

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman +1 authors

QLoRA enables efficient finetuning of large language models using 4-bit quantization and Low Rank Adapters, achieving high performance with reduced memory usage.

QLoRALow Rank AdaptersLoRA4-bit quantizedHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

StarCoder: may the source be with you!

Raymond Li, Loubna Ben Allal, Yangtian Zi +64 authors

StarCoder, a 15.5B parameter LLM trained on 1 trillion tokens, outperforms other open Code LLMs across multiple languages and fine-tuned Python, with safety enhancements and publicly available under the Open Responsible AI Model license.

34Large Language ModelsCode LLMsHF ↗arXiv ↗
05

LIMA: Less Is More for Alignment

Chunting Zhou, Pengfei Liu, Puxin Xu +12 authors

A 65B parameter LLaMa language model trained with minimal instructional data matches or outperforms models with extensive human preference modeling in most cases, indicating that pretraining is predominantly responsible for knowledge acquisition.

27large language modelsunsupervised pretrainingHF ↗arXiv ↗
06

RWKV: Reinventing RNNs for the Transformer Era

Bo Peng, Eric Alcaide, Quentin Anthony +27 authors

A new architecture, RWKV, combines the parallelizable training of Transformers with the efficient inference of RNNs, achieving linear scaling and matching performance.

21Transformersrecurrent neural networks (RNNs)HF ↗arXiv ↗
08

SoundStorm: Efficient Parallel Audio Generation

Zalán Borsos, Matt Sharifi, Damien Vincent +3 authors

SoundStorm, a non-autoregressive audio generation model, delivers high-quality and consistent audio two orders of magnitude faster than autoregressive methods.

15SoundStormnon-autoregressiveHF ↗arXiv ↗
10

Voyager: An Open-Ended Embodied Agent with Large Language Models

Guanzhi Wang, Yuqi Xie, Yunfan Jiang +5 authors

Voyager is an LLM-powered agent in Minecraft that autonomously explores, learns, and discovers skills through a curriculum, skill library, and prompting mechanism, demonstrating superior lifelong learning and task-solving capabilities.

13LLM-poweredembodied lifelong learning agentHF ↗arXiv ↗
11

LDM3D: Latent Diffusion Model for 3D

Gabriela Ben Melech Stan, Diana Wofk, Scottie Fox +8 authors

Latent Diffusion Model for 3D (LDM3D) generates RGBD images from text prompts and is used to create immersive 360-degree-view experiences.

13Latent Diffusion Model for 3DLDM3DHF ↗arXiv ↗
16

MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers

Lili Yu, Dániel Simig, Colin Flaherty +3 authors

Megabyte, a multi-scale decoder architecture, enables efficient byte-level modeling of long sequences through sub-quadratic self-attention, larger feedforward layers, and improved parallelism, achieving performance competitive with subword models and state-of-the-art results.

10Megabytemulti-scale decoder architectureHF ↗arXiv ↗
17

Personalize Segment Anything Model with One Shot

Renrui Zhang, Zhengkai Jiang, Ziyu Guo +5 authors

A training-free and fine-tuning variant of the Segment Anything Model (SAM), PerSAM and PerSAM-F, achieves personalized image and video segmentation using a single reference image and minimal fine-tuning, improving performance on personalized and dreambooth applications.

10Segment Anything ModelSAMHF ↗arXiv ↗
22

PaLM 2 Technical Report

Rohan Anil, Andrew M. Dai, Orhan Firat +125 authors

PaLM 2, a Transformer-based language model, improves multilingual and reasoning capabilities with enhanced efficiency and performance across various tasks compared to its predecessor.

9Transformer-based modelmixture of objectivesHF ↗arXiv ↗
23

Recommender Systems with Generative Retrieval

Shashank Rajput, Nikhil Mehta, Anima Singh +10 authors

A generative retrieval model using Semantic IDs and a Transformer sequence-to-sequence approach improves recommendation system performance and generalization, especially for cold-start items.

9dual-encoder modelApproximate Nearest NeighborHF ↗arXiv ↗
24

Generating Images with Multimodal Language Models

Jing Yu Koh, Daniel Fried, Ruslan Salakhutdinov

A method to integrate text-only large language models with image encoders and decoders through embedding space mapping achieves multimodal capabilities including image retrieval, generation, and dialogue.

8large language modelsLLMsHF ↗arXiv ↗
29

Model Dementia: Generated Data Makes Models Forget

Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao +3 authors

Model dementia, caused by the use of model-generated content in training, affects the quality of generative models like LLMs, VAEs, and GMMs, highlighting the need for genuine human interaction data.

7Variational AutoencodersGaussian Mixture ModelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号