TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

December 2023
本月最热265

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko +5 authors

Efficient inference for large language models on devices with limited DRAM by optimizing data transfer and access from flash memory.

large language models (LLMs)flash memoryDRAMinference cost modelHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Magicoder: Source Code Is All You Need

Yuxiang Wei, Zhe Wang, Jiawei Liu +2 authors

Magicoder, using OSS-Instruct to incorporate open-source code snippets, achieves superior performance on coding benchmarks while reducing bias in synthetic data generation.

83Large Language ModelsLLMsHF ↗arXiv ↗
07

LLM360: Towards Fully Transparent Open-Source LLMs

Zhengzhong Liu, Aurick Qiao, Willie Neiswanger +25 authors

LLM360 initiative promotes full transparency and reproducibility in LLM training by open-sourcing training code, data, model checkpoints, and intermediate results.

57Large Language ModelsLLaMAHF ↗arXiv ↗
09

AppAgent: Multimodal Agents as Smartphone Users

Chi Zhang, Zhao Yang, Jiaxuan Liu +5 authors

A novel LLM-based multimodal agent learns to operate smartphone apps through autonomous exploration or imitation, demonstrating proficiency across diverse tasks.

54large language modelsmultimodal agentHF ↗arXiv ↗
10

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +939 authors

Gemini, a family of multimodal models, achieves state-of-the-art performance across various benchmarks, including human-expert performance on MMLU, through advanced cross-modal reasoning and language understanding.

51multimodal modelscross-modal reasoningHF ↗arXiv ↗
12

StemGen: A music generation model that listens

Julian D. Parker, Janne Spijkervet, Katerina Kosta +6 authors

A transformer-based, non-autoregressive model generates musically coherent audio by responding to context, achieving audio quality comparable to text-conditioned models.

48non-autoregressivetransformer-basedHF ↗arXiv ↗
14

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Yixin Song, Zeyu Mi, Haotong Xie +1 authors

PowerInfer, a high-speed LLM inference engine for personal computers, enhances efficiency using hotspot neuron analysis, GPU-CPU hybrid computation, adaptive predictors, and neuron-aware sparse operators, achieving performance close to server-grade GPUs.

46Large Language Model (LLM)inference engineHF ↗arXiv ↗
15

Kandinsky 3.0 Technical Report

Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev +6 authors

Kandinsky 3.0, a large-scale text-to-image model based on latent diffusion, improves quality and realism through a larger architecture and advanced text understanding.

45latent diffusionU-NetHF ↗arXiv ↗
19

SparQ Attention: Bandwidth-Efficient LLM Inference

Luka Ribar, Ivan Chelombiev, Luke Hudlass-Galley +3 authors

SparQ Attention reduces memory bandwidth requirements in LLM attention blocks, enhancing inference throughput without accuracy loss.

40SparQ Attentiongenerative large language models (LLMs)HF ↗arXiv ↗
24

Generative Multimodal Models are In-Context Learners

Quan Sun, Yufeng Cui, Xiaosong Zhang +8 authors

A large-scale generative multimodal model with 37 billion parameters demonstrates strong few-shot in-context learning and achieves state-of-the-art performance on multimodal tasks through scaling-up and instruction tuning.

36task-agnostic in-context learninggenerative multimodal modelHF ↗arXiv ↗
27

Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

Zeyi Sun, Ye Fang, Tong Wu +6 authors

Alpha-CLIP enhances CLIP by adding an auxiliary alpha channel for attentive region suggestion, enabling precise control over image content across various tasks like open-world recognition and multimodal generation.

34CLIPContrastive Language-Image Pre-trainingHF ↗arXiv ↗
30

Analyzing and Improving the Training Dynamics of Diffusion Models

Tero Karras, Miika Aittala, Jaakko Lehtinen +3 authors

Modifications to network layers in the ADM diffusion model architecture improve training stability and synthesis quality, reducing FID from 2.41 to 1.81, and a novel method for post-hoc EMA parameter tuning is introduced.

33diffusion modelsADM diffusion modelHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号