TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本年最热264

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko +5 authors

Efficient inference for large language models on devices with limited DRAM by optimizing data transfer and access from flash memory.

large language models (LLMs)flash memoryDRAMinference cost modelHF ↗arXiv ↗

403 篇论文 · 按点赞排序

02

Llama 2: Open Foundation and Fine-Tuned Chat Models

Hugo Touvron, Louis Martin, Kevin Stone +65 authors

Llama 2, a series of pretrained and fine-tuned large language models, achieves superior performance in dialogue tasks compared to open-source alternatives and offers safety enhancements.

252pretrained language modelsfine-tuningHF ↗arXiv ↗
03

GAIA: a benchmark for General AI Assistants

Grégoire Mialon, Clémentine Fourrier, Craig Swift +3 authors

GAIA benchmarks general AI assistants using real-world questions that challenge both reasoning and multi-modality handling, showcasing a significant gap between human and AI performance.

249multi-modality handlingweb browsingHF ↗arXiv ↗
07

Simple and Controllable Music Generation

Jade Copet, Felix Kreuk, Itai Gat +5 authors

MusicGen, a single-stage transformer language model, generates high-quality music conditioned on text or melodic features using efficient token interleaving patterns, outperforming existing models on text-to-music benchmarks.

167Language Modeltransformer LMHF ↗arXiv ↗
08

Textbooks Are All You Need

Suriya Gunasekar, Yi Zhang, Jyoti Aneja +16 authors

A new compact Transformer-based large language model for code, phi-1, achieves high accuracy on coding benchmarks despite having fewer parameters than competing models.

159Transformer-basedHumanEvalHF ↗arXiv ↗
11

Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation

Shuai Yang, Yifan Zhou, Ziwei Liu +1 authors

A novel framework adapts image diffusion models for video by generating key frames with hierarchical constraints and propagating them using patch matching and blending, achieving high-quality and temporally-coherent videos.

113diffusion modelstext-guided video-to-video translationHF ↗arXiv ↗
13

MVDream: Multi-view Diffusion for 3D Generation

Yichun Shi, Peng Wang, Jianglong Ye +3 authors

MVDream generates geometrically consistent multi-view images from text prompts using pre-trained image diffusion models and Score Distillation Sampling, improving 3D generation stability and supporting personalized generation.

106multi-view diffusion modeltext promptHF ↗arXiv ↗
17

Textbooks Are All You Need II: phi-1.5 technical report

Yuanzhi Li, Sébastien Bubeck, Ronen Eldan +3 authors

A new 1.3 billion parameter Transformer-based language model, phi-1.5, demonstrates comparable performance to much larger models on common sense reasoning and complex tasks despite the absence of web data.

92Transformer-based language modelsTinyStoriesHF ↗arXiv ↗
20

Vision Transformers Need Registers

Timothée Darcet, Maxime Oquab, Julien Mairal +1 authors

Additional input tokens in Vision Transformers mitigate artifacts in feature maps, enhancing performance and enabling smoother visual processing.

86Transformersfeature mapsHF ↗arXiv ↗
22

Language Modeling Is Compression

Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne +9 authors

Large language models demonstrate strong compression capabilities, outperforming domain-specific compressors and offering new insights into scaling laws, tokenization, and in-context learning.

85self-supervised modelslarge language modelsHF ↗arXiv ↗
25

Magicoder: Source Code Is All You Need

Yuxiang Wei, Zhe Wang, Jiawei Liu +2 authors

Magicoder, using OSS-Instruct to incorporate open-source code snippets, achieves superior performance on coding benchmarks while reducing bias in synthetic data generation.

83Large Language ModelsLLMsHF ↗arXiv ↗
28

CodePlan: Repository-level Coding using LLMs and Planning

Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade +6 authors

CodePlan automates repository-level coding tasks, such as package migration and temporal code edits, using a planning framework that leverages LLMs with context derived from code repositories and change analysis.

80Large Language ModelsLLMsHF ↗arXiv ↗
29

Large Language Models as Optimizers

Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors

OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.

79derivative-based algorithmsOptimization by PROmpting (OPRO)HF ↗arXiv ↗
1 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号