TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热192

GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Jiawei Zhao, Zhenyu Zhang, Beidi Chen +3 authors

Gradient Low-Rank Projection (GaLore) improves memory efficiency for training large language models without sacrificing performance, enabling pre-training of 7B models on consumer GPUs.

low-rank adaptation (LoRA)Gradient Low-Rank Projection (GaLore)memory efficiencypre-trainingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

StarCoder 2 and The Stack v2: The Next Generation

Anton Lozhkov, Raymond Li, Loubna Ben Allal +63 authors

StarCoder2, a large language model for code developed through a collaboration with Software Heritage, outperforms other models of similar size on various benchmarks and matches or outperforms larger models in specific areas.

160Large Language ModelsCode LLMsHF ↗arXiv ↗
06

SaulLM-7B: A pioneering Large Language Model for Law

Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf +8 authors

SaulLM-7B, a large language model with 7 billion parameters, excels in legal text comprehension and generation using instructional fine-tuning on a legal corpus.

93large language modellegal domainHF ↗arXiv ↗
07

Stealing Part of a Production Language Model

Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +10 authors

A model-stealing attack is introduced that can extract detailed information such as the embedding projection layer from black-box language models with minimal cost.

90model-stealing attackembedding projection layerHF ↗arXiv ↗
08

The Unreasonable Ineffectiveness of the Deeper Layers

Andrey Gromov, Kushal Tirumala, Hassan Shapourian +2 authors

Layer pruning of pre-trained LLMs with parameter-efficient finetuning methods shows minimal performance degradation and significant resource savings in both finetuning and inference.

82layer-pruningopen-weight pretrained LLMsHF ↗arXiv ↗
12

LLM Agent Operating System

Kai Mei, Zelong Li, Shuyuan Xu +3 authors

AIOS, an operating system embedding large language models, addresses resource allocation, context switching, and concurrency challenges for intelligent agents, demonstrating reliability and efficiency.

73large language modelintelligent agentsHF ↗arXiv ↗
14

RAFT: Adapting Language Model to Domain Specific RAG

Tianjun Zhang, Shishir G. Patil, Naman Jain +4 authors

Retrieval Augmented FineTuning (RAFT) enhances pre-trained large language models' in-domain performance by filtering out irrelevant documents and citing relevant passages.

72Retrieval Augmented FineTuningRAFTHF ↗arXiv ↗
16

Yi: Open Foundation Models by 01.AI

01. AI, Alex Young, Bei Chen +28 authors

The Yi model family, based on transformer architecture, showcases strong performance across benchmarks and modalities through optimized data and scalable infrastructure.

66language modelsmultimodal modelsHF ↗arXiv ↗
21

Evolutionary Optimization of Model Merging Recipes

Takuya Akiba, Makoto Shing, Yujin Tang +2 authors

An evolutionary algorithm automates model merging, discovering effective combinations of open-source models without additional training data or compute, achieving state-of-the-art performance on benchmarks and culturally-aware tasks.

58evolutionary algorithmsfoundation modelsHF ↗arXiv ↗
23

ViTAR: Vision Transformer with Any Resolution

Qihang Fan, Quanzeng You, Xiaotian Han +5 authors

ViTAR enhances Vision Transformers' scalability across resolutions through dynamic token integration and fuzzy positional encoding, improving accuracy and reducing computational costs.

56Vision Transformersdynamic resolution adjustmentHF ↗arXiv ↗
29

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin +105 authors

Gemma, a family of lightweight and high-performing language models, outperforms similarly sized open models across text-based tasks and emphasizes the importance of responsible model development and safety.

52lightweightstate-of-the art open modelsHF ↗arXiv ↗
30

Chronos: Learning the Language of Time Series

Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen +14 authors

Chronos, a framework using pretrained transformer-based models for time series forecasting, outperforms classical methods and achieves comparable zero-shot performance on unseen datasets using tokenized time series data.

51pretrained probabilistic time series modelsChronosHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号