TensorX
返回文献探索

Paper · arXiv 2411.12364

Ultra-Sparse Memory Network

Zihao Huang, Qiyang Min, Hongzhi Huang, Defa Zhu, Yutao Zeng, Ran Guo, Xun Zhou

23 upvotesNovember 19, 2024arXiv 预印本
AI 摘要

UltraMem introduces a large-scale, ultra-sparse memory layer to Transformer models, reducing inference latency and improving performance.

Mixture of ExpertsTransformer modelsUltraMemultra-sparse memory layerinference latencymodel performancescaling laws

Abstract

It is widely acknowledged that the performance of Transformer models is exponentially related to their number of parameters and computational complexity. While approaches like Mixture of Experts (MoE) decouple parameter count from computational complexity, they still face challenges in inference due to high memory access costs. This work introduces UltraMem, incorporating large-scale, ultra-sparse memory layer to address these limitations. Our approach significantly reduces inference latency while maintaining model performance. We also investigate the scaling laws of this new architecture, demonstrating that it not only exhibits favorable scaling properties but outperforms traditional models. In our experiments, we train networks with up to 20 million memory slots. The results show that our method achieves state-of-the-art inference speed and model performance within a given computational budget.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号