TensorX
返回文献探索

Paper · arXiv 2311.10768

Memory Augmented Language Models through Mixture of Word Experts

Cicero Nogueira dos Santos, James Lee-Thorp, Isaac Noble, Chung-Ching Chang, David Uthus

18 upvotesNovember 15, 2023arXiv 预印本
AI 摘要

A Mixture of Word Experts (MoWE) model demonstrates superior performance in NLP tasks compared to T5 models and traditional MoE models, using memory-augmented techniques and knowledge-rich routing functions.

Mixture-of-ExpertsMoEMixture of Word ExpertsMoWEdense modelsFLOPslanguage modelsknowledge-intensive tasksmemory augmented modelsparse memory

Abstract

Scaling up the number of parameters of language models has proven to be an effective approach to improve performance. For dense models, increasing model size proportionally increases the model's computation footprint. In this work, we seek to aggressively decouple learning capacity and FLOPs through Mixture-of-Experts (MoE) style models with large knowledge-rich vocabulary based routing functions and experts. Our proposed approach, dubbed Mixture of Word Experts (MoWE), can be seen as a memory augmented model, where a large set of word-specific experts play the role of a sparse memory. We demonstrate that MoWE performs significantly better than the T5 family of models with similar number of FLOPs in a variety of NLP tasks. Additionally, MoWE outperforms regular MoE models on knowledge intensive tasks and has similar performance to more complex memory augmented approaches that often require to invoke custom mechanisms to search the sparse memory.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号