TensorX
返回文献探索

Paper · arXiv 2403.11901

Larimar: Large Language Models with Episodic Memory Control

Payel Das, Subhajit Chaudhury, Elliot Nelson, Igor Melnyk, Sarath Swaminathan, Sihui Dai, Aurélie Lozano, Georgios Kollias, Vijil Chenthamarakshan, Jiří, Navrátil, Soham Dan, Pin-Yu Chen

33 upvotesMarch 18, 2024arXiv 预印本
AI 摘要

Larimar, a brain-inspired architecture with distributed episodic memory, enhances LLMs for efficient, fast, and accurate knowledge updates, fact editing, and context length generalization.

Larimarbrain-inspired architecturedistributed episodic memoryknowledge updatesone-shot updatesfact editingselective fact forgettinginput context length generalizationLLM-agnosticsequential editingspeed-ups

Abstract

Efficient and accurate updating of knowledge stored in Large Language Models (LLMs) is one of the most pressing research challenges today. This paper presents Larimar - a novel, brain-inspired architecture for enhancing LLMs with a distributed episodic memory. Larimar's memory allows for dynamic, one-shot updates of knowledge without the need for computationally expensive re-training or fine-tuning. Experimental results on multiple fact editing benchmarks demonstrate that Larimar attains accuracy comparable to most competitive baselines, even in the challenging sequential editing setup, but also excels in speed - yielding speed-ups of 4-10x depending on the base LLM - as well as flexibility due to the proposed architecture being simple, LLM-agnostic, and hence general. We further provide mechanisms for selective fact forgetting and input context length generalization with Larimar and show their effectiveness.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Larimar: Large Language Models with Episodic Memory Control | TensorX