Evaluating Very Long-Term Conversational Memory of LLM Agents
Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov +3 authors
A pipeline combining LLMs and human annotation generates long-term dialogues to evaluate model performance in understanding lengthy conversations and temporal dynamics.