I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
Andrey Galichin, Alexey Dontsov, Polina Druzhinina +4 authors
A study uses Sparse Autoencoders to identify and enhance reasoning features in Large Language Models, providing insights into the mechanisms driving their reasoning capabilities.