CoT Overthinking: Do LLMs Lose Answers?
Project Overview
In this project, we investigate how an LLM's answer to a multiple choice question changes over the course of its reasoning process.
What are the key results of your project?
Out of the 481 questions that the model got wrong, it got the right answer at some point in its reasoning for 108 of those questions - 22.8%. Of those 108 questions, 47 were correct at multiple deciles before landing on the incorrect answer.
We ran two control experiments: random and shuffle. For the shuffle experiment we randomised the order of the tokens in the chain of thought. The accuracy plateaued at 17%. For the random experiment we inserted random tokens into the chain of thought. Here the accuracy plateaued at 10% (equivalent to guessing).
We find that the accuracy of the model monotonically increases at each decile - this is expected.
An unexpected trend we found was that sometimes models get the answer to questions correct, but then lose this correct answer. We briefly investigate what potential causes of this through a sparse autoencoder analysis.
GitHub Repository: https://github.com/tobypullan/CoTOverthinking
How do you feel you have benefitted from completing this internship?
This internship has given me the opportunity to do my first proper AI safety research project. It has given me valuable experience on how to do research, in particular on choosing good research directions.
It has made me consider a career in academia, and made me want to do a PhD even more!