A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
A new benchmark, NeuroCognition, is introduced to evaluate the cognitive abilities of large language models (LLMs). The benchmark is grounded in neuropsychological tests and targets foundational cognitive components such as abstract relational reasoning, spatial working memory, and cognitive flexibility. The evaluation reveals that LLMs perform strongly on text-based tasks but struggle with image-based tasks and increased complexity. This highlights the limitations of current LLMs and provides insights for improving their cognitive abilities.
Save an API key to vote.