FuguReport

NeuroCogMap Reveals Cognitive Organization of Large Language Models

Authors Zhongxiang Sun, Haolang Lu, Qiang Ma, Qi Li, Qipeng Wang, Liang Pang, Chenyu Liu, Qiankun Li, Hao Sun, Kun Wang, Yi Zeng, Jun Xu, Guoqi Li, Ji-Rong Wen
Affiliations Nanyang Technological University / Chinese Academy of Sciences / Renmin University of China / The University of Hong Kong / Imperial College London
Categories Method / Model Analysis / Mapping cognitive organization in LLMs, Application / Cognitive Modeling / System-level functional parcellation, Evaluation / Model Evaluation / Cross-model feature consistency
License CC BY 4.0

Abstract Overview

NeuroCogMap is a cognitive-neuroscience-inspired framework that organizes internal sparse features of large language models into functional parcels, links those parcels to interpretable functions and cognitive capabilities, and arranges them into a four-level hierarchy spanning perception to application. Using SAE-derived features from models including Gemma and Llama, the authors report that the resulting parcels are semantically coherent, stable at an intermediate granularity, and partly conserved across models. They then use this organization to analyze pathological behaviors such as hallucination and refusal failure, showing that different failures correspond to distinct disruptions in representational versus behavioral-control systems. Beyond model pathology, the framework is used to improve prediction of human cortical responses during naturalistic language comprehension and to derive latent strategy signals that help refine classical cognitive models of decision-making.

Novelty

The paper's main novelty is a system-level mapping framework for LLMs that borrows concepts from cognitive neuroscience, especially functional parcellation, to define interpretable internal units and higher-order cognitive organization. It is also distinctive in using the same internal map across multiple downstream analyses: failure diagnosis and intervention, comparison with human cortical function, and cognitively guided model discovery.

Results

The authors report that a 270-parcel atlas yields coherent within-parcel semantics, predictive functional descriptions, selective parcel-capability links, and partial cross-model correspondence. NeuroCogMap-derived signatures detect hallucination and refusal failure better than several uncertainty-, probing-, or perturbation-based baselines, and parcel-level interventions improve factual accuracy or refusal performance in the tested settings. The framework also achieves the strongest reported cortical-response prediction among the compared representations on the LeBel story-listening dataset and supports improved held-out fit for a refined two-step decision-making model.

Key Points

  1. NeuroCogMap constructs a multi-level internal atlas of LLM cognition by clustering sparse features into functional parcels, annotating them with cognitive functions, linking them to capabilities, and organizing them into a hierarchy.
  2. The mapped organization is used to distinguish different failure mechanisms, with hallucination tied to representational disruptions and refusal failure tied to shifts from safety-related control toward procedural execution; these signatures also support detection and targeted intervention.
  3. NeuroCogMap is connected to human cognition by improving cortical-response prediction during language comprehension and by exposing latent decision strategies that help refine classical cognitive models.

References

This page was created using generative AI such as GPT-5, Claude Opus 4, Gemini 3, Gemini 3.1 Flash Image, and their higher-end successor versions. No guarantee can be made regarding its contents.