Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
- URL: http://arxiv.org/abs/2510.04285v1
- Date: Sun, 05 Oct 2025 16:55:58 GMT
- Title: Probing Geometry of Next Token Prediction Using Cumulant Expansion of the Softmax Entropy
- Authors: Karthik Viswanathan, Sang Eon Park,
- Abstract summary: We introduce a cumulant-expansion framework for quantifying how large language models internalize higher-order statistical structure.<n>We track cumulants in GPT-2 and Pythia models on Pile-10K prompts.
- Score: 0.4329197710438657
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: We introduce a cumulant-expansion framework for quantifying how large language models (LLMs) internalize higher-order statistical structure during next-token prediction. By treating the softmax entropy of each layer's logit distribution as a perturbation around its "center" distribution, we derive closed-form cumulant observables that isolate successively higher-order correlations. Empirically, we track these cumulants in GPT-2 and Pythia models on Pile-10K prompts. (i) Structured prompts exhibit a characteristic rise-and-plateau profile across layers, whereas token-shuffled prompts remain flat, revealing the dependence of the cumulant profile on meaningful context. (ii) During training, all cumulants increase monotonically before saturating, directly visualizing the model's progression from capturing variance to learning skew, kurtosis, and higher-order statistical structures. (iii) Mathematical prompts show distinct cumulant signatures compared to general text, quantifying how models employ fundamentally different processing mechanisms for mathematical versus linguistic content. Together, these results establish cumulant analysis as a lightweight, mathematically grounded probe of feature-learning dynamics in high-dimensional neural networks.
Related papers
- InfoNCE Induces Gaussian Distribution [7.8922077372145685]
A loss in contrastive training is InfoNCE and its variants.<n>We show that the InfoNCE objective induces Gaussian structure in representations that emerge from contrastive training.<n>The resulting Gaussian model enables principled analytical treatment of learned representations and is expected to support a wide range of applications in contrastive learning.
arXiv Detail & Related papers (2026-02-27T13:35:58Z) - Maximum entropy based testing in network models: ERGMs and constrained optimization [1.9116784879310027]
We develop a constrained entropy-maximization problem on the space of networks.<n>The resulting test statistics are defined through the Lagrange multipliers associated with the constrained optimization problem.<n>We show that the proposed Lagrange-multiplier framework connects naturally to classical score tests for constrained maximum likelihood estimation.
arXiv Detail & Related papers (2026-02-24T12:35:08Z) - In-Context Linear Regression Demystified: Training Dynamics and Mechanistic Interpretability of Multi-Head Softmax Attention [52.159541540613915]
We study how multi-head softmax attention models are trained to perform in-context learning on linear data.<n>Our results reveal that in-context learning ability emerges from the trained transformer as an aggregated effect of its architecture and the underlying data distribution.
arXiv Detail & Related papers (2025-03-17T02:00:49Z) - Spatial Reasoning with Denoising Models [49.83744014336816]
We introduce a framework to perform reasoning over sets of continuous variables via denoising generative models.<n>For the first time, that order of generation can successfully be predicted by the denoising network itself.<n>Using these findings, we can increase the accuracy of specific reasoning tasks from 1% to >50%.
arXiv Detail & Related papers (2025-02-28T14:08:30Z) - On learning higher-order cumulants in diffusion models [6.610338540492242]
We study the behaviour of higher-order cumulants, or connected n-point functions, under both the forward and backward process.<n>We derive explicit expressions for the moment- and cumulant-generating functionals.<n>We confirm our results in an exactly solvable toy model with nonzero cumulants and in scalar lattice field theory.
arXiv Detail & Related papers (2024-10-28T16:57:02Z) - Convergence of Score-Based Discrete Diffusion Models: A Discrete-Time Analysis [56.442307356162864]
We study the theoretical aspects of score-based discrete diffusion models under the Continuous Time Markov Chain (CTMC) framework.<n>We introduce a discrete-time sampling algorithm in the general state space $[S]d$ that utilizes score estimators at predefined time points.<n>Our convergence analysis employs a Girsanov-based method and establishes key properties of the discrete score function.
arXiv Detail & Related papers (2024-10-03T09:07:13Z) - Bayesian Circular Regression with von Mises Quasi-Processes [57.88921637944379]
In this work we explore a family of expressive and interpretable distributions over circle-valued random functions.<n>For posterior inference, we introduce a new Stratonovich-like augmentation that lends itself to fast Gibbs sampling.<n>We present experiments applying this model to the prediction of wind directions and the percentage of the running gait cycle as a function of joint angles.
arXiv Detail & Related papers (2024-06-19T01:57:21Z) - Learn2Extend: Extending sequences by retaining their statistical
properties with mixture models [7.15769102504304]
This paper addresses the challenge of extending general finite sequences of real numbers within a subinterval of the real line.
Our focus lies on preserving the gap distribution and pair correlation function of these point sets.
Leveraging advancements in deep learning applied to point processes, this paper explores the use of an auto-regressive textitSequence Extension Mixture Model.
arXiv Detail & Related papers (2023-12-03T21:05:50Z) - Structured Radial Basis Function Network: Modelling Diversity for
Multiple Hypotheses Prediction [51.82628081279621]
Multi-modal regression is important in forecasting nonstationary processes or with a complex mixture of distributions.
A Structured Radial Basis Function Network is presented as an ensemble of multiple hypotheses predictors for regression problems.
It is proved that this structured model can efficiently interpolate this tessellation and approximate the multiple hypotheses target distribution.
arXiv Detail & Related papers (2023-09-02T01:27:53Z) - Optimal regularizations for data generation with probabilistic graphical
models [0.0]
Empirically, well-chosen regularization schemes dramatically improve the quality of the inferred models.
We consider the particular case of L 2 and L 1 regularizations in the Maximum A Posteriori (MAP) inference of generative pairwise graphical models.
arXiv Detail & Related papers (2021-12-02T14:45:16Z) - Regularization of Mixture Models for Robust Principal Graph Learning [0.0]
A regularized version of Mixture Models is proposed to learn a principal graph from a distribution of $D$-dimensional data points.
Parameters of the model are iteratively estimated through an Expectation-Maximization procedure.
arXiv Detail & Related papers (2021-06-16T18:00:02Z) - Understanding Neural Abstractive Summarization Models via Uncertainty [54.37665950633147]
seq2seq abstractive summarization models generate text in a free-form manner.
We study the entropy, or uncertainty, of the model's token-level predictions.
We show that uncertainty is a useful perspective for analyzing summarization and text generation models more broadly.
arXiv Detail & Related papers (2020-10-15T16:57:27Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.