Publication:

Spectral Scaling: Structured Sequence Mixers in In-Context Regression

datacite.rightsrestricted
dc.contributor.advisorHazan, Elad
dc.contributor.authorDevre, H. Yagiz
dc.date.accessioned2026-07-23T16:43:24Z
dc.date.available2026-07-23T16:43:24Z
dc.date.issued2026-04-27
dc.description.abstractIn-context regression is the task in which a sequence model reads labeled examples from its context window and predicts labels for new query points. During this process, no weight update is performed. For deep linear attention, recent work has shown that the forward pass reduces to a single matrix Γ. This reduction preconditions an in-context gradient descent on the training examples. The resulting scaling laws in depth, context length, and width follow from the spectra of Γ and the data. This thesis asks what happens when the architecture is a structured sequence mixer instead. The family includes state-space models, spectral sequence models, semiseparable operators, and hybrids. We show that, in fact, this reduction survives in the structured sequence mixer setting but only under an algebraic condition we call GD-compatibility. Mixers that satisfy this condition collapse to a reduced model through a pair of operators we denote (AS, BS). These operators separate token-side structure from feature-side preconditioning. In this thesis, we study this behavior in two different cases. First, we look at the case in which AS is circulant, and show that the discrete Fourier transform diagonalizes the dynamics exactly. The loss decouples across frequency modes, and gradient flow stays inside the circulant algebra. Depth stops being a bottleneck once the inverse filter can be cached in the weights. The only remaining limit becomes the number of controllable spectral modes. Second, we study basis drift, modeled by a band-random-rotation ensemble. This ensemble interpolates between the fully fixed and fully random regimes. Haar symmetry forces the effective operator into a block-scalar class with one scalar per band. That restriction creates a finite-depth error for pure spectral models, and motivates a basis-adaptive component. Under additional assumptions on how accurately that component estimates the true basis, the excess risk decomposes separably. The five terms are optimization time, context length, spectral bottleneck, spectral depth, and basis-adaptive depth. Every exact claim in the thesis is verified numerically at machine precision. The code is publicly available in the following repository: github.com/ygzdvr/spectral-icl
dc.identifier.urihttps://theses-dissertations.princeton.edu/handle/88435/dsp01x346d7697
dc.language.isoen_US
dc.rights.accessRightsLimited Access: This thesis can be viewed on computers in the Mudd Manuscript Library.
dc.titleSpectral Scaling: Structured Sequence Mixers in In-Context Regression
dc.typePrinceton University Senior Theses
dspace.entity.typePublication
dspace.workflow.startDateTime2026-04-27T20:44:23.052Z
pu.contributor.authorid920347259
pu.date.classyear2026
pu.departmentComputer Science
pu.minorMathematics
pu.minorStatistics and Machine Learning
pu.minorFinance

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
hd0216_written_final_report (6).pdf
Size:
1.18 MB
Format:
Adobe Portable Document Format
Download

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
100 B
Format:
Item-specific license agreed to upon submission
Description:
Download