Publication: A Predictive Mechanistic Benchmark for Language Modeling in Transformers and State-Space Models
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Sequence modeling architectures drive state-of-the-art results across language, vision, and science, yet no mechanistic evaluation suite exists for identifying the capabilities each design acquires or lacks. Transformers offer content-addressable retrieval at O(N^2) cost, while state space models such as Mamba and FlashSTU run at O(N) cost with weaker token recall, and layer-alternating hybrids combine both primitives with poorly characterized capability profiles. This thesis makes two contributions. We introduce a mechanistic benchmark of 28 synthetic tasks isolating retrieval, aggregation, and compound capabilities, runnable in under one GPU-hour, whose rankings correlate with downstream language-modeling loss. We then propose Hydra, a mixing block that colocates attention, Mamba, and STU heads within a single layer. At 150M and 1B parameters on OLMo-2, Hydra consistently matches or surpasses Transformer and layer-alternating baselines, with its advantage widening at scale. The benchmark proves a useful tool for rapidly innovating architectures like Hydra without the cost of large-scale pretraining.