Publication:

Identifying Concepts Used by Human-Like Neural Network Chess Engines

datacite.rightsrestricted
dc.contributor.advisorGriffiths, Tom
dc.contributor.authorLi, Issac
dc.date.accessioned2026-07-23T17:39:05Z
dc.date.available2026-07-23T17:39:05Z
dc.date.issued2026
dc.description.abstractAs neural networks reach expert-level performance in many domains, there is growing interest in understanding the internal representations that guide their decisions. Models trained to mimic human decisions, rather than strive for optimal performance, offer a particularly interesting testbed for interpretability. Identifying what information these networks encode can generate hypotheses about what information humans rely on when performing similar tasks. Chess presents a well-suited domain to investigate this as neural network-based chess engines have achieved superhuman performance in the game, and previous work has demonstrated that human-interpretable chess concepts are decodable from their latent representations. In this thesis, I extend this line of work by applying the technique of linear probing to Maia chess engines, a family of neural network-based chess engines that are trained to emulate human play, and Leela Chess Zero, a superhuman chess engine trained through self-play. I find that human-interpretable chess concepts are linearly decodable from the internal activations. In the Maia models, decodability increases with human-emulated skill level for most concepts and also varies with network depth. Preliminary evidence supports the idea that concept complexity influences the relationship between layer depth and decodability: the simpler the concept, the earlier it peaks in the network’s activation. These results are mirrored in Leela despite it never having seen a human game. Altogether, these results suggest that human understandable chess concepts are useful representations that are employed by neural network-based chess engines and how task representations may shift with expertise.
dc.identifier.urihttps://theses-dissertations.princeton.edu/handle/88435/dsp016d5701142
dc.language.isoen_US
dc.titleIdentifying Concepts Used by Human-Like Neural Network Chess Engines
dc.typePrinceton University Senior Theses
dspace.entity.typePublication
dspace.workflow.startDateTime2026-06-23T18:11:34.784Z
pu.contributor.authorid920315202
pu.date.classyear2026
pu.departmentComputer Science

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
tl1719_written_final_report.pdf
Size:
20.69 MB
Format:
Adobe Portable Document Format
Download

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
100 B
Format:
Item-specific license agreed to upon submission
Description:
Download