Publication: Multimodal Representation Learning between Histology and Spatial Transcriptomics
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
This thesis studies multimodal representation learning between histology and spatial transcriptomics (ST), with the goal of learning a shared latent space that aligns tissue morphology with spatial gene expression. A central challenge is that ST spots do not correspond exactly to histology patches, making strict one-to-one supervision biologically imperfect. To address this, we compare three patch-spot correspondence definitions: 1) a strict spot-centered baseline, 2) a grid-based center-only baseline, and 3) two neighborhood-aware grid variants, spot-kNN and distance-softmax. Histology patches are embedded with a frozen foundation encoder, while ST gene vectors are mapped into the shared space with a trainable multilayer perceptron. Across held-out multislide experiments, grid-based correspondence improves retrieval substantially over the strict spot-centered baseline, showing that broader local visual context is important for cross-slide alignment. Neighborhood-aware objectives further improve retrieval behavior when evaluated through rank, spatial locality, and gene-expression similarity, with the distance-softmax model showing the strongest overall neighborhood-level alignment. However, exact paired-spot retrieval remains weak across all settings, batching interventions do not meaningfully improve performance, and visually ambiguous histologic regions remain a major source of struggle. These results suggest that ST-histology alignment is better framed as a neighborhood-aware local compatibility problem than as strict point-to-point matching.