Title Information
Title
A Deep Learning Framework for Protein Expression Prediction from Histology Images
Type of Resource (primo)
dissertations
Name: Personal
Name Part
Liu, Ian Yiheng
Role
Role Term: Text
creator
Name: Personal
Name Part
Ma, Ying
Role
Role Term: Text
Advisor
Name: Personal
Name Part
Duan, Fenghai
Role
Role Term: Text
Reader
Name: Corporate
Name Part
Brown University. Department of Biostatistics
Role
Role Term: Text
sponsor
Origin Information
Copyright Date
2026
Physical Description
Extent
xv, 61 p.
digitalOrigin
born digital
Note: thesis
Thesis (Sc. M.)--Brown University, 2026
Genre (aat)
theses
Abstract
Hematoxylin and eosin (H&E) histology is widely available, but it captures morphology rather than molecular phenotype. Spatial proteomics measures tissue protein organization, yet is limited by cost, throughput, marker panels, and paired training data. This thesis tests whether deep-learning models can infer spot-level protein expression from H&E morphology and whether routing predictions through a learned protein latent space improves transfer to held-out tissue. Two public 10x Genomics Visium CytAssist FFPE human tonsil datasets were used. Models were trained and evaluated on 33 antibody targets shared by both datasets in two directions: train on one section and validate on the other, then reverse the roles. The proposed method is a two-stage latent alignment framework. Stage 1 learns modality-specific image and protein representations with ResNet-50 image and regularized protein variational autoencoders. Stage 2 maps histology embeddings into the learned protein latent space and decodes them into protein predictions. It was compared with a DeepPT-style direct-prediction baseline, ridge and CatBoost probes, and ablations of augmentation and image-encoder fine-tuning. Across pooled held-out transfers, the proposed model improved marker-level spatial association relative to DeepPT: Pearson correlation increased from 0.180 to 0.255, Spearman correlation from 0.172 to 0.258, and marker SSIM from 0.053 to 0.106. Gains were strongest in the forward transfer direction, while reverse-transfer performance and spot-level error were mixed. CatBoost achieved the lowest pooled spot RMSE, and ridge achieved the highest marker correlations. Ablations indicated that image-encoder fine-tuning was important for marker-level recovery, whereas augmentation improved section-shift robustness. The most recoverable proteins were associated with visible tissue architecture or immune organization, including KRT5, CXCR5, PAX5, CD3E, CD19, PDCD1, CD274, CEACAM8, and ACTA2. Predicted maps captured broad spatial structure but often smoothed local variation and compressed dynamic range. These findings support H\&E-derived virtual spatial proteomics as a marker-specific, section-sensitive tool for hypothesis generation and assay prioritization, not as a replacement for measured protein assays.
Subject
Topic
Machine Learning
Subject (fast) (authorityURI="http://id.worldcat.org/fast", valueURI="http://id.worldcat.org/fast/00957675")
Topic
Histology, Pathological
Subject
Topic
Computer Vision
Subject
Topic
Deep Learning
Subject
Topic
Representation Learning
Subject
Topic
spatial transcriptomics
Subject
Topic
spatial proteomics
Subject
Topic
histology images
Subject
Topic
computational pathology
Subject
Topic
spatial omics
Subject
Topic
multimodal learning
Subject
Topic
cross-modal alignment
Subject
Topic
variational autoencoder
Subject
Topic
morphology-to-omics prediction
Subject
Topic
digital twin
Language
Language Term (ISO639-2B)
English
Record Information
Record Content Source (marcorg)
RPB
Record Creation Date (encoding="iso8601")
20260516