I work on representation learning for biological sequences. In
particular, my past work has focused on foundation models for
mRNA (Orthrus, mRNABench)
and I am currently working on foundation models for
DNA with MSR. My research leverages
knowledge of evolutionary structure to design scalable pretraining objectives
that are both sample efficient and highly predictive of downstream biological properties.
Selected Work
Orthrus: toward evolutionary and functional RNA foundation models
Philip Fradkin*, Ruian Shi*, Taykhoom Dalal*,
Keren Isaev, Brendan J. Frey, Leo J. Lee, Quaid Morris, Bo Wang
Nature Methods, 2026
Orthrus leverages splice isoforms and orthology alongside contrastive
learning to pretrain a state-of-the-art mRNA foundation model, which
is broadly predctive of a wide range of mRNA properties and functions.
It is extremely sample and parameter efficient, and can group mRNAs by
function in its latent space, the first such demonstration (to our knowledge).
mRNABench: a curated benchmark for mature mRNA property and function prediction
Ruian Shi*, Taykhoom Dalal*, Philip Fradkin*,
Divya Koyyalagunta, et al.
bioRxiv, 2025
mRNABench evaluates 75 nucleotide foundation models on 79
mature-mRNA prediction tasks, the only benchmark of its kind.
We analyze results through the lens of parameter scaling, pretraining
objectives, compressibility of genomic regions, and data splitting strategies.
PYPE: a pipeline for phenome-wide association and Mendelian randomization in investigator-driven biobank scale analysis
Taykhoom Dalal, Chirag J. Patel
Cell Patterns, 2024
PYPE runs phenome-wide association and Mendelian randomization
analyses, annotates variants and genes, and generates standard plots
from biobank-scale data.