Assistant Professor, Computer Science & Laboratory Medicine and Pathobiology
Canada CIFAR AI Chair at the Vector Institute
Tier II Canada Research Chair in Computational Medicine
Research Direction
Data collected from natural phenomena characterize computation occurring within complex processes governed by known and unknown laws. Neural networks are powerful tools that can learn to compress computation happening in nature. My group advances fundamental research in machine learning to understand, identify, and create controllable and reliable artificial intelligence systems.
Research Threads
- Causal machine learning. Natural phenomena follow laws, physics among them, that explain cause and effect. Humans generalize remarkably rapidly because we can manipulate or intuit those laws. Teaching neural networks how to do the same is critical to abduction and generalization, yet traditional tools for causality have resisted the bitter lesson. Our group has built and scaled the first causal foundation models (e.g. CausalPFN, IV-ICL and SurvivalPFN) which enable zero-shot prediction of causal effects, bounds on causal effects, and survival outcomes via in-context learning. [1] [2] [3] [4] [5]
- Generative models. Generative models compress the computation in nature directly. The more efficiently, in terms of FLOPs, we can compress that computation, the faster we can scale over data and parameters. We are interested in understanding how to minimize the distance between energy (via compute) and bits using generative models like MDM-Prime-v2 Deep Markov Models [1] [2] [3] [4]
- Large language models. Large language models compress computation expressed in natural language text. By prompting we can index into this computation. But to use LLMs reliably in health and biology, we need the indexed computation to mimic human intent. We also need new ways for LLMs to quickly acquire domain knowledge and ensure their outputs can be trusted in high-stakes settings. We develop methods like Contextual Fine-Tuning that teach LLMs how to learn quickly from new corpora, tools like AutoElicit to help experimentalists use LLMs as proxies for expert opinions, agentic systems AMG-RAG for question answering when knowledge evolves, and study causal methods to the development and evaluation of language models. [1] [2]
- Decision making in healthcare and biology. Solving decision-making problems in health at scale requires tools that can extract signal from large images, (HIPT), from long time series (SurF) and from tabular data (InterpreTabNet), and composing them correctly (CoMET). Deployment in medicine also requires creating tools to detect model failures like Detectron, and D3M. [1] [2] [3] [4] [5]
Students
PhD Students
- Vahid Balazadeh-Meresht
- Ethan Choi
- Viet Nguyen
- Lance Chao
- Yeongbin Seo
- Michael Cooper (co-sup w/ Mike Brudno)
- Jerry Ji (co-sup w/ Anna Goldenberg)
- Yujia Ma (co-sup w/ Chris Maddison)
- Steven Palayew (co-sup w/ Mike Wainberg)
- Mohammad Adnan (PhD student at the University of Calgary sup. w/ Yani Ioannou)
Postdoctoral Fellows
MSc Students
Links
Selected Publications
A selected list of representative papers is available below. For a full list, see my Google Scholar profile.
-
SurvivalPFN: Amortizing Survival Prediction via In-Context Bayesian InferencearXiv preprint arXiv:2605.15488, 2026🏆 Best Paper Award (Foundation Models for Structured Data, ICML 2026)
-
SurF: A Generative Model for Multivariate Irregular Time Series ForecastingarXiv preprint arXiv:2605.14069, 2026
-
IV-ICL: Bounding Causal Effects with Instrumental Variables via In-Context LearningarXiv preprint arXiv:2605.12924, 2026
-
Causal methods for LLM development and evaluationACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2026
-
SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingInternational Conference on Machine Learning (ICML), 2026
-
Frequentist Consistency of Prior-Data Fitted Networks for Causal InferenceInternational Conference on Machine Learning (ICML), 2026
-
Mitigating Privacy Risk via Forget Set-Free UnlearningInternational Conference on Learning Representations (ICLR), 2026
-
Can we generate portable representations for clinical time series data using LLMs?International Conference on Learning Representations (ICLR), 2026
-
CausalPFN: Amortized Causal Effect Estimation via In-Context LearningNeural Information Processing Systems (NeurIPS), 2025(Spotlight)
-
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking.Neural Information Processing Systems (NeurIPS), 2025
-
Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models .International Conference on Computer Vision (ICCV), 2025
-
Reliably Detecting Model Failures in Deployment Without LabelsNeural Information Processing Systems (NeurIPS), 2025
-
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical KnowledgeFindings of ACL 2025
-
Sparse Training from Random Initialization: Aligning Lottery Ticket Masks using Weight SymmetryInternational Conference on Machine Learning (ICML) 2025
-
AutoElicit: Using Large Language Models for Expert Prior Elicitation in Predictive ModellingInternational Conference on Machine Learning (ICML) 2025
-
Diverse Prototypical Ensembles Improve Robustness to Subpopulation ShiftInternational Conference on Machine Learning (ICML) 2025
-
ExOSITO: Explainable Off-Policy Learning with Side Information for Intensive Care Unit Blood Test OrdersConference on Health, Inference and Learning (CHIL), 2025
-
Teaching LLMs How to Learn with Contextual Fine-TuningInternational Conference on Learning Representations (ICLR), 2025
-
End-To-End Causal Effect Estimation from Unstructured Natural Language DataNeural Information Processing Systems (NeurIPS), 2024
-
Sequential Decision Making with Expert Demonstrations under Unobserved HeterogeneityNeural Information Processing Systems (NeurIPS), 2024 ( *: equal contribution )
-
Long-Term Allograft Survival in Liver Transplant RecipientsMachine Learning for Healthcare (MLHC), 2024
-
NeRF-US: Removing Ultrasound Imaging Artifacts from Neural Radiance Fields in the WildMachine Learning for Healthcare (MLHC), 2024
-
InterpreTabNet: Distilling Predictive Signals from Tabular Data by Salient Feature InterpretationInternational Conference on Machine Learning (ICML), 2024(Spotlight)
-
A Geometric Explanation of the Likelihood OOD Detection ParadoxInternational Conference on Machine Learning (ICML), 2024
-
Structured Neural Networks for Density Estimation and Causal InferenceNeural Information Processing Systems (NeurIPS), 2023
-
Copula-Based Deep Survival Models for Dependent CensoringUncertainty in Artificial Intelligence (UAI), 2023
-
DuETT: Dual Event Time Transformer for Electronic Health RecordsMachine Learning for Healthcare (MLHC), 2023
-
Machine learning in computational histopathology: Challenges and OpportunitiesGenes Cells and Chromosomes, 2023
-
A Learning Based Hypothesis Test for Harmful Covariate ShiftInternational Conference on Learning Representations (ICLR), 2023
-
Anamnesic Neural Differential Equations with Orthogonal Polynomial ProjectionsInternational Conference on Learning Representations (ICLR), 2023
-
Partial Identification with Implicit Generative ModelsNeural Information Processing Systems (NeurIPS), 2022(Spotlight)
-
HiCu: Leveraging Hierarchy for Curriculum Learning in Automated ICD CodingMachine Learning for Healthcare (MLHC), 2022
-
Large Images as Long Documents: Hierarchical ViTs with Self-Supervised Pretraining in Gigapixel Image PyramidsComputer Vision and Pattern Recognition (CVPR), 2022(Oral)
-
Hierarchical Optimal Transport for Comparing Histopathology DatasetsMedical Imaging with Deep Learning (MIDL), 2022
-
Using Time-Series Privileged Information for Provably Efficient Learning of Prediction ModelsArtificial Intelligence and Statistics (AISTATS), 2022
-
Clustering Interval-Censored Time-Series for Disease PhenotypingAssociation for the Advancement of Artificial Intelligence (AAAI), 2022
-
Mitigating bias in estimating epidemic severity due to heterogeneity of epidemic on-set and data aggregationAnnals of Epidemiology (In Press), 2021
-
Neural Pharmacodynamic State Space ModelingInternational Conference on Machine Learning (ICML), 2021
-
Max-Margin learning with the Bayes factorUncertainty in Artificial Intelligence (UAI), 2018
-
Representation Learning Approaches to Detect False Arrhythmia Alarms from ECG DynamicsMachine Learning for Healthcare (MLHC), 2018
-
Variational Autoencoders for Collaborative FilteringWorld Wide Web Conference (WWW), 2018
-
On the challenges of learning with inference networks on sparse, high-dimensional dataArtificial Intelligence and Statistics (AISTATS), 2018
-
Structured Inference Networks for Nonlinear State Space ModelsAssociation for the Advancement of Artificial Intelligence (AAAI), 2017(Oral)
-
Barrier Frank-Wolfe for Marginal InferenceNeural Information Processing Systems (NeurIPS), 2015