About Me: I’m an Assistant Professor of Computer Science at Johns Hopkins, and a part-time Member of Technical Staff at Abridge. Previously, I was a postdoc at Carnegie Mellon University with Zack Lipton, and obtained my PhD in Computer Science at MIT with David Sontag.
News
- Sep 2026 Two papers accepted at NeurIPS 2026: “Fixed Size Active Statistical Inference” with Erik Skalnes, and “Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation” with Daniel Jeong and colleagues in the Evaluation and Datasets Track.
- Jul 2026 Jonathan Zhang is heading to Amsterdam to present “Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness” at UAI 2026.
- Jul 2026 Pranav Mani presents “No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference” at ICML 2026.
- Jun 2026 Andrew Wang presents “Revisiting Performance Claims for Chest X-Ray Models Using Clinical Context”, joint work with Jiashuo Zhang, at CHIL 2026.
Research Overview
My group develops methods for principled and efficient evaluation and monitoring of AI systems, with an emphasis on applications in healthcare. We draw on a broad methodological toolkit spanning statistics, causal inference, and machine learning, and collaborate closely with clinical researchers in medicine, nursing, and public health. Our work can be seen as covering three areas, detailed below. A few representative papers are listed under each; see the Papers page for the full list.
Methods for scalable, automated assessment of AI systems
It is often challenging to scalably define what “good” looks like for a generative AI system in a high-expertise, non-verifiable domain like healthcare. Our group develops new methods and interrogates existing methods for trying to do so at scale, leveraging e.g., existing clinician documentation (in the case of e.g., radiology report generation), or expert-written rubrics.
Evaluating Rubric Generation with Interventional Transfer
preprintcode
Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation
Neural Information Processing Systems, Evaluation and Datasets Track (NeurIPS), 2026
paper
Methods for efficient statistical evaluation of AI systems
Even when it is clear what a “good” output looks like, determining the “ground truth” performance of a system often requires time-intensive expert annotation. Our group develops statistical methods that make the most of limited labeled data, by e.g., carefully selecting a small number of expert labels in combination with LLM-as-judge or other automated measures of performance, while still providing valid statistical guarantees.
Fixed Size Active Statistical Inference
Neural Information Processing Systems (NeurIPS), 2026
No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference
International Conference on Machine Learning (ICML), 2026
paper
Methods for causal evaluation of evolving AI systems
Even when an AI system scores well according to retrospective evaluations, it may fail to improve outcomes when deployed in the wild. Randomized evaluations (e.g., RCTs / AB testing) are an important tool for establishing these effects, but running randomized trials for every update to an AI system is not feasible. With that in mind, our group develops causal inference methods for assessing the real-world impact of AI systems, including their effect on downstream decisions and how to validate them on an ongoing basis after deployment.
Bounding the Causal Impact of ML-assisted Decision-Making via Counterfactual Correctness
Conference on Uncertainty in Artificial Intelligence (UAI), 2026
paper
Just Trial Once: Ongoing Causal Validation of Machine Learning Models
Conference on Uncertainty in Artificial Intelligence (UAI), 2025
Oral Presentation (3% of submissions, 9% of accepted papers)
paperposter
Research Group
Meet the current members of the group and our collaborators on the People page. If you are interested in working with me as a PhD student or postdoc, please see this page for more information.