Data efficiency
How can models learn meaningful biomedical representations when expert labels and patient data are limited?
Research
I’m drawn to questions where better machine learning can support better decisions for people. Biomedical data make that responsibility concrete: labels are limited, populations shift, errors carry different costs, and every useful result must stand up to careful scrutiny.
Purpose
I care about methods that do more than improve a benchmark. The deeper task is to understand when a model works, where it fails, and whether its behavior is clear enough to support responsible use. That makes data quality, evaluation, uncertainty, and reproducibility part of the research itself.
Current questions
How can models learn meaningful biomedical representations when expert labels and patient data are limited?
How can evaluation reveal whether a model will remain reliable across patients, devices, sites, and populations?
Can inspectable structure be built into learning itself, so explanations reflect how the model reaches a decision?
How can calibrated uncertainty help a model recognize when human review matters most?
Which leakage checks, stress tests, and subgroup analyses should be standard before a result is considered convincing?
Which representations can carry useful knowledge across signals, clinical variables, and eventually other biological modalities?
Emerging direction
I’m beginning to explore whether the same principles—data efficiency, interpretable structure, robust validation, and reusable representations—can extend to higher-dimensional biological and multimodal data. This is a developing research direction, rather than an established publication area.