I develop models that connect perception with higher-level understanding. My past work focused on video understanding and visual perception, and I'm now expanding into multimodal models that follow events through time, ground language in visual evidence, and learn coherent reasoning strategies.
I recently completed my Ph.D. in Computer Science at Dartmouth College. I work across video understanding, visual grounding, reinforcement-learning post-training, and transparent reasoning evaluation.
I'm also interested in high-performance AI infrastructure and efficient inference systems, including continuous batching, multimodal prefix caching, and native model serving on Apple Silicon that reaches hundreds of tokens per second.
My work has appeared at CVPR, ICCV, ECCV, ICML, WACV, and other leading venues, as well as in PLOS Digital Health, the Journal of Pathology Informatics, Frontiers in Medical Technology, and other journals.
I earned my Ph.D. at Dartmouth College, advised by SouYoung Jin, and founded Wiqonn, an AI lab growing ambitious AI work in Latin America.