AI research publications

Research publications by Wayner Barrios, from multimodal reasoning and video understanding to visual grounding, medical AI, and efficient inference. Find papers, summaries, and links to the open source implementations.

Questions behind the research

These publications examine how models connect visual observations with language and reasoning. My work on video asks how a language query can identify a relevant moment in a long sequence. The visual grounding work studies which information an instruction should draw from an image. More recent evaluation work asks whether an explanation contains the steps needed to justify its final answer.

For a closer look at the methods and their scope, read the MoDA visual grounding guide and the CRYSTAL reasoning guide. The vLLM-MLX inference guide follows the systems side of this research, explaining batching, caching, and how to interpret the reported Apple Silicon benchmarks. Each guide links the original paper and implementation; the publication list below also includes earlier work and collaborations in medical AI.

Publications

Research thatruns.

Earlier publications 4 papers
WACVW2019

Minding the Gaps in a Video Action Analysis Pipeline

J. Chen, J. Liu, J. Liang, T. Y. Hu, W. Ke, Wayner Barrios, D. Huang, A. G. Hauptmann

Analyzes the training and testing gaps between proposal, classification, and localization modules in event detection pipelines, then introduces practical fixes.

Complete publication recordGoogle Scholar

Building Wiqonn and researching multimodal perception, video understanding, and efficient inference.

Get in touch