CRYSTAL
ECCV 2026 benchmark, dataset, and pip-installable metrics for evaluating ordered reasoning traces and training with process rewards.
Open source AI projects and contributions by Wayner Barrios: inference engines, research benchmarks, model training tools, and geospatial platforms. Read the research publications behind the models and benchmarks.
If you want to run language or multimodal models locally on a Mac, start with the vLLM-MLX inference guide. It explains the serving workflow, the role of continuous batching and caching, and the measurements that matter when comparing models on your own hardware. The repository remains the reference for installation and supported features.
If you are evaluating the steps in a model's explanation, explore CRYSTAL. Its benchmark and metrics separate final-answer correctness from the coverage and ordering of reasoning steps. For work on instruction-guided visual features, the MoDA guide introduces the adapter and points to the model-family-specific training and evaluation resources.
The remaining projects cover training workflows, video research, development tools, and contributions to geospatial software. Entries marked as contributions describe my work within a larger project. Their repository links lead to the upstream projects, where you can inspect the code, documentation, and development history. The research publications provide context for the implementations associated with papers.
Open source
My research and tools first, then contributions to open-source communities — video, geospatial, and LLM infrastructure.
ECCV 2026 benchmark, dataset, and pip-installable metrics for evaluating ordered reasoning traces and training with process rewards.
Official implementation of MoDA (ICML 2026) — instruction-guided channel modulation for fine-grained visual grounding in instructional MLLMs, with under 1% additional FLOPs.
High-throughput, OpenAI- and Anthropic-compatible LLM and MLLM inference for Apple Silicon, with continuous batching, tool calling, vision, audio, and video support.
Building Wiqonn and researching multimodal perception, video understanding, and efficient inference.
Get in touch