3 papers
cs.CV2026
Building a Precise Video Language with Human-AI Oversight
Zhiqiu Lin, Chancharik Mitra, Siyuan Cen +13
Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable ov…
cs.CV2025
Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision
Tianqin Li, George Liu, Tai Sing Lee
Despite remarkable progress in computer vision, modern recognition systems remain fundamentally limited by their dependence on rich, redundant visual inputs. In contrast, humans ca…
cs.CL2025
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
Cheng-Ting Chou, George Liu, Jessica Sun +4
Deterministically controlling the target generation language of large multilingual language models (LLMs) remains a fundamental challenge, particularly in zero-shot settings where…