3 papers
cs.CV2026
Video Models Can Reason with Verifiable Rewards
Tinghui Zhu, Sheng Zhang, James Y. Huang +5
Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable re…
cs.CV2026
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
Junbin Xiao, Shenglang Zhang, Pengxiang Zhu +1
We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the came…
cs.CL2025
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
James Y. Huang, Wenxuan Zhou, Nan Xu +5
The ability of Large Language Models (LLMs) to generate structured outputs that follow arbitrary schemas is crucial to a wide range of downstream tasks that require diverse structu…