9 papers
ChronoVision: Temporal Reasoning via Latent State Reconstruction
Yifan Shen, Jian Xu, Boyi Li +6
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stem…
Decoding Children's Gait Behavior
Yifan Shen, Boyi Li, Meihuan Huang +12
We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We specifically target the ambulator…
The 1st AI Children Challenge
Boyi Li, Yifan Shen, Houze Yang +7
The First AI Children Challenge aims to advance real-world applications of computer vision and AI in child healthcare, child education, and pediatrics. The 2026 CV4CHL edition feat…
: Unifying Generation and Self-Verification for Parallel Reasoners
Harman Singh, Xiuyu Li, Kusha Sareen +14
Test-time scaling for complex reasoning tasks shows that leveraging inference-time compute, by methods such as independently sampling and aggregating multiple solutions, results in…
DexImit: Learning Bimanual Dexterous Manipulation from Monocular Human Videos
Juncheng Mu, Sizhe Yang, Yiming Bao +6
Data scarcity fundamentally limits the generalization of bimanual dexterous manipulation, as real-world data collection for dexterous hands is expensive and labor-intensive. Human…
TADS: Task-Aware Data Selection for Multi-Task Multimodal Pre-Training
Guanjie Cheng, Boyi Li, Lingyu Sun +4
Large-scale multimodal pre-trained models like CLIP rely heavily on high-quality training data, yet raw web-crawled datasets are often noisy, misaligned, and redundant, leading to…