14 papers
ChronoVision: Temporal Reasoning via Latent State Reconstruction
Yifan Shen, Jian Xu, Boyi Li +6
Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stem…
FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision
Tongyan Wang, Zhengyuan Li, Muhan Lin +5
Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form description…
Decoding Children's Gait Behavior
Yifan Shen, Boyi Li, Meihuan Huang +12
We introduce a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video. We specifically target the ambulator…
The 1st AI Children Challenge
Boyi Li, Yifan Shen, Houze Yang +7
The First AI Children Challenge aims to advance real-world applications of computer vision and AI in child healthcare, child education, and pediatrics. The 2026 CV4CHL edition feat…
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
Tianjiao Yu, Xinzhuo Li, Yifan Shen +4
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existi…
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
Tianjiao Yu, Xinzhuo Li, Muntasir Wahed +4
Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most text-to-3D methods overlook the semantic…