1 citations · 1 across the 9 of their papers we have counts for
11 papers
SAGE-Yoga: Multi-Cue Learning for Yoga Pose Classification and Joint-Level Correction
Hung Le Chi, Khanh Minh Huynh, Long Nghia Tran Pham +3
Automated yoga analysis requires both accurate pose classification and interpretable feedback on pose execution. However, existing methods often rely on a single visual prediction,…
Action-Aligned Retrieval with Pairwise Multimodal Reranking for Text-Based Person Anomaly Search
Thanh-Khoi Nguyen, Thanh-Nhan Vo, Trong-Thuan Nguyen +1
Text-based person anomaly search requires distinguishing individuals based on fine-grained, context-dependent behaviors rather than mere appearance. Existing methods struggle to ca…
TreeSoc: Tree-Structured Dynamic Reasoning and Tool Synergy for Soccer Video Understanding
Thanh-Nhan Vo, Thanh-Khoi Nguyen, Trong-Thuan Nguyen +2
Automated understanding of complex soccer scenarios from video remains a significant challenge for contemporary vision-language models (VLMs), which suffer from shallow cross-modal…
SAGA: Stable Acceleration Guidance for Autoregressive Video Generation
Thanh-Nhan Vo, Trong-Thuan Nguyen, Trung-Hoang Le +2
Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors…
LOGOS: Language-guided Oriented Object Detection in Aerial Scenes
Trong-Thuan Nguyen, Minh-Triet Tran
Object detection in geospatial scenes, such as satellite and aerial imagery, poses significant challenges due to the varying orientations and densities of objects, as well as the c…
SoccerNet 2026 Challenges Results
Anthony Cioppa, Silvio Giancola, Håkan Ardö +102
The SoccerNet 2026 Challenges constitute the sixth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in sports video underst…