6 papers
Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models
Minseok Kang, Hyunwoo Kim, Chanyoung Kim +3
Vision-language models (VLMs) have achieved remarkable generalization across diverse multimodal tasks through large-scale pre-training, yet their rapidly increasing computational a…
Mitigating State Aliasing in Vision-Language-Action Models via Inverse Dynamics Learning
Kyujin Lee, Injae Kim, Jihwan Park +3
Vision-Language-Action (VLA) models have emerged as a promising framework that unifies perception, reasoning, and control for robot manipulation by adapting pretrained vision-langu…
LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
Chanyoung Kim, Minwoo Kim, Minseok Kang +2
Vision-Language-Action (VLA) models achieve strong performance in robotic manipulation by leveraging pre-trained vision-language backbones. However, in downstream robotic settings,…
Visual Diversity and Region-aware Prompt Learning for Zero-shot HOI Detection
Chanhyeong Yang, Taehoon Song, Jihwan Park +1
Zero-shot Human-Object Interaction detection aims to localize humans and objects in an image and recognize their interaction, even when specific verb-object pairs are unseen during…
Super-class guided Transformer for Zero-Shot Attribute Classification
Sehyung Kim, Chanhyeong Yang, Jihwan Park +2
Attribute classification is crucial for identifying specific characteristics within image regions. Vision-Language Models (VLMs) have been effective in zero-shot tasks by leveragin…
SoccerNet 2024 Challenges Results
Anthony Cioppa, Silvio Giancola, Vladimir Somers +81
The SoccerNet 2024 challenges represent the fourth annual video understanding challenges organized by the SoccerNet team. These challenges aim to advance research across multiple t…