10 citations · 12 across the 20 of their papers we have counts for
12 papers · 1 filter
Detector-in-the-Loop Tracking: Active Memory Rectification for Stable Glottic Opening Localization
Huayu Wang, Bahaa Alattar, Cheng-Yen Yang +5
Temporal stability in glottic opening localization remains challenging due to the complementary weaknesses of single-frame detectors and foundation-model trackers: the former lacks…
Reasoning Matters for 3D Visual Grounding
Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai +3
The recent development of Large Language Models (LLMs) with strong reasoning ability has driven research in various domains such as mathematics, coding, and scientific discovery. M…
Warehouse Spatial Question Answering with LLM Agent
Hsiang-Wei Huang, Jen-Hao Cheng, Kuang-Ming Chen +8
Spatial understanding has been a challenging task for existing Multi-modal Large Language Models~(MLLMs). Previous methods leverage large-scale MLLM finetuning to enhance MLLM's sp…
ToSA: Token Merging with Spatial Awareness
Hsiang-Wei Huang, Wenhao Chai, Kuang-Ming Chen +2
Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual t…
Adapting SAM 2 for Visual Object Tracking: 1st Place Solution for MMVPR Challenge Multi-Modal Tracking
Cheng-Yen Yang, Hsiang-Wei Huang, Pyong-Kun Kim +5
We present an effective approach for adapting the Segment Anything Model 2 (SAM2) to the Visual Object Tracking (VOT) task. Our method leverages the powerful pre-trained capabiliti…
TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
Jen-Hao Cheng, Vivian Wang, Huayu Wang +11
Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. Existing methods either compress vid…