activity
20232026
most citedSAMURAI: Adapting Segment Anything Model for Zero-Shot Visual Tracking with Motion-Aware Memory

10 citations · 12 across the 20 of their papers we have counts for

collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Detector-in-the-Loop Tracking: Active Memory Rectification for Stable Glottic Opening Localization

Huayu Wang, Bahaa Alattar, Cheng-Yen Yang +5

Temporal stability in glottic opening localization remains challenging due to the complementary weaknesses of single-frame detectors and foundation-model trackers: the former lacks…

cs.CV2026

Reasoning Matters for 3D Visual Grounding

Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai +3

The recent development of Large Language Models (LLMs) with strong reasoning ability has driven research in various domains such as mathematics, coding, and scientific discovery. M…

cs.CV2025

Warehouse Spatial Question Answering with LLM Agent

Hsiang-Wei Huang, Jen-Hao Cheng, Kuang-Ming Chen +8

Spatial understanding has been a challenging task for existing Multi-modal Large Language Models~(MLLMs). Previous methods leverage large-scale MLLM finetuning to enhance MLLM's sp…

cs.CV2025

ToSA: Token Merging with Spatial Awareness

Hsiang-Wei Huang, Wenhao Chai, Kuang-Ming Chen +2

Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual t…

cs.CV2025

Adapting SAM 2 for Visual Object Tracking: 1st Place Solution for MMVPR Challenge Multi-Modal Tracking

Cheng-Yen Yang, Hsiang-Wei Huang, Pyong-Kun Kim +5

We present an effective approach for adapting the Segment Anything Model 2 (SAM2) to the Visual Object Tracking (VOT) task. Our method leverages the powerful pre-trained capabiliti…

cs.CV2025

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Jen-Hao Cheng, Vivian Wang, Huayu Wang +11

Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. Existing methods either compress vid…