activity
20242026
collaborators

13 papers

cs.CV2026

Detector-in-the-Loop Tracking: Active Memory Rectification for Stable Glottic Opening Localization

Huayu Wang, Bahaa Alattar, Cheng-Yen Yang +5

Temporal stability in glottic opening localization remains challenging due to the complementary weaknesses of single-frame detectors and foundation-model trackers: the former lacks…

cs.CV2026

Reasoning Matters for 3D Visual Grounding

Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai +3

The recent development of Large Language Models (LLMs) with strong reasoning ability has driven research in various domains such as mathematics, coding, and scientific discovery. M…

cs.CV2025

UniHPR: Unified Human Pose Representation via Singular Value Contrastive Learning

Zhongyu Jiang, Wenhao Chai, Lei Li +3

In recent years, there has been a growing interest in developing effective alignment pipelines to generate unified representations from different modalities for multi-modal fusion…

cs.CV2025

Warehouse Spatial Question Answering with LLM Agent

Hsiang-Wei Huang, Jen-Hao Cheng, Kuang-Ming Chen +8

Spatial understanding has been a challenging task for existing Multi-modal Large Language Models~(MLLMs). Previous methods leverage large-scale MLLM finetuning to enhance MLLM's sp…

cs.CV2025

ToSA: Token Merging with Spatial Awareness

Hsiang-Wei Huang, Wenhao Chai, Kuang-Ming Chen +2

Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual t…

cs.LG2025

Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache Compression

Kunjun Li, Zigeng Chen, Cheng-Yen Yang +1

Visual Autoregressive (VAR) modeling has garnered significant attention for its innovative next-scale prediction approach, which yields substantial improvements in efficiency, scal…