collaborators

6 papers

cs.CV2026

ProTPS: Prototype-Guided Text Prompt Selection for Continual Learning

Jie Mei, Li-Leng Peng, Keith Fuller +1

For continual learning, text-prompt-based methods leverage text encoders and learnable prompts to encode semantic features for sequentially arrived classes over time. A common chal…

cs.CV2025

ToSA: Token Merging with Spatial Awareness

Hsiang-Wei Huang, Wenhao Chai, Kuang-Ming Chen +2

Token merging has emerged as an effective strategy to accelerate Vision Transformers (ViT) by reducing computational costs. However, existing methods primarily rely on the visual t…

cs.CV2025

TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Jen-Hao Cheng, Vivian Wang, Huayu Wang +11

Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. Existing methods either compress vid…

cs.CV2025

CityGen: Infinite and Controllable City Layout Generation

Jie Deng, Wenhao Chai, Jianshu Guo +6

The recent surge in interest in city layout generation underscores its significance in urban planning and smart city development. The task involves procedurally or automatically ge…

cs.CV2025

MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object Detection

Hou-I Liu, Christine Wu, Jen-Hao Cheng +8

Monocular 3D object detection (Mono3D) holds noteworthy promise for autonomous driving applications owing to the cost-effectiveness and rich visual context of monocular camera sens…

cs.CV2025

Pointmap Association and Piecewise-Plane Constraint for Consistent and Compact 3D Gaussian Segmentation Field

Wenhao Hu, Wenhao Chai, Shengyu Hao +4

Achieving a consistent and compact 3D segmentation field is crucial for maintaining semantic coherence across views and accurately representing scene structures. Previous 3D scene…