7 papers
Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation
Artheme Gauthier-Villar, Guodong Ding, Angela Yao
Dataset condensation for action segmentation synthesizes compact, informative representations of long, untrimmed video datasets. The existing approach relies on Variational Autoenc…
LightAVSeg: Lightweight Audio-Visual Segmentation
Qing Zhong, Guodong Ding, Lingqiao Liu +3
Audio-Visual Segmentation (AVS) targets pixel level localization of sounding emitting objects in videos. However, existing models rely on dense cross-modal attention with quadratic…
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
Guodong Ding, Angela Yao
This paper tackles compositional personalization of vision-language models (VLMs). In this problem, multiple user-defined concepts must be recognized or described jointly at test t…
MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation
Gengliang Li, Rongyu Chen, Bin Li +2
Ensuring factual consistency and reliable reasoning remains a critical challenge for medical vision-language models. We introduce MEDFACT-R1, a two-stage framework that integrates…
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
Qing Zhong, Peng-Tao Jiang, Wen Wang +3
Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos.…
Condensing Action Segmentation Datasets via Generative Network Inversion
Guodong Ding, Rongyu Chen, Angela Yao
This work presents the first condensation approach for procedural video datasets used in temporal action segmentation. We propose a condensation framework that leverages generative…