8 papers
STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation
Syed Ariff Syed Hesham, Yun Liu, Guolei Sun +4
Video reasoning segmentation demands pixel-accurate object tracking across hundreds of frames under complex natural language queries, producing dense spatiotemporal tokens whose qu…
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
Henghui Ding, Chang Liu, Shuting He +2
Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) g…
Evaluating SAM2 for Video Semantic Segmentation
Syed Hesham Syed Ariff, Yun Liu, Guolei Sun +4
The Segmentation Anything Model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object…
Open-set Anomaly Segmentation in Complex Scenarios
Song Xia, Yi Yu, Henghui Ding +4
Precise segmentation of out-of-distribution (OoD) objects, herein referred to as anomalies, is crucial for the reliable deployment of semantic segmentation models in open-set, safe…
Transferable Adversarial Attacks on SAM and Its Downstream Models
Song Xia, Wenhan Yang, Yi Yu +4
The utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical…
Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference Systems
Song Xia, Yi Yu, Wenhan Yang +5
By locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data…