6 papers
FD: A Dedicated Framework for Fine-Grained Dataset Distillation
Hongxu Ma, Guang Li, Shijie Wang +5
Dataset distillation (DD) compresses a large training set into a small synthetic set, reducing storage and training cost, and has shown strong results on general benchmarks. Decoup…
Towards an Effective Action-Region Tracking Framework for Fine-grained Video Action Recognition
Baoli Sun, Yihan Wang, Xinzhu Ma +3
Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture…
Referring Video Object Segmentation with Cross-Modality Proxy Queries
Baoli Sun, Xinzhu Ma, Ning Wang +2
Referring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. Th…
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
Zhuqiang Lu, Zhenfei Yin, Mengwei He +4
Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual conten…
Propagating Sparse Depth via Depth Foundation Model for Out-of-Distribution Depth Completion
Shenglun Chen, Xinzhu Ma, Hong Zhang +2
Depth completion is a pivotal challenge in computer vision, aiming at reconstructing the dense depth map from a sparse one, typically with a paired RGB image. Existing learning bas…
Is a Pure Transformer Effective for Separated and Online Multi-Object Tracking?
Chongwei Liu, Haojie Li, Zhihui Wang +1
Recent advances in Multi-Object Tracking (MOT) have demonstrated significant success in short-term association within the separated tracking-by-detection online paradigm. However,…