7 papers
STDec: Spatio-Temporal Stability Guided Decoding for dLLMs
Yuzhe Chen, Jiale Cao, Xuyang Liu +3
Diffusion Large Language Models (dLLMs) have achieved rapid progress, viewed as a promising alternative to the autoregressive paradigm. However, most dLLM decoders still adopt a gl…
SNNSIR: A Simple Spiking Neural Network for Stereo Image Restoration
Ronghua Xu, Jin Xie, Jing Nie +2
Spiking Neural Networks (SNNs), characterized by discrete binary activations, offer high computational efficiency and low energy consumption, making them well-suited for computatio…
Multi-Granularity Language-Guided Training for Multi-Object Tracking
Yuhao Li, Jiale Cao, Muzammal Naseer +4
Most existing multi-object tracking methods typically learn visual tracking features via maximizing dis-similarities of different instances and minimizing similarities of the same…
SSLFusion: Scale & Space Aligned Latent Fusion Model for Multimodal 3D Object Detection
Bonan Ding, Jin Xie, Jing Nie +1
Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial i…
CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation
Lin Sun, Jiale Cao, Jin Xie +2
Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level…
iSeg: An Iterative Refinement-based Framework for Training-free Segmentation
Lin Sun, Jiale Cao, Jin Xie +2
Stable diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers hav…