activity
20242026
collaborators

7 papers

cs.CL2026

STDec: Spatio-Temporal Stability Guided Decoding for dLLMs

Yuzhe Chen, Jiale Cao, Xuyang Liu +3

Diffusion Large Language Models (dLLMs) have achieved rapid progress, viewed as a promising alternative to the autoregressive paradigm. However, most dLLM decoders still adopt a gl…

cs.CV2025

SNNSIR: A Simple Spiking Neural Network for Stereo Image Restoration

Ronghua Xu, Jin Xie, Jing Nie +2

Spiking Neural Networks (SNNs), characterized by discrete binary activations, offer high computational efficiency and low energy consumption, making them well-suited for computatio…

cs.CV2025

Multi-Granularity Language-Guided Training for Multi-Object Tracking

Yuhao Li, Jiale Cao, Muzammal Naseer +4

Most existing multi-object tracking methods typically learn visual tracking features via maximizing dis-similarities of different instances and minimizing similarities of the same…

cs.CV2025

SSLFusion: Scale & Space Aligned Latent Fusion Model for Multimodal 3D Object Detection

Bonan Ding, Jin Xie, Jing Nie +1

Multimodal 3D object detection based on deep neural networks has indeed made significant progress. However, it still faces challenges due to the misalignment of scale and spatial i…

cs.CV2024

CLIPer: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic Segmentation

Lin Sun, Jiale Cao, Jin Xie +2

Contrastive Language-Image Pre-training (CLIP) exhibits strong zero-shot classification ability on various image-level tasks, leading to the research to adapt CLIP for pixel-level…

cs.CV2024

iSeg: An Iterative Refinement-based Framework for Training-free Segmentation

Lin Sun, Jiale Cao, Jin Xie +2

Stable diffusion has demonstrated strong image synthesis ability to given text descriptions, suggesting it to contain strong semantic clue for grouping objects. The researchers hav…