17 papers
SpikingMOT: A Spike-Driven Multi-Object Tracker
Yiding Sun, Xiangyang Yang, Dongxu Zhang +7
Multi-object tracking (MOT) plays a fundamental role in visual perception, where accurate trajectory prediction is essential for reliable target association under complex motion pa…
Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning
Leichao Dong, Dongxu Zhang, Yiding Sun +4
Large reasoning models often solve problems through long chain-of-thought (CoT) traces, yet much of this computation is spent on redundant derivations, repeated self-verification,…
SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models
Dongxu Zhang, Yiding Sun, Zihao Guo +5
Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may…
GaussFusion: Towards Multimodal 3D Gaussian Pretraining
Zhixuan You, Jihua Zhu, Yiding Sun +5
3D Gaussian Splatting provides an explicit representation that jointly models geometry and appearance, serving as a scalable foundation for 3D representation learning. Existing pre…
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Kai Tang, Jinhao You, Bohua Zhang +6
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain su…
Tri-Efficient Transfer Learning for Point Cloud Videos
Yiding Sun, Dongxu Zhang, Jihua Zhu +6
While point cloud foundation models have significantly advanced point cloud video understanding, existing parameter-efficient fine-tuning (PEFT) methods still suffer from two criti…