15 papers
CPath: Class-Conditional Pathway Decoupling for Vision-Language Incremental Object Detection
Lecheng Xu, Feifei Shao, Ouyangzi Ye +5
Incremental Object Detection (IOD) aims to enable detectors to continuously learn novel categories while preserving previously acquired knowledge. However, existing methods suffer…
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Jinsheng Quan, Jianhua Li, Siyi Xie +7
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approa…
PromptPath: Prompt-Adaptive Computational Pathways for In-Context Learning
Hangrui Zhang, Feifei Shao, Yawei Luo +6
In-context learning (ICL) has attracted increasing attention for enabling models to perform new tasks using only a few ``input--output'' prompt examples. However, existing approach…
CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging
Jie Cao, Zhenxuan Fan, Zhuonan Wang +8
Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEFT). However, existing PEFT metho…
RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control
Youcan Xu, Jiaxin Shi, Zhen Wang +5
Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcas…
GateMOT: Q-Gated Attention for Dense Object Tracking
Mingjin Lv, Zelin Liu, Feifei Shao +4
While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tracking: its quadratic all-to…