collaborators

13 papers

cs.CV2026

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

Jinsheng Quan, Jianhua Li, Siyi Xie +7

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approa…

cs.CV2026

PromptPath: Prompt-Adaptive Computational Pathways for In-Context Learning

Hangrui Zhang, Feifei Shao, Yawei Luo +6

In-context learning (ICL) has attracted increasing attention for enabling models to perform new tasks using only a few ``input--output'' prompt examples. However, existing approach…

cs.CL2026

CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Merging

Jie Cao, Zhenxuan Fan, Zhuonan Wang +8

Large language models (LLMs) achieve remarkable performance on diverse downstream and domain-specific tasks via parameter-efficient fine-tuning (PEFT). However, existing PEFT metho…

cs.CV2026

RealCam: Real-Time Novel-View Video Generation with Interactive Camera Control

Youcan Xu, Jiaxin Shi, Zhen Wang +5

Camera-controlled video-to-video (V2V) generation enables dynamic viewpoint synthesis from monocular footage, holding immense potential for interactive filmmaking and live broadcas…

cs.CV2026

GateMOT: Q-Gated Attention for Dense Object Tracking

Mingjin Lv, Zelin Liu, Feifei Shao +4

While large models demonstrate the strong representational power of vanilla attention, this core mechanism cannot be directly applied to Dense Object Tracking: its quadratic all-to…

cs.CV2026

DVAR: Adversarial Multi-Agent Debate for Video Authenticity Detection

Hongyuan Qi, Feifei Shao, Ming Li +2

The rapid evolution of video generation technologies poses a significant challenge to media forensics, as conventional detection methods often fail to generalize beyond their train…