4 papers
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
Qi Wang, Yanrui Yu, Ye Yuan +2
Reinforcement fine-tuning (RFT) has shown great promise in achieving humanlevel reasoning capabilities of Large Language Models (LLMs), and has recently been extended to MLLMs. Nev…
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
Yanrui Yu, Tianfei Zhou, Jiaxin Sun +4
In modern urban environments, camera networks generate massive amounts of operational footage -- reaching petabytes each day -- making scalable video analytics essential for effici…
Prompt-Driven Continual Graph Learning
Qi Wang, Tianfei Zhou, Ye Yuan +1
Continual Graph Learning (CGL), which aims to accommodate new tasks over evolving graph data without forgetting prior knowledge, is garnering significant research interest. Mainstr…
Image Segmentation in Foundation Model Era: A Survey
Tianfei Zhou, Wang Xia, Fei Zhang +5
Image segmentation is a long-standing challenge in computer vision, studied continuously over several decades, as evidenced by seminal algorithms such as N-Cut, FCN, and MaskFormer…