collaborators

8 papers

cs.CV2026

BUS: Brain-Inspired Unsupervised Self-Reflection via Backward Prediction for Multimodal Reasoning

Jiacheng Yang, Tongying Xiao, Yunkai Dang +7

Current Vision-Language Models (VLMs) often struggle to handle complex visual tasks that require consistent and fine-grained reasoning. Recent methods aim to train models to facili…

cs.CV2026

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework

Jiacheng Yang, Anqi Chen, Yunkai Dang +5

Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens increases quadratically with resoluti…

cs.CV2026

VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning

Zhe Gao, Shiyu Shen, Taifeng Chai +7

Existing Multimodal Large Language Models (MLLMs) often suffer from hallucinations in long video understanding (LVU), primarily due to the imbalance between textual and visual toke…

cs.CV2026

Prompt-Free Universal Region Proposal Network

Qihong Tang, Changhan Liu, Shaofeng Zhang +3

Identifying potential objects is critical for object recognition and analysis across various computer vision applications. Existing methods typically localize potential objects by…

cs.CV2025

FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing

Yunkai Dang, Donghao Wang, Jiacheng Yang +7

Large vision-language models (VLMs) exhibit strong performance across various tasks. However, these VLMs encounter significant challenges when applied to the remote sensing domain…

cs.LG2025

LibContinual: A Comprehensive Library towards Realistic Continual Learning

Wenbin Li, Shangge Liu, Borui Kang +7

A fundamental challenge in Continual Learning (CL) is catastrophic forgetting, where adapting to new tasks degrades the performance on previous ones. While the field has evolved wi…