4 papers
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
Xuanpu Zhao, Zhentao Tan, Dianmo Sheng +6
To enhance the perception and reasoning capabilities of multimodal large language models in complex visual scenes, recent research has introduced agent-based workflows. In these wo…
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
Zhiqian Zhang, Xu Zhao, Xiaoqing Xu +5
In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models stil…
NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results
Yuqian Fu, Xingyu Qiu, Bin Ren +59
Cross-Domain Few-Shot Object Detection (CD-FSOD) poses significant challenges to existing object detection and few-shot detection models when applied across domains. In conjunction…
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
Xiantao Hu, Ying Tai, Xu Zhao +5
Multimodal tracking has garnered widespread attention as a result of its ability to effectively address the inherent limitations of traditional RGB tracking. However, existing mult…