activity
20242026
collaborators

10 papers

cs.CV2026

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

Ziyu Ma, Shidong Yang, Yuxiang Ji +5

High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when the image is processed as a w…

cs.CL2026

Learning Agentic Policy from Action Guidance

Yuxiang Ji, Zengbin Wang, Yong Wang +6

Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training signals emerge only within its…

cs.LG2026

Tree Search for LLM Agent Reinforcement Learning

Yuxiang Ji, Ziyu Ma, Yong Wang +3

Recent advances in reinforcement learning (RL) have significantly enhanced the agentic capabilities of large language models (LLMs). In long-term and multi-turn agent tasks, existi…

cs.CV2026

Thinking with Map: Reinforced Parallel Map-Augmented Agent for Geolocalization

Yuxiang Ji, Yong Wang, Ziyu Ma +6

The image geolocalization task aims to predict the location where an image was taken anywhere on Earth using visual clues. Existing large vision-language model (LVLM) approaches le…

cs.CV2025

Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability

Boyong He, Yuxiang Ji, Zhuoyue Tan +1

Detectors often suffer from performance drop due to domain gap between training and testing data. Recent methods explore diffusion models applied to domain generalization (DG) and…

cs.CV2025

VisLanding: Monocular 3D Perception for UAV Safe Landing via Depth-Normal Synergy

Zhuoyue Tan, Boyong He, Yuxiang Ji +1

This paper presents VisLanding, a monocular 3D perception-based framework for safe UAV (Unmanned Aerial Vehicle) landing. Addressing the core challenge of autonomous UAV landing in…