3 papers
cs.LG2026
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Xiongbin Wu, Zhihao Luo, Shanzhe Lei +7
Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception…
cs.RO2026
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
Zhihao Luo, Wentao Yan, Jingyu Gong +5
Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and tra…
cs.CV2025
Synthetic-to-Real Camouflaged Object Detection
Zhihao Luo, Luojun Lin, Zheng Lin
Due to the high cost of collection and labeling, there are relatively few datasets for camouflaged object detection (COD). In particular, for certain specialized categories, the av…