5 papers
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Hongxi Yan, Ziyue Huang, Shichao Fan +1
Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual actions. Existing critic-free methods propag…
Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images
Shuai Yang, Ziyue Huang, Jiaxin Chen +2
Open-vocabulary object detection in remote sensing commonly relies on text-only prompting to specify target categories, implicitly assuming that inference-time category queries can…
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
Yongchao Feng, Yajie Liu, Shuai Yang +13
Vision-Language Model (VLM) have gained widespread adoption in Open-Vocabulary (OV) object detection and segmentation tasks. Despite they have shown promise on OV-related tasks, th…
A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
Ziyue Huang, Hongxi Yan, Qiqi Zhan +7
The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data int…
OpenRSD: Towards Open-prompts for Object Detection in Remote Sensing Images
Ziyue Huang, Yongchao Feng, Shuai Yang +3
Remote sensing object detection has made significant progress, but most studies still focus on closed-set detection, limiting generalization across diverse datasets. Open-vocabular…