9 papers
AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture
Zi Ye, Yibin Wen, Xiaoya Fan +10
Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multi…
AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
Yibin Wen, Qingmei Li, Zi Ye +14
Recent advancements in Vision-Language Models (VLMs) have significantly impacted various industries. In agriculture, these multimodal capabilities hold great promise for applicatio…
ConsistNav: Closing the Action Consistency Gap in Zero-Shot Object Navigation with Semantic Executive Control
Haosen Wang, Zhenyang Li, Yinqiang Zhang +9
Zero-shot object navigation has advanced rapidly with open-vocabulary detectors, image--text models, and language-guided exploration. However, even after current methods detect a p…
Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework
Xinyu Zhang, Zurong Mai, Qingmei Li +14
Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evidence beyond the visible range remains la…
GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset
Zhiwei Zhang, Zi Ye, Yibin Wen +4
Agricultural parcels serve as basic units for conducting agricultural practices and applications, which is vital for land ownership registration, food security assessment, soil ero…
GALA: A GlobAl-LocAl Approach for Multi-Source Active Domain Adaptation
Juepeng Zheng, Peifeng Zhang, Yibin Wen +3
Domain Adaptation (DA) provides an effective way to tackle target-domain tasks by leveraging knowledge learned from source domains. Recent studies have extended this paradigm to Mu…