3 papers
cs.AI2025
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
Junzhe Chen, Tianshu Zhang, Shiyu Huang +6
Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time enviro…
cs.CV2025
Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
Runwei Guan, Ningwei Ouyang, Tianhao Xu +10
Automated waterway environment perception is crucial for enabling unmanned surface vessels (USVs) to understand their surroundings and make informed decisions. Most existing waterw…
cs.AI2025
Building a Stable Planner: An Extended Finite State Machine Based Planning Module for Mobile GUI Agent
Fanglin Mo, Junzhe Chen, Haoxuan Zhu +1
Mobile GUI agents execute user commands by directly interacting with the graphical user interface (GUI) of mobile devices, demonstrating significant potential to enhance user conve…