Publications (8)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Ziming Cheng, Binrui Xu, Lisheng Gong +14
With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously…
Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions
Ziming Cheng, Zhiyuan Huang, Junting Pan +2
Graphical user interfaces (GUI) automation agents are emerging as powerful tools, enabling humans to accomplish increasingly complex tasks on smart devices. However, users often in…
MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences
Qihao Wang, Ziming Cheng, Shuo Zhang +12
While autonomous software engineering (SWE) agents are reshaping programming paradigms, they currently suffer from a "closed-world" limitation: they attempt to fix bugs from scratc…
SpiritSight Agent: Advanced GUI Agent with One Look
Zhiyuan Huang, Ziming Cheng, Junting Pan +2
Graphical User Interface (GUI) agents show amazing abilities in assisting human-computer interaction, automating human user's navigation on digital devices. An ideal GUI agent is e…
Radiative cooling capacity on Earth
Cunhai Wang, Hao Chen, Yanyan Feng +3
By passively dissipating thermal emission into the ultracold deep space, radiative cooling (RC) is an environment-friendly means for gaining cooling capacity, paving a bright futur…
Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training
Yan Zeng, Wangchunshu Zhou, Ao Luo +2
In this paper, we introduce Cross-View Language Modeling, a simple and effective pre-training framework that unifies cross-lingual and cross-modal pre-training with shared architec…