7 papers
A Task-State Representation for Long-Horizon Mobile GUI Agents
Yujie Zheng, Zikang Liu, Xin Zhao +1
While long-horizon mobile GUI agents typically rely on thought-action-observation loops, they struggle to separate persistent task states from transient screen observations. As exe…
Towards Long-horizon Agentic Multimodal Search
Yifan Du, Zikang Liu, Jinbiao Peng +5
Multimodal deep search agents have shown great potential in solving complex tasks by iteratively collecting textual and visual evidence. However, managing the heterogeneous informa…
Think Before Writing: Feature-Level Multi-Objective Optimization for Generative Citation Visibility
Zikang Liu, Peilan Xu
Generative answer engines expose content through selective citation rather than ranked retrieval, fundamentally altering how visibility is determined. This shift calls for new opti…
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
Yifan Li, Yukai Gu, Yingqian Min +6
Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation…
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
Zikang Liu, Junyi Li, Wayne Xin Zhao +3
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) promise human-like interaction with software applications, yet long-horizon tasks remain c…
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
Zikang Liu, Kun Zhou, Wayne Xin Zhao +3
Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success,…