2 citations · 2 across the 4 of their papers we have counts for
4 papers
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
Yifan Li, Yukai Gu, Yingqian Min +6
Recent breakthroughs in video generation have demonstrated an emerging capability termed Chain-of-Frames (CoF) reasoning, where models resolve complex tasks through the generation…
PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
Zikang Liu, Junyi Li, Wayne Xin Zhao +3
Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) promise human-like interaction with software applications, yet long-horizon tasks remain c…
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
Zikang Liu, Kun Zhou, Wayne Xin Zhao +3
Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success,…
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
Yifan Du, Zikang Liu, Yifan Li +7
Recently, slow-thinking reasoning systems, built upon large language models (LLMs), have garnered widespread attention by scaling the thinking time during inference. There is also…