8 papers
Source-Agnostic Image Translation Based on Latent Aware Adaptive Masking
Tomislav Dobrički, Byung-Woo Hong
In this work, we propose a source-agnostic framework that dynamically refines a binary mask throughout the reverse diffusion process by computing the discrepancies of a pretrained…
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
Peiwen Sun, Xudong Lu, Huadai Liu +10
While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and multi-screen collaboration, inh…
Benchmark Everything Everywhere All at Once
Shiyun Xiong, Dongming Wu, Peiwen Sun +5
Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensiv…
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Sihan Yang, Runsen Xu, Yiman Xie +10
Spatial intelligence is essential for multimodal large language models (MLLMs) operating in the complex physical world. Existing benchmarks, however, probe only single-image relati…
OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
Zhenyu Wu, Jingjing Xie, Zehao Li +8
With VLM-powered computer-using agents (CUAs) becoming increasingly capable at graphical user interface (GUI) navigation and manipulation, reliable step-level decision-making has e…
ScaleCUA: Scaling Open-Source Computer Use Agents with Cross-Platform Data
Zhaoyang Liu, Jingjing Xie, Zichen Ding +27
Vision-Language Models (VLMs) have enabled computer use agents (CUAs) that operate GUIs autonomously, showing great potential, yet progress is limited by the lack of large-scale, o…