11 citations · 11 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Vision as Unified Multimodal Generation
Xiaoyang Han, Jianhua Li, Kewang Deng +14
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal…
cs.CV2025
ConsistCompose: Unified Multimodal Layout Control for Image Composition
Xuanke Shi, Boxuan Li, Xiaoyang Han +4
Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with imag…
cs.CV2024
GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices
Quanfeng Lu, Wenqi Shao, Zitao Liu +7
Autonomous Graphical User Interface (GUI) navigation agents can enhance user experience in communication, entertainment, and productivity by streamlining workflows and reducing man…