collaborators

5 papers

cs.CV2026

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

Junke Wang, Xiao Wang, Jiacheng Pan +16

This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework.…

cs.CV2026

Vector Map as Language: Toward Unified Remote Sensing Vector Mapping

Yinglong Yan, Yunkai Yang, Haoyi Wang +7

Remote sensing vector mapping aims to generate structured maps of geospatial entities, such as buildings, roads, and water bodies, from remote sensing imagery. In practice, vector…

cs.CV2026

OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning

Jiahao Wang, An Ping, Yanghai Wang +13

While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…

cs.CL2025

DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents

Yibin Xu, Liang Yang, Hao Chen +3

The limitation of graphical user interface (GUI) data has been a significant barrier to the development of GUI agents today, especially for the desktop / computer use scenarios. To…

cs.CL2025

SEKI: Self-Evolution and Knowledge Inspiration based Neural Architecture Search via Large Language Models

Zicheng Cai, Yaohua Tang, Yutao Lai +3

We introduce SEKI, a novel large language model (LLM)-based neural architecture search (NAS) method. Inspired by the chain-of-thought (CoT) paradigm in modern LLMs, SEKI operates i…