5 papers
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
Junke Wang, Xiao Wang, Jiacheng Pan +16
This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework.…
Vector Map as Language: Toward Unified Remote Sensing Vector Mapping
Yinglong Yan, Yunkai Yang, Haoyi Wang +7
Remote sensing vector mapping aims to generate structured maps of geospatial entities, such as buildings, roads, and water bodies, from remote sensing imagery. In practice, vector…
OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
Jiahao Wang, An Ping, Yanghai Wang +13
While Omni-modal Large Language Models (OLLMs) have demonstrated impressive capabilities in jointly processing audio and visual streams, their ability to strictly adhere to complex…
DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents
Yibin Xu, Liang Yang, Hao Chen +3
The limitation of graphical user interface (GUI) data has been a significant barrier to the development of GUI agents today, especially for the desktop / computer use scenarios. To…
SEKI: Self-Evolution and Knowledge Inspiration based Neural Architecture Search via Large Language Models
Zicheng Cai, Yaohua Tang, Yutao Lai +3
We introduce SEKI, a novel large language model (LLM)-based neural architecture search (NAS) method. Inspired by the chain-of-thought (CoT) paradigm in modern LLMs, SEKI operates i…