4 papers
Reducing Pretraining-Generation Mismatch in Diffusion Language Models
Xiaocheng Lu, Huabin Liu, Song Guo +1
Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language mode…
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants
Junan Hu, Jian Liu, Jingxiang Lai +6
Graphical User Interface (GUI) agents have emerged as a promising paradigm for intelligent systems that perceive and interact with graphical interfaces visually. Yet supervised fin…
BoxSeg: Quality-Aware and Peer-Assisted Learning for Box-supervised Instance Segmentation
Jinxiang Lai, Wenlong Wu, Jiawei Zhan +5
Box-supervised instance segmentation methods aim to achieve instance segmentation with only box annotations. Recent methods have demonstrated the effectiveness of acquiring high-qu…
Spider: Any-to-Many Multimodal LLM
Jinxiang Lai, Jie Zhang, Jun Liu +3
Multimodal LLMs (MLLMs) have emerged as an extension of Large Language Models (LLMs), enabling the integration of various modalities. However, Any-to-Any MLLMs are limited to gener…