12 papers
MiniWorld: Democratizing the Training of Video World Models from Scratch
Yian Zhao, Ruochong Zheng, Hongcan Guo +3
Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through autoregressive state transitions…
DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration
Wenkai Wang, Tao Xiong, Jingchen Ni +6
Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents…
SimSD: Simple Speculative Decoding in Diffusion Language Models
Junxia Cui, Haotian Ye, Runchu Tian +9
Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decodi…
Continuous Latent Diffusion Language Model
Hongcan Guo, Qinyu Zhao, Yian Zhao +8
Large language models have achieved remarkable success under the autoregressive paradigm, yet high-quality text generation need not be tied to a fixed left-to-right order. Existing…
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
Wenkai Wang, Xiyun Li, Hongcan Guo +5
Graphical User Interface (GUI) grounding requires mapping natural language instructions to precise pixel coordinates. However, due to visually homogeneous elements and dense layout…
Disentangling Deception and Hallucination Failures in LLMs
Haolang Lu, Hongrui Peng, WeiYe Fu +5
Failures in large language models (LLMs) are often analyzed from a behavioral perspective, where incorrect outputs in factual question answering are commonly associated with missin…