3 papers
cs.RO2026
Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation
Chuyao Fu, Xiaowei Chi, Yuhan Rui +14
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing a…
cs.CV2026
From Spatial to Spectral: An Efficient, Frequency-Guided Feature Representation Learner for Small Object Detection
Yuhan Rui, Shihan Qiao, Yibin Lou +7
Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operations of spatial-domain detectors that indisc…
cs.AI2026
ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents
Yincheng Zhou, Athena Zhuoming Zhong, Shijie Zhang +3
As multimodal agents move from interface understanding to real software control, successful trajectory discovery in live desktop environments becomes a key challenge. GUI tasks req…