26 papers
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
Jiarui Yang, Jiale Zhange, Jiawei Li +5
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by…
Beyond Illumination: A Conditional Mutual Information-Guided Network for Low-Light Image Enhancement
Ya-nan Guan, Shaonan Zhang, Tao Dai +5
Low-light image enhancement (LLIE) seeks to restore structural fidelity, natural color rendition, and proper exposure from images captured under inadequate lighting conditions. Rec…
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
Niu Lian, Tongbo Chen, Alan Chen +10
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction. However, unified mul…
FinMamba: Market-Aware Graph Enhanced Multi-Level Mamba for Stock Movement Prediction
Yifan Hu, Peiyuan Liu, Yuante Li +5
Recently, combining stock features with inter-stock correlations has become a common and effective approach for stock movement prediction. However, financial data presents signific…
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
Wen Huang, Jiarui Yang, Tao Dai +4
Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing too…
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…