20 citations · 44 across the 37 of their papers we have counts for
8 papers · 1 filter
GDLAM: Group-Disentangled Latent Action Model for Highly Disentangled Embodied Pretraining
Jiarui Yang, Jiawei Li, Jiale Zhang +5
Latent action models (LAMs) learn action-related representations from action-free videos via self-supervised future prediction, offering a scalable paradigm for embodied intelligen…
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
Niu Lian, Tongbo Chen, Zhehao Yu +8
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction. However, unified mul…
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
Jiarui Yang, Jiale Zhange, Jiawei Li +5
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by…
Beyond Illumination: A Conditional Mutual Information-Guided Network for Low-Light Image Enhancement
Ya-nan Guan, Shaonan Zhang, Tao Dai +5
Low-light image enhancement (LLIE) seeks to restore structural fidelity, natural color rendition, and proper exposure from images captured under inadequate lighting conditions. Rec…
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…
Forecasting as Rendering: A 2D Gaussian Splatting Framework for Time Series Forecasting
Yixin Wang, Yifan Hu, Peiyuan Liu +3
Time series forecasting remains a challenging problem due to the intricate entanglement of intra-period fluctuations and inter-period trends. While recent advances have attempted t…