8 papers
Layer barriers for colour-biased tight Hamilton cycles
Zijian Deng, Qinfei Tang, Caihong Yang
We construct a family of layer barriers for colour-biased tight Hamilton cycles in uniform hypergraphs. For every and every , we give a red--blue col…
Backpropagating Through Simulation: Analytic Policy Gradients for Sample and Learning Efficient Differentiable Continuous Control
Yueci Deng
Model-free reinforcement learning algorithms such as Proximal Policy Optimization (PPO) treat the environment as a black box, estimating policy gradients from sampled rewards; this…
From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation
Sheng Xu, Ruixing Jin, Huayi Zhou +6
Although robotic manipulation has made significant progress, reliable execution remains challenging because task failures are inevitable in dynamic and unstructured environments. T…
Adaptive TD-Lambda for Cooperative Multi-agent Reinforcement Learning
Yue Deng, Zirui Wang, Yin Zhang
TD() in value-based MARL algorithms or the Temporal Difference critic learning in Actor-Critic-based (AC-based) algorithms synergistically integrate elements from Monte-Carlo s…
DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks
Yueci Deng, Guiliang Liu, Kui Jia
Deploying generative World-Action Models for manipulation is severely bottlenecked by redundant pixel-level reconstruction, memory scaling, and sequential inferenc…
EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards
Ruixiang Wang, Qingming Liu, Yueci Deng +3
Video generative models are increasingly used as world models for robotics, where a model generates a future visual rollout conditioned on the current observation and task instruct…