9 papers
Decoupling the Declarative from the Procedural in Vision-Language-Action Models
Nikolaos Tsagkas, Andreas Sochopoulos, Chris Xiaoxuan Lu +2
Deploying generalist robotic agents in the real world requires transferable skills. Specifically, a policy trained to clone a behavior from object-specific demonstrations must gene…
M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
Junqiao Fan, Yunjiao Zhou, Yizhuo Yang +6
Human mesh reconstruction (HMR) provides direct insights into body-environment interaction, which enables various immersive applications. While existing large-scale HMR datasets re…
Physically Accurate Differentiable Inverse Rendering for Radio Frequency Digital Twin
Xingyu Chen, Xinyu Zhang, Kai Zheng +4
Digital twins, virtual simulated replicas of physical scenes, are transforming system design across industries. However, their potential in radio frequency (RF) systems has been li…
Attentive Feature Aggregation or: How Policies Learn to Stop Worrying about Robustness and Attend to Task-Relevant Visual Cues
Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier +4
The adoption of pre-trained visual representations (PVRs), leveraging features from large-scale vision models, has become a popular paradigm for training visuomotor policies. Howev…
The Temporal Trap: Entanglement in Pre-Trained Visual Representations for Visuomotor Policy Learning
Nikolaos Tsagkas, Andreas Sochopoulos, Duolikun Danier +2
The integration of pre-trained visual representations (PVRs) has significantly advanced visuomotor policy learning. However, effectively leveraging these models remains a challenge…
Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse
Zhekai Duan, Yuan Zhang, Shikai Geng +3
Embodied Chain-of-Thought (ECoT) reasoning enhances vision-language-action (VLA) models by improving performance and interpretability through intermediate reasoning steps. However,…