1 citations · 1 across the 9 of their papers we have counts for
9 papers
OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling
Yikun Miao, Fangqi Zhu, Quanxin Shou +6
Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts levera…
RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy
Zhengyang Yan, Junhao Li, Fangqi Zhu +6
Flow-matching Vision-Language-Action (VLA) policies have shown strong potential for robotic manipulation but often suffer from compounding errors caused by distribution shifts duri…
AdaTok: Self-Budgeting Image Tokenization with Quality-Preserving Dynamic Tokens
Xiaocheng Lu, Yuxi Chen, Jie Zhang +5
Image tokenizers, from 2D grids to recent 1D sequences, typically encode every image with the same fixed number of tokens. Yet visual complexity is highly heterogeneous, so a unifo…
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
Dazhao Du, Jian Liu, Jialong Qin +7
Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues and language priors rather…
3D Generation for Embodied AI and Robotic Simulation: A Survey
Tianwei Ye, Yifan Mao, Minwen Liao +6
Embodied AI and robotic systems increasingly depend on scalable, diverse, and physically grounded 3D content for simulation-based training and real-world deployment. While 3D gener…
HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
Quanxin Shou, Fangqi Zhu, Shawn Chen +9
Vision-Language-Action (VLA) models have shown strong performance in robotic manipulation, but often struggle in long-horizon or out-of-distribution scenarios due to the lack of ex…