works on

From the 1 of 21 linked papers with an AI index.

collaborators

21 papers

cs.RO2026

DLAM: Distributional Latent Actions with Temporal Constraints

Zuojin Tang, Feifan Luo, Haoyun Liu +10

The paper introduces DLAM, a distributional latent-action model that encodes video transitions as diagonal Gaussians with temporal constraints, improving reconstruction consistency…

cs.CV2026

HumanOmni-Speaker: Identifying Who said What and When

Detao Bai, Zhiheng Ma, Xihan Wei

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, mult…

cs.CV2026

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Ronghan Chen, Yandan Yang, Zuojin Tang +18

Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack expl…

cs.LG2026

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

Cong Wan, Zeyu Guo, Zijian Cai +6

Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics…

cs.CV2026

ReMoT: Reinforcement Learning with Motion Contrast Triplets

Cong Wan, Zeyu Guo, Jiangyang Li +5

We present ReMoT, a unified training paradigm to systematically address the fundamental shortcomings of VLMs in spatio-temporal consistency -- a critical failure point in navigatio…

cs.RO2026

Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models

Haoyun Liu, Jianzhuang Zhao, Xinyuan Chang +11

Despite the rapid progress of vision-language-action (VLA) models, the prevailing practice of predicting action chunks as discrete waypoints remains structurally misaligned with th…