5 papers
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
Kairos Team, Fei Wang, Shan You +21
We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully si…
Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +3
We present Seirênes, a self-play RL framework that transforms contextual interference from a failure mode of LLM reasoning into an internal training signal for co-evolving more re…
DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +3
The "thinking with images" paradigm represents a pivotal shift in the reasoning of Vision Language Models (VLMs), moving from text-dominant chain-of-thought to image-interactive re…
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
Zhiyuan Wu, Shuai Wang, Li Chen +5
Video diffusion models (VDMs) perform attention computation over the 3D spatio-temporal domain. Compared to large language models (LLMs) processing 1D sequences, their memory consu…
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
Chi Zhang, Haibo Qiu, Qiming Zhang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced the reasoning capabilities of Large Language Models (LLMs) and is now being applied to Vision-Langu…