1 citations · 1 across the 11 of their papers we have counts for
8 papers · 1 filter
DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
Cong Wan, Zeyu Guo, Zijian Cai +6
Raw multimodal streams are abundant but noisy, redundant, and unaligned with any particular training objective. Turning them into supervision today means either brittle heuristics…
CoRe: A Comprehensive Framework for Cross-Image Comparative Reasoning in Vision-Language Models
Lin Peng, Cong Wan, Zeyu Guo +2
Cross-image comparative reasoning remains challenging for vision-language models (VLMs), especially when correct prediction requires fine-grained attribute grounding and globally c…
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
Kairos Team, Fei Wang, Shan You +21
We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physical world model should not aim to fully si…
Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning
Cong Wan, Ying He, Zhongzhan Huang +1
Test-time Scaling (TTS) has emerged as a pivotal research direction for enhancing model performance by dynamically allocating computational resources during inference. Recent advan…
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
Jiangyang Li, Cong Wan, Changjie Wu +8
Reliable spatial reasoning remains a core bottleneck for vision-language models (VLMs). Existing mainstream training paradigms for spatial reasoning largely rely on outcome alignme…
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs
Jianchao Zhao, Huoren Yang, Yusong Hu +6
Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades under local deployment conditio…