3 papers
cs.RO2026
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
Peize Li, Ruimeng Zhang, Ru Zhang +3
Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real…
cs.CV2026
Towards Spatial Supersensing in the Wild
Tianjun Gu, Tianyu Xin, Kuan Zhang +12
The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…
cs.CV2026
SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Zhiyuan Ma, Zhengfeng Shi, Yuning An +6
While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical rea…