2 papers
cs.CV2026
What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
Kunlin Cai, Rui Song, Jinghuai Zhang +7
Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not whether a single video look…
cs.CV2025
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
Yuliang Cai, Jesse Thomason, Mohammad Rostami
Vision-language models (VLMs), such as CLIP, have demonstrated strong performance across a range of downstream tasks. However, CLIP is still limited in negation understanding: the…