1 citations · 1 across the 4 of their papers we have counts for
4 papers · 1 filter
What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
Kunlin Cai, Rui Song, Jinghuai Zhang +7
Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not whether a single video look…
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
Yuliang Cai, Jesse Thomason, Mohammad Rostami
Vision-language models (VLMs), such as CLIP, have demonstrated strong performance across a range of downstream tasks. However, CLIP is still limited in negation understanding: the…
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
Yuliang Cai, Mohammad Rostami
Large vision-language models (VLMs) have shown significant performance boost in various application domains. However, adopting them to deal with several sequentially encountered ta…
Dynamic Transformer Architecture for Continual Learning of Multimodal Tasks
Yuliang Cai, Mohammad Rostami
Transformer neural networks are increasingly replacing prior architectures in a wide range of applications in different data modalities. The increasing size and computational deman…