works on

From the 1 of 10 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Shufan Li, Konstantinos Kallidromitis, Hritik Bansal +7

LaViDa introduces a diffusion-based vision-language model that combines a vision encoder with discrete diffusion to enable fast parallel decoding and controllable multimodal genera…

cs.CV2026

HoneyBee: Data Recipes for Vision-Language Reasoners

Hritik Bansal, Devendra Singh Sachan, Kai-Wei Chang +4

Recent advances in vision-language models (VLMs) have made them highly effective at reasoning tasks. However, the principles underlying the construction of performant VL reasoning…

cs.CV2025

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Hritik Bansal, Clark Peng, Yonatan Bitton +3

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However,…

cs.CV2024

TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation

Hritik Bansal, Yonatan Bitton, Michal Yarom +3

Most of these text-to-video (T2V) generative models often produce single-scene video clips that depict an entity performing a particular action (e.g., 'a red panda climbing a tree'…

cs.CV2024

VideoPhy: Evaluating Physical Commonsense for Video Generation

Hritik Bansal, Zongyu Lin, Tianyi Xie +7

Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of…