From the 1 of 4 linked papers with an AI index.
4 papers
Evidence-Backed Video Question Answering
Shijie Wang, Honglu Zhou, Ziyang Wang +5
The paper introduces Evidence-Backed Video Question Answering (E-VQA), a task where models must provide both a textual answer and precise spatio‑temporal visual evidence (temporal…
Future Optical Flow Prediction Improves Robot Control & Video Generation
Kanchana Ranasinghe, Honglu Zhou, Yu Fang +7
Future motion representations, such as optical flow, offer immense value for control and generative tasks. However, forecasting generalizable spatially dense motion representations…
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
Yu Fang, Kanchana Ranasinghe, Le Xue +10
Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they…
Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
Artemis Panagopoulou, Le Xue, Honglu Zhou +6
Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressiv…