From the 1 of 12 linked papers with an AI index.
12 papers
Evidence-Backed Video Question Answering
Shijie Wang, Honglu Zhou, Ziyang Wang +5
The paper introduces Evidence-Backed Video Question Answering (E-VQA), a task where models must provide both a textual answer and precise spatio‑temporal visual evidence (temporal…
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
Ziyang Wang, Honglu Zhou, Shijie Wang +6
Long video understanding (LVU) is challenging because answering real-world queries often depends on sparse, temporally dispersed cues buried in hours of mostly redundant and irrele…
Future Optical Flow Prediction Improves Robot Control & Video Generation
Kanchana Ranasinghe, Honglu Zhou, Yu Fang +7
Future motion representations, such as optical flow, offer immense value for control and generative tasks. However, forecasting generalizable spatially dense motion representations…
Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
Yu Fang, Kanchana Ranasinghe, Le Xue +10
Vision-Language-Action (VLA) models have achieved remarkable progress in robotic manipulation by mapping multimodal observations and instructions directly to actions. However, they…
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
Shrikant Kendre, Austin Xu, Honglu Zhou +3
Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing…
BLIP3o-NEXT: Next Frontier of Native Image Generation
Jiuhai Chen, Le Xue, Zhiyang Xu +12
We present BLIP3o-NEXT, a fully open-source foundation model in the BLIP3 series that advances the next frontier of native image generation. BLIP3o-NEXT unifies text-to-image gener…