5 papers
From Detection to Understanding: TAR and TAR-Bench for Multi-Task Traffic Anomaly Reasoning
Han Zhang, Yilin Zhao, Zaid Pervaiz Bhat +5
We present TAR (Traffic Anomaly Reasoning) and TAR-Bench datasets, resources for training and evaluating video-language models beyond anomaly detection. TAR contains 44,040 chain-o…
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
Han Zhang, Wanting Jiang, Tomasz Kornuta +2
Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what…
Mind-of-Director: Multi-modal Agent-Driven Film Previsualization via Collaborative Decision-Making
Shufeng Nan, Mengtian Li, Sixiao Zheng +3
We present Mind-of-Director, a multi-modal agent-driven framework for film previz that models the collaborative decision-making process of a film production team. Given a creative…
Decision-Level Ordinal Modeling for Multimodal Essay Scoring with Large Language Models
Han Zhang, Jiamin Su, Li liu
Automated essay scoring (AES) predicts multiple rubric-defined trait scores for each essay, where each trait follows an ordered discrete rating scale. Most LLM-based AES methods ca…
HALO: Human-Aligned End-to-end Image Retargeting with Layered Transformations
Yiran Xu, Siqi Xie, Zhuofang Li +13
Image retargeting aims to change the aspect-ratio of an image while maintaining its content and structure with less visual artifacts. Existing methods still generate many artifacts…