works on

From the 1 of 8 linked papers with an AI index.

activity
20242026
collaborators

8 papers

cs.CV2026

Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering

Haotian Xia, Zilin Xiao, Junbo Zou +2

Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence across time. Existing methods typically rely o…

cs.AI2026

SportD: How do VLMs physically strategize?

Jasin Cekinmez, Addison J. Wu, Haotian Xia +11

The paper introduces SportD, a benchmark that tests whether vision‑language models can choose optimal shoot or pass actions in soccer situations, comparing model choices to a value…

cs.CL2026

StoryAlign: Evaluating and Training Reward Models for Story Generation

Haotian Xia, Hao Peng, Yunjia Qi +4

Story generation aims to automatically produce coherent, structured, and engaging narratives. Although large language models (LLMs) have significantly advanced text generation, sto…

cs.CV2026

DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning

Junbo Zou, Haotian Xia, Zhen Ye +5

Sports video understanding requires perceiving high-speed dynamics, complex rules, and long temporal contexts. Yet, current Multimodal Large Language Models (MLLMs) remain narrowly…

cs.CV2026

SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports

Haotian Xia, Haonan Ge, Junbo Zou +16

Deeply understanding sports requires an intricate blend of fine-grained visual perception and rule-based reasoning - a challenge that pushes the limits of current multimodal models…

cs.CE2025

VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding

Zhaowei Liu, Xin Guo, Haotian Xia +11

Multimodal large language models (MLLMs) hold great promise for automating complex financial analysis. To comprehensively evaluate their capabilities, we introduce VisFinEval, the…