34 papers
BasketEvent: Understanding Who Did What and When in Basketball Videos
Yu Zhang, Jiayuan Rao, Haoning Wu +1
Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears. However, exist- ing metho…
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
Yeyao Ma, Chen Li, Xiaosong Zhang +2
Post-training of flow matching models-aligning the output distribution with a high-quality target-is mathematically equivalent to imitation learning. While Supervised Fine-Tuning m…
A Vision-language Framework for Comparative Reasoning in Radiology
Tengfei Zhang, Ziheng Zhao, Xiaoman Zhang +5
Medical imaging artificial intelligence has achieved strong performance in isolated image interpretation, but remains poorly aligned with radiological practice, where diagnosis and…
Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
Cheng Liang, Pengcheng Qiu, Ya Zhang +3
Large language models (LLMs) are increasingly proposed as clinical agents, yet static, single-turn benchmarks cannot capture how a model dynamically delivers care across an encount…
SoccerMaster: A Vision Foundation Model for Soccer Understanding
Haolin Yang, Jiayuan Rao, Haoning Wu +1
Soccer understanding has recently garnered growing research interest due to its domain-specific complexity and unique challenges. Unlike prior works that typically rely on isolated…
POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
Yikun Liu, Yuan Liu, Le Tian +6
Large Multimodal Models (LMMs) excel at visual perception but struggle with real-time, knowledge-intensive queries due to their reliance on static parametric knowledge. While multi…