11 papers
MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
Junhao Ruan, Abudukeyumu Abudula, Bei Li +8
Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational retrieval benchmarks suffer…
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Yuqi Tang, Yang Shi, Zhuoran Zhang +21
Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural dist…
Controlling Decision Drift in Multimodal Sentiment Analysis with Missing Modalities
Chenglizhao Chen, Yuchen Cao, Xinyu Liu +3
Multimodal sentiment analysis relies on textual, acoustic, and visual signals, yet real-world data often suffer from modality missing and quality imbalance. Existing methods genera…
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
V Team, Wenyi Hong, Xiaotao Gu +94
We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…
Understanding Real-World Traffic Safety through RoadSafe365 Benchmark
Xinyu Liu, Darryl C. Jacob, Yuxin Liu +4
Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we in…
"Do I Trust the AI?" Towards Trustworthy AI-Assisted Diagnosis: Understanding User Perception in LLM-Supported Reasoning
Yuansong Xu, Yichao Zhu, Haokai Wang +7
Large language models (LLMs) have shown considerable potential in supporting medical diagnosis. However, their effective integration into clinical workflows is hindered by physicia…