7 papers
Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks
Lining Hu, Ting Liu, Yuzhuo Fu
Offline root-cause-analysis (RCA) benchmarks commonly rank methods by a single pooled top-1 accuracy across multiple subsystems, and engineers often read the pooled winner as a rec…
From Motion Signals to Insights: A Unified Framework for Student Behavior Analysis and Feedback in Physical Education Classes
Xian Gao, Jiacheng Ruan, Jingsheng Gao +4
Analyzing student behavior in educational scenarios is crucial for enhancing teaching quality and student engagement. Existing AI-based models often rely on classroom video footage…
GoAI: Enhancing AI Students' Learning Paths and Idea Generation via Graph of AI Ideas
Xian Gao, Zongyun Zhang, Ting Liu +1
With the rapid advancement of artificial intelligence technology, AI students are confronted with a significant "information-to-innovation" gap: they must navigate through the rapi…
ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
Xian Gao, Jiacheng Ruan, Zongyun Zhang +3
Academic paper review is a critical yet time-consuming task within the research community. With the increasing volume of academic publications, automating the review process has be…
EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language Models
Zongyun Zhang, Jiacheng Ruan, Xian Gao +2
Industrial Anomaly Detection (IAD) is critical to ensure product quality during manufacturing. Although existing zero-shot defect segmentation and detection methods have shown effe…
VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward Models
Jiacheng Ruan, Wenzhen Yuan, Xian Gao +6
Although large visual-language models (LVLMs) have demonstrated strong performance in multimodal tasks, errors may occasionally arise due to biases during the reasoning process. Re…