7 papers
Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ
Zongyun Zhang, Jiacheng Ruan, Xian Gao +5
Although multimodal large language models (MLLMs) have shown substantial potential in visual understanding and graphic code generation, editing scientific figures through code pres…
BasketHAR: A Multimodal Dataset for Human Activity Recognition and Sport Analysis in Basketball Training Scenarios
Xian Gao, Haoyue Zhang, Zongyun Zhang +3
Human Activity Recognition (HAR) involves the automatic identification of user activities and has gained significant research interest due to its broad applicability. Most HAR syst…
OnlineMate: An LLM-Based Multi-Agent Companion System for Cognitive Support in Online Learning
Xian Gao, Zongyun Zhang, Ting Liu +1
In online learning environments, students often lack personalized peer interactions, which are crucial for cognitive development and learning engagement. Although previous studies…
MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
Xian Gao, Jiacheng Ruan, Zongyun Zhang +3
With the rapid growth of academic publications, peer review has become an essential yet time-consuming responsibility within the research community. Large Language Models (LLMs) ha…
EIAD: Explainable Industrial Anomaly Detection Via Multi-Modal Large Language Models
Zongyun Zhang, Jiacheng Ruan, Xian Gao +2
Industrial Anomaly Detection (IAD) is critical to ensure product quality during manufacturing. Although existing zero-shot defect segmentation and detection methods have shown effe…
GoAI: Enhancing AI Students' Learning Paths and Idea Generation via Graph of AI Ideas
Xian Gao, Zongyun Zhang, Ting Liu +1
With the rapid advancement of artificial intelligence technology, AI students are confronted with a significant "information-to-innovation" gap: they must navigate through the rapi…