7 papers
GRACE: Grounded Reasoning via Adapter Composition and Evidence-Aware Calibration for Educational Visual Question Answering
Xinjin Li, Yudi Xia, Xi Zhao +8
Educational visual question answering, or VQA, requires models to solve curriculum-oriented multiple-choice questions using both language and visual evidence. Compared with convent…
Chain-Aware Encoding for Microservice Trace Anomaly Detection
Yiliu Xu, Ziwei Hong, Zhongheng Yang +2
Microservice traces can be structurally anomalous even when every span returns normally -- a payment flow that silently skips a risk check looks fine to any per-span monitor. Seque…
MOSAIC: Orchestrating Collaborative Knowledge Tracing with Hierarchical Semantic Alignment
Xinjin Li, Mengyue Wang, Yuzhen Lin +4
Knowledge Tracing (KT) is important for personalized education but traditionally suffers from two key limitations: a reliance on shallow ID-based representations that neglect seman…
FAST: A Synergistic Framework of Attention and State-space Models for Spatiotemporal Traffic Prediction
Xinjin Li, Jinghan Cao, Mengyue Wang +5
Traffic forecasting requires modeling complex temporal dynamics and long-range spatial dependencies over large sensor networks. Existing methods typically face a trade-off between…
Task-Specific Efficiency Analysis: When Small Language Models Outperform Large Language Models
Jinghan Cao, Yu Ma, Xinjin Li +2
Large Language Models achieve remarkable performance but incur substantial computational costs unsuitable for resource-constrained deployments. This paper presents the first compre…
CATCH: A Modular Cross-domain Adaptive Template with Hook
Xinjin Li, Yulie Lu, Jinghan Cao +3
Recent advances in Visual Question Answering (VQA) have demonstrated impressive performance in natural image domains, with models like LLaVA leveraging large language models (LLMs)…