9 papers
Spiking Graph Predictive Coding for Reliable OOD Generalization
Jing Ren, Jiapeng Du, Bowen Li +6
Graphs provide a powerful basis for modeling Web-based relational data, with expressive GNNs to support the effective learning in dynamic web environments. However, real-world depl…
SOGPTSpotter: Detecting ChatGPT-Generated Answers on Stack Overflow
Suyu Ma, Chunyang Chen, Hourieh Khalajzadeh +1
Stack Overflow is a popular Q&A platform where users ask technical questions and receive answers from a community of experts. Recently, there has been a significant increase in the…
Improving Methodologies for LLM Evaluations Across Global Languages
Akriti Vij, Benjamin Chua, Darshini Ramiah +43
As frontier AI models are deployed globally, it is essential that their behaviour remains safe and reliable across diverse linguistic and cultural contexts. To examine how current…
Improving Methodologies for Agentic Evaluations Across Domains: Leakage of Sensitive Information, Fraud and Cybersecurity Threats
Ee Wei Seah, Yongsen Zheng, Naga Nikshith +67
The rapid rise of autonomous AI systems and advancements in agent capabilities are introducing new risks due to reduced oversight of real-world interactions. Yet agent testing rema…
Environment-Aware Code Generation: How far are We?
Tongtong Wu, Rongyi Chen, Wenjie Du +6
Recent progress in large language models (LLMs) has improved code generation, but most evaluations still test isolated, small-scale code (e.g., a single function) under default or…
MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems
Jieshan Chen, Suyu Ma, Qinghua Lu +2
Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on grou…