From the 1 of 6 linked papers with an AI index.
6 papers
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
Yijiang Li, Huiqi Zou, Bingyang Wang +1
The paper presents CEDI, a framework that evaluates vision‑language models through multi‑turn, interactive dialogues between the model, an automated examiner, and a grader, uncover…
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?
Jingheng Ye, Huiqi Zou, Simon Yu +1
AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates…
Generative Personality Simulation via Theory-Informed Structured Interview
Pengda Wang, Huiqi Zou, Han Jiang +5
Despite their potential as human proxies, LLMs often fail to generate heterogeneous data with human-like diversity, thereby diminishing their value in advancing social science rese…
Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots
Huiqi Zou, Pengda Wang, Zihan Yan +2
A chatbot's personality design is key to interaction quality. As chatbots evolved from rule-based systems to those powered by large language models (LLMs), evaluating the effective…
Practitioners' Expectations on Log Anomaly Detection
Xiaoxue Ma, Yishu Li, Jacky Keung +5
Log anomaly detection has become a common practice for software engineers to analyze software system behavior. Despite significant research efforts in log anomaly detection over th…
On the Influence of Data Resampling for Deep Learning-Based Log Anomaly Detection: Insights and Recommendations
Xiaoxue Ma, Huiqi Zou, Pinjia He +4
Numerous Deep Learning (DL)-based approaches have gained attention in software Log Anomaly Detection (LAD), yet class imbalance in training data remains a challenge, with anomalies…