2 papers
cs.HC2026
CentaurTA Studio: A Self-Improving Human-Agent Collaboration System for Thematic Analysis
Lei Wang, Min Huang, Eduard Dragut
Thematic analysis is difficult to scale: manual workflows are labor-intensive, while fully automated pipelines often lack controllability and transparent evaluation. We present \te…
cs.DC2026
ACE-Bench: A Lightweight Benchmark for Evaluating Azure SDK Usage Correctness
Wenxing Zhu, Simeng Qi, Junkui Chen +7
We present ACE-Bench (Azure SDK Coding Evaluation Benchmark), an execution-free benchmark that provides fast, reproducible pass or fail signals for whether large language model (LL…