3 papers
cs.SE2026
HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection
Luke Patterson, Li Wang, Adam Faulkner
Thanks to the rapid adoption of AI code assistants powered by large language models (LLMs), industry codebases are, increasingly, a hybrid of AI- and human-authored code. For risk…
cs.AI2026
Deconstructing Instruction-Following: A New Benchmark for Granular Evaluation of Large Language Model Instruction Compliance Abilities
Alberto Purpura, Li Wang, Sahil Badyal +2
Reliably ensuring Large Language Models (LLMs) follow complex instructions is a critical challenge, as existing benchmarks often fail to reflect real-world use or isolate complianc…
cs.AI2026
Enhancing LLM Instruction Following: An Evaluation-Driven Multi-Agentic Workflow for Prompt Instructions Optimization
Alberto Purpura, Li Wang, Sahil Badyal +2
Large Language Models (LLMs) often generate substantively relevant content but fail to adhere to formal constraints, leading to outputs that are conceptually correct but procedural…