3 papers
cs.SE2026
What Does an Agentic Software Engineering Benchmark Measure? Profiling Task Demands and Agent Behaviour Beyond What Category Labels Reveal
Radin Shayanfar, Keheliya Gallaba, Ahmed E. Hassan
Agentic software engineering benchmarks are typically summarized by nominal category labels such as "bug fix" or "feature implementation," yet benchmarks carrying the same label ar…
cs.CL2025
CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment
Radin Shayanfar, Chu Fei Luo, Rohan Bhambhoria +2
Building Task-Oriented Dialogue (TOD) systems that generalize across different tasks remains a challenging problem. Data-driven approaches often struggle to transfer effectively to…
cs.CL2024
Misinformation with Legal Consequences (MisLC): A New Task Towards Harnessing Societal Harm of Misinformation
Chu Fei Luo, Radin Shayanfar, Rohan Bhambhoria +2
Misinformation, defined as false or inaccurate information, can result in significant societal harm when it is spread with malicious or even innocuous intent. The rapid online info…