From the 1 of 4 linked papers with an AI index.
4 papers
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Lanbo Lin, Jiayao Liu, Tianyuan Yang +5
JADE is a two‑layer evaluation system that encodes expert knowledge as predefined evaluation skills and adds claim‑level, evidence‑gated assessment to more reliably evaluate AI age…
How Do We Research Human-Robot Interaction in the Age of Large Language Models? A Systematic Review
Yufeng Wang, Yuan Xu, Anastasia Nikolova +4
Advances in large language models (LLMs) are profoundly reshaping the field of human-robot interaction (HRI). While prior work has highlighted the technical potential of LLMs, few…
DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented Interfaces
Yuan Xu, Shaowen Xiang, Yizhi Song +2
Large Language Models are reshaping task automation, yet remain limited in complex, multi-step real-world tasks that require aligning with vague user intent and enabling dynamic us…
FrontierCS: Evolving Challenges for Evolving Intelligence
Qiuyang Mang, Wenhao Chai, Zhifei Li +48
We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…