From the 1 of 4 linked papers with an AI index.
4 papers
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang +23
The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…
Interactive Task Alignment as a POMDP
Andy Dai, Zexue He, Zhenyu Zhang +2
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, explorator…
Interactive Evaluation Requires a Design Science
Keyang Xuan, Peiyang Song, Pan Lu +10
AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other…
Bridging Dual Knowledge Graphs for Multi-Hop Question Answering in Construction Safety
Yuxin Zhang, Xi Wang, Mo Hu +1
Information retrieval and question answering from safety regulations are essential for automated construction compliance checking but are hindered by the linguistic and structural…