2 citations · 2 across the 3 of their papers we have counts for
5 papers · 1 filter
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang +23
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used throughout commercial AI products, but are r…
Interactive Task Alignment as a POMDP
Andy Dai, Zexue He, Zhenyu Zhang +2
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous. Users arrive with incomplete, explorator…
Interactive Evaluation Requires a Design Science
Keyang Xuan, Peiyang Song, Pan Lu +10
AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other…
Bridging Dual Knowledge Graphs for Multi-Hop Question Answering in Construction Safety
Yuxin Zhang, Xi Wang, Mo Hu +1
Information retrieval and question answering from safety regulations are essential for automated construction compliance checking but are hindered by the linguistic and structural…
Responsible AI in Construction Safety: Systematic Evaluation of Large Language Models and Prompt Engineering
Farouq Sammour, Jia Xu, Xi Wang +2
Construction remains one of the most hazardous sectors. Recent advancements in AI, particularly Large Language Models (LLMs), offer promising opportunities for enhancing workplace…