4 papers · 1 filter
Can Editing LLMs Inject Harm?
Canyu Chen, Baixiang Huang, Zekun Li +12
Large Language Models (LLMs) have emerged as a new information channel. Meanwhile, one critical but under-explored question is: Is it possible to bypass the safety alignment and in…
SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
Zekun Li, Shinda Huang, Jiangtian Wang +8
As language agents increasingly automate critical tasks, their ability to follow domain-specific standard operating procedures (SOPs), policies, and constraints when taking actions…
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
Zekun Li, Xianjun Yang, Kyuri Choi +11
Scientific figure interpretation is a crucial capability for AI-driven scientific assistants built on advanced Large Vision Language Models. However, current datasets and benchmark…
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
Zekun Li, Zhiyu Zoey Chen, Mike Ross +7
Large language models (LLMs) are increasingly prevalent in conversational systems due to their advanced understanding and generative capabilities in general contexts. However, thei…