3 papers
cs.AI2026
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
Shouju Wang, Haopeng Zhang
As language-model agents evolve from passive chatbots into proactive assistants that handle personal data, evaluating their adherence to social norms becomes increasingly critical,…
cs.CL2025
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models
Kaiying Kevin Lin, Hsiyu Chen, Haopeng Zhang
While large language models (LLMs) have demonstrated impressive performance across a wide range of natural language processing (NLP) tasks in high-resource languages, their capabil…
cs.LG2025
MermaidFlow: Redefining Agentic Workflow Generation via Safety-Constrained Evolutionary Programming
Chengqi Zheng, Jianda Chen, Yueming Lyu +5
Despite the promise of autonomous agentic reasoning, existing workflow generation methods frequently produce fragile, unexecutable plans due to unconstrained LLM-driven constructio…