2 citations · 4 across the 13 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems
Hanglong Lv, Dawei Zhu, Lei Li +11
Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. O…
cs.CL2024
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
Botian Jiang, Lei Li, Xiaonan Li +5
The rapid advancement of Multimodal Large Language Models (MLLMs) has been accompanied by the development of various benchmarks to evaluate their capabilities. However, the true na…
cs.CL2024
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
Jiahui Gao, Renjie Pi, Tianyang Han +5
The deployment of multimodal large language models (MLLMs) has demonstrated remarkable success in engaging in conversations involving visual inputs, thanks to the superior power of…