Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses
Kerui Peng, Feifei Li, Xingyu Fan +1
When large language models are fine-tuned to generate persona- or tone-conditioned responses, their output diversity is severely limited--a failure we term Cross-Style Collapse. We…
cs.CL2026
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
Doyoung Kim, Zhiwei Ren, Jie Hao +11
We introduce WildAGTEval, a benchmark designed to evaluate large language model (LLM) agents' function-calling capabilities under realistic API complexity. Unlike prior work that a…