2 papers
cs.CL2026
TriEx: A Game-based Tri-View Framework for Explaining Internal Reasoning in Multi-Agent LLMs
Ziyi Wang, Chen Zhang, Wenjun Peng +2
Explainability for Large Language Model (LLM) agents is especially challenging in interactive, partially observable settings, where decisions depend on evolving beliefs and other a…
cs.SE2026
ProxyWar: Dynamic Assessment of LLM Code Generation in Game Arenas
Wenjun Peng, Xinyu Wang, Qi Wu
Large language models (LLMs) have revolutionized automated code generation, yet the evaluation of their real-world effectiveness remains limited by static benchmarks and simplistic…