3 papers
cs.AI2026
Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition
Jinnuo Liu, Yue Peng, Jinhan Niu +1
Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a function name: models must coordin…
cs.MA2026
ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems
Jinnuo Liu, Chuke Liu, Hua Shen
Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value alignment is typically evaluated for is…
cs.LG2026
From Navigation to Refinement: Revealing the Two-Stage Nature of Flow-based Diffusion Models through Oracle Velocity
Haoming Liu, Jinnuo Liu, Yanhao Li +5
Flow-based diffusion models have emerged as a leading paradigm for training generative models across images and videos. However, their memorization-generalization behavior remains…