3 citations · 6 across the 3 of their papers we have counts for
8 papers
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
Edward Y. Chang, Longling Geng, Emily J. Chang
LLMs increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidenc…
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
Longling Geng, Andy Ouyang, Theodore Wu +10
Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, aban…
Epistemic Regret Minimization: Label-Free Causal Critique Beyond Outcome Reward
Edward Y. Chang, Longling Geng
Large language models can answer causal questions correctly for the wrong reasons. Current RL methods reward \emph{what} a model concludes but ignore \emph{why}, reinforcing correl…
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
Edward Y. Chang, Longling Geng
Inference-time scaling can amplify reasoning pathologies: sycophancy, rung collapse, and premature certainty. We present RAudit, a diagnostic protocol for auditing LLM reasoning wi…
ALAS: Transactional and Dynamic Multi-Agent LLM Planning
Longling Geng, Edward Y. Chang
Large language models enable flexible multi-agent planning but remain fragile in practice: verification is often circular, state changes are not tracked for repair, and small fault…
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
Longling Geng, Edward Y. Chang
This benchmark suite provides a comprehensive evaluation framework for assessing both individual LLMs and multi-agent systems in Real-world planning and scheduling scenarios. The s…