most citedCausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs

3 citations · 3 across the 3 of their papers we have counts for

collaborators
Showing cs.AIShow all

8 papers · 1 filter

cs.AI2026

TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views

Edward Y. Chang

World models let agents plan against predicted physical state, but that state drifts; re-observation is costly and delayed, and repair can fail. We present TRACE-RealWorld (TRW), t…

cs.AI2026

Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows

Edward Y. Chang, Longling Geng, Emily J. Chang

LLMs increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, conflicting, or destructive of the evidenc…

cs.AI20263 cited

CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs

Longling Geng, Andy Ouyang, Theodore Wu +10

Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, aban…

cs.AI20262 cited

RAudit: A Blind Auditing Protocol for Large Language Model Reasoning

Edward Y. Chang, Longling Geng

Inference-time scaling can amplify reasoning pathologies: sycophancy, rung collapse, and premature certainty. We present RAudit, a diagnostic protocol for auditing LLM reasoning wi…

cs.AI2025

REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks

Longling Geng, Edward Y. Chang

This benchmark suite provides a comprehensive evaluation framework for assessing both individual LLMs and multi-agent systems in Real-world planning and scheduling scenarios. The s…

cs.AI2025

SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning

Edward Y. Chang, Longling Geng

This paper introduces SagaLLM, a structured multi-agent architecture designed to address four foundational limitations of current LLM-based planning systems: unreliable self-valida…