3 citations · 6 across the 5 of their papers we have counts for
8 papers · 1 filter
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
Edward Y. Chang, Longling Geng
LLMs increasingly generate workflow actions and repairs that may be well formed yet stale, infeasible, conflicting, or destructive of their own evidence. We introduce Agentic Trans…
TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views
Edward Y. Chang
World models let agents plan against predicted physical state, but that state drifts; re-observation is costly and delayed, and repair can fail. We present TRACE-RealWorld (TRW), t…
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
Longling Geng, Andy Ouyang, Theodore Wu +10
Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: confusing association with intervention, aban…
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
Edward Y. Chang, Longling Geng
Inference-time scaling can amplify reasoning pathologies: sycophancy, rung collapse, and premature certainty. We present RAudit, a diagnostic protocol for auditing LLM reasoning wi…
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
Edward Y. Chang, Longling Geng
Large language models (LLMs) excel at rapid generation of text and multimodal content, yet they falter on transaction-style planning that demands ACID-like guarantees and real-time…
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
Edward Y. Chang, Longling Geng
This paper introduces SagaLLM, a structured multi-agent architecture designed to address four foundational limitations of current LLM-based planning systems: unreliable self-valida…