3 citations · 6 across the 3 of their papers we have counts for
6 papers
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
Edward Y. Chang, Longling Geng
Inference-time scaling can amplify reasoning pathologies: sycophancy, rung collapse, and premature certainty. We present RAudit, a diagnostic protocol for auditing LLM reasoning wi…
ALAS: Transactional and Dynamic Multi-Agent LLM Planning
Longling Geng, Edward Y. Chang
Large language models enable flexible multi-agent planning but remain fragile in practice: verification is often circular, state changes are not tracked for repair, and small fault…
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
Edward Y. Chang, Longling Geng
Large language models (LLMs) excel at rapid generation of text and multimodal content, yet they falter on transaction-style planning that demands ACID-like guarantees and real-time…
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
Edward Y. Chang, Longling Geng
This paper introduces SagaLLM, a structured multi-agent architecture designed to address four foundational limitations of current LLM-based planning systems: unreliable self-valida…
REALM-Bench: A Benchmark for Evaluating Multi-Agent Systems on Real-world, Dynamic Planning and Scheduling Tasks
Longling Geng, Edward Y. Chang
This benchmark suite provides a comprehensive evaluation framework for assessing both individual LLMs and multi-agent systems in Real-world planning and scheduling scenarios. The s…
MACI: Multi-Agent Collaborative Intelligence for Adaptive Reasoning and Temporal Planning
Edward Y. Chang
Artificial intelligence requires deliberate reasoning, temporal awareness, and effective constraint management, capabilities traditional LLMs often lack due to their reliance on pa…