18 papers
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
Dongxin Guo, Jikun Wu, Siu Ming Yiu
AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and in…
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces. We present QD-LLM, a framework for parameter-efficient neuroe…
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Gradient-based preference optimization methods for large language model (LLM) alignment suffer from preference collapse, converging to narrow behavioral modes while neglecting pref…
Capacity-Controlled Global Attention for Graph Transformers
Yang Liu, Dongxin Guo, Tom Zheng +3
Global self-attention drives modern graph transformers, yet the softmax at its core imposes a structural constraint rarely examined directly: every attention row is non-negative an…
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases but, on the evidence we present, because…
Model Collapse as Cultural Evolution
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Model collapse, the progressive degradation of LLMs trained on their own outputs, has been characterized statistically but lacks a linguistic explanation for which structures degra…