3 papers
cs.AI2026
Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory
Aijun Yang, Qianxue Guo, Ziyi Huang +3
Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, providing complete execution trajectories to th…
cs.PF2026
WaveTune: Wave-aware Bilinear Modeling for Efficient GPU Kernel Auto-tuning
Kaixuan Zhang, Chutong Ding, Shiyou Qian +6
The rapid adoption of Large Language Models (LLMs) has made GPU inference efficiency an increasingly critical system concern. The runtime of LLM workloads is largely dominated by t…
cs.LG2025
GLOW: Graph-Language Co-Reasoning for Agentic Workflow Performance Prediction
Wei Guan, Jian Cao, Jinyu Cai +3
Agentic Workflows (AWs) have emerged as a promising paradigm for solving complex tasks. However, the scalability of automating their generation is severely constrained by the high…