2 papers
cs.AR2026
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures
Di Mu, Tengyuan Jin, Zhenkun Wang +6
Modern deep learning workloads increasingly comprise heterogeneous computation graphs that combine compute-intensive operators with memory-intensive subgraphs. Existing deep learni…
cs.AI2025
Demystifying and Enhancing the Efficiency of Large Language Model Based Search Agents
Tiannuo Yang, Zebin Yao, Bowen Jin +4
Large Language Model (LLM)-based search agents have shown remarkable capabilities in solving complex tasks by dynamically decomposing problems and addressing them through interleav…