5 papers
Mechanistic Attention Guidance for Agent Memory Refinement
Yechao Hong, Haiquan Qiu, Yaqing Wang +1
Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely inco…
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
Haiquan Qiu, Quanming Yao
The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models. However, this progress is often hindered by notorious trai…
Scaling GraphLLM with Bilevel-Optimized Sparse Querying
Yangzhe Peng, Haiquan Qiu, Quanming Yao +1
LLMs have recently shown strong potential in enhancing node-level tasks on text-attributed graphs (TAGs) by providing explanation features. However, their practical use is severely…
Spectral Alignment as Predictor of Loss Explosion in Neural Network Training
Haiquan Qiu, You Wu, Yingjie Tan +2
Loss explosions in training deep neural networks can nullify multi-million dollar training runs. Conventional monitoring metrics like weight and gradient norms are often lagging an…
Neural Symbolic Regression of Complex Network Dynamics
Haiquan Qiu, Shuzhi Liu, Quanming Yao
Complex networks describe important structures in nature and society, composed of nodes and the edges that connect them. The evolution of these networks is typically described by d…