6 papers
Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation Methods
Sarthak Mahapatra, Zihan Zhou, Khatoon Khedri +2
AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training on constrained hardware. We b…
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
MiniCPM Team, Wenhao An, Yingfa Chen +44
The evolution of large language models (LLMs) towards applications with ultra-long contexts faces challenges posed by the high computational and memory costs of the Transformer arc…
Order Matters in Retrosynthesis: Structure-aware Generation via Reaction-Center-Guided Discrete Flow Matching
Chenguang Wang, Zihan Zhou, Lei Bai +1
Template-free retrosynthesis methods treat the task as black-box sequence generation, limiting learning efficiency, while semi-template approaches rely on rigid reaction libraries…
Efficient Causal Structure Learning via Modular Subgraph Integration
Haixiang Sun, Pengchao Tian, Zihan Zhou +3
Learning causal structures from observational data remains a fundamental yet computationally intensive task, particularly in high-dimensional settings where existing methods face c…
Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
Jian Xiong, Jingbo Zhou, Zihan Zhou +6
Latent learning, classically theorized by Tolman, shows that biological agents (e.g., rats) can acquire internal representations of their environment without rewards, enabling rapi…
GradES: Significantly Faster Training in Transformers with Gradient-Based Early Stopping
Qifu Wen, Xi Zeng, Zihan Zhou +4
Early stopping monitors global validation loss and halts all parameter updates simultaneously, which is computationally costly for large transformers due to the extended time requi…