2 papers
cs.AI2026
CurvZO: Adaptive Curvature-Guided Sparse Zeroth-Order Optimization for Efficient LLM Fine-Tuning
Shuo Wang, Ziyu Chen, Ming Tang
Fine-tuning large language models (LLMs) with backpropagation achieves high performance but incurs substantial memory overhead, limiting scalability on resource-constrained hardwar…
cs.CL2026
JEPA-Reasoner: Decoupling Latent Reasoning from Token Generation
Bingyang Kelvin Liu, Ziyu Patrick Chen, David P. Woodruff
Current autoregressive language models couple high-level reasoning and low-level token generation into a single sequential process, making the reasoning trajectory vulnerable to co…