15 papers
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving
Shihao Ji, HongXi Li, Zihui Song +1
Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically…
A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature--Variance Trade-offs
Shihao Ji, Mingyu Li, Zihui Song
Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order gradient-free methods. We develop a kine…
Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification
Shihao Ji, Haotao Tan, Zihui Song +1
Large Language Models (LLMs) are increasingly used with formal interactive theorem provers such as Lean 4. Scaling these systems with reinforcement learning or search methods requi…
Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context
Shihao Ji, Mingyu Li, Zihui Song
The quadratic complexity of self-attention remains a bottleneck for Large Language Models (LLMs) processing ultra-long contexts. The Naive Bayes Cognitive Engine (NBCE) parallelize…
Data Scaling as Progressive Coverage of a Predictive Contribution Spectrum
Zihui Song, Shihao Ji, Hongxi Li +2
We investigate the hypothesis that real-data scaling laws are governed by progressive coverage of a latent predictive contribution spectrum rather than by token-frequency tails alo…
L-MoE: End-to-End Training of a Lightweight Mixture of Low-Rank Adaptation Experts
Shihao Ji, Zihui Song
The Mixture of Experts (MoE) architecture enables the scaling of Large Language Models (LLMs) to trillions of parameters by activating a sparse subset of weights for each input, ma…