3 papers
cs.AI2026
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Ali Ansari, Haoran Sun, Andy Zeyi Liu +48
Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still strugg…
cs.LG2026
Semigroup-JEPA: Latent Dynamics Consistency for Zero-Shot Physics Generalization
Andy Zeyi Liu, Haoran Sun, Lucas Baker +2
Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning, but their capability to learn…
cs.LG2026
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
Andy Zeyi Liu, Michael Zhang, Ilana Greenberg +3
Steering large language models (LLMs) is usually done by either instruction prompting or activation steering. Prompting often gives strong control, but caches guidance tokens at ev…