13 papers
Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning
Jinhu Qi, Minda Hu, Wentao Zhang +4
Large language models (LLMs) increasingly handle in-context learning (ICL) tasks where a long, novel context defines the rules, knowledge, and output schema for a series of questio…
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and…
Geometric Collapse: When Vision Models Fail to Verify Physical Causality
Wentao Zhang, Jinhu Qi, Weiqiang Jin +3
Recent progress in large-scale self-supervised learning has improved dense geometric prediction, but it remains unclear whether such scaling yields inference-time physical plausibi…
When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation
Haowei Guo, Baolong Bi, Ruicheng Zhang +2
Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- which governs the \emph{temporal…
One-Shot Klein Cutting Planes for Lipschitz Geodesically Convex Optimization in Hyperbolic Space
Yutong Zhang, Yaoran Yang, Yifan Zhu +1
Motivated by the COLT 2023 open problem of Criscitiello, Martínez-Rubio, and Boumal on deterministic first-order methods for Lipschitz geodesically convex optimization on Hadamard…
TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning
Chusen Li, Zhou Liu, Shuigeng Zhou +1
Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain difficult to combine. Directl…