From the 1 of 12 linked papers with an AI index.
12 papers
TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning
Jinhu Qi, Wentao Zhang, Siu Man Ng +4
The paper introduces TREK, a benchmark and deterministic evaluation kit for testing large language model agents on complex travel itinerary planning, requiring joint satisfaction o…
Geometric Collapse: When Vision Models Fail to Verify Physical Causality
Wentao Zhang, Jinhu Qi, Weiqiang Jin +3
Recent progress in large-scale self-supervised learning has improved dense geometric prediction, but it remains unclear whether such scaling yields inference-time physical plausibi…
When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation
Haowei Guo, Baolong Bi, Ruicheng Zhang +2
Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- which governs the \emph{temporal…
One-Shot Klein Cutting Planes for Lipschitz Geodesically Convex Optimization in Hyperbolic Space
Yutong Zhang, Yaoran Yang, Yifan Zhu +1
Motivated by the COLT 2023 open problem of Criscitiello, MartÃnez-Rubio, and Boumal on deterministic first-order methods for Lipschitz geodesically convex optimization on Hadamard…
TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning
Chusen Li, Zhou Liu, Shuigeng Zhou +1
Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain difficult to combine. Directl…
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
Yifan Dai, Zhenhua Wu, Bohan Zeng +18
Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evid…