collaborators

5 papers

cs.LG2026

Emergence of Superposition: Unveiling the Training Dynamics of Chain of Continuous Thought

Hanlin Zhu, Shibo Hao, Zhiting Hu +3

Previous work shows that the chain of continuous thought (continuous CoT) improves the reasoning capability of large language models (LLMs) by enabling implicit parallel thinking,…

cs.LG2025

Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought

Hanlin Zhu, Shibo Hao, Zhiting Hu +3

Large Language Models (LLMs) have demonstrated remarkable performance in many applications, including challenging reasoning problems via chain-of-thoughts (CoTs) techniques that ge…

cs.LG2025

Spectral Journey: How Transformers Predict the Shortest Path

Andrew Cohen, Andrey Gromov, Kaiyu Yang +1

Decoder-only transformers lead to a step-change in capability of large language models. However, opinions are mixed as to whether they are really planning or reasoning. A path to m…

cs.LG2025

LLM Pretraining with Continuous Concepts

Jihoon Tack, Jack Lanchantin, Jane Yu +7

Next token prediction has been the standard training objective used in large language model pretraining. Representations are learned as a result of optimizing for token-level perpl…

cs.CL2024

Towards Full Delegation: Designing Ideal Agentic Behaviors for Travel Planning

Song Jiang, Da JU, Andrew Cohen +5

How are LLM-based agents used in the future? While many of the existing work on agents has focused on improving the performance of a specific family of objective and challenging ta…