activity
20242026
collaborators

8 papers

cs.CL2026

Chronos: Learning Temporal Dynamics of Reasoning Chains for Test-Time Scaling

Kai Zhang, Jiayi Liao, Chengpeng Li +3

Test-Time Scaling (TTS) has emerged as an effective paradigm for improving the reasoning performance of large language models (LLMs). However, existing methods -- most notably majo…

cs.DC2026

MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era

Lei Zhang, Mouxiang Chen, Ruisheng Cao +16

The rapid development of interactive and autonomous AI systems signals our entry into the agentic era. Training and evaluating agents on complex agentic tasks such as software engi…

cs.CL2025

AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval

Kai Zhang, Xinyuan Zhang, Ejaz Ahmed +11

Accurate recall from large scale memories remains a core challenge for memory augmented AI assistants performing question answering (QA), especially in similarity dense scenarios w…

cs.HC2025

Practicing a Second Language Without Fear: Mixed Reality Agents for Interactive Group Conversation

Mariana Fernandez-Espinosa, Kai Zhang, Jad Bendarkawi +8

Developing speaking proficiency in a second language can be cognitively demanding and emotionally taxing, often triggering fear of making mistakes or being excluded from larger gro…

cs.CL2025

R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search

Yibo Wang, Haotian Luo, Huanjin Yao +8

Chain-of-Thought (CoT) reasoning enhances large language models (LLMs) by enabling step-by-step problem-solving, yet its extension to Long-CoT introduces substantial computational…

cs.CL2025

Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought

Tencent Hunyuan Team, Ao Liu, Botong Zhou +248

As Large Language Models (LLMs) rapidly advance, we introduce Hunyuan-TurboS, a novel large hybrid Transformer-Mamba Mixture of Experts (MoE) model. It synergistically combines Mam…