activity
20242026
collaborators
Showing cs.CLShow all

16 papers · 1 filter

cs.CL2026

Structured Thoughts For Improved Reasoning And Context Pruning

Zain Sarwar, Supriyo Chakraborty, Berkcan Kapusuzoglu +5

Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient. In this work, we introduce Structured T…

cs.CL2026

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

Chia-Hsuan Lee, Sihui Dai, Mingyang Zhou +5

Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction that consume tokens without im…

cs.CL2026

SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision

Chia-Hsuan Lee, Zelei Cheng, Yu Wang +4

On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent rollouts yield noisy gradie…

cs.CL2026

Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications

Paresh Dashore, Shreyas Kulkarni, Uttam Gurram +5

Large language model (LLM)-based multi-agent systems demonstrate strong performance on complex reasoning and task execution, enabling broad enterprise applications. However, produc…

cs.CL2026

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

Genta Indra Winata, Amartya Chakraborty, Yuzhen Lin +12

Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain li…

cs.CL2026

MemGym: a Long-Horizon Memory Environment for LLM Agents

Wujiang Xu, Yu Wang, Kai Mei +8

Memory is a central capability for LLM agents operating across long-horizon tasks. Existing memory benchmarks predominantly evaluate retention of personalized information in multi-…