4 papers
Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
Yongmin Kim, Shota Takashiro, Yusuke Iwasawa +2
Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference.…
Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity
Debanjan Dutta, Anish Chakrabarty, Swagatam Das
Chain of Thought (CoT) lifts the expressive ceiling of bounded-depth Transformers, with characterizations tying the number of CoT steps to circuit complexity classes. What remains…
On the Existence of Universal Simulators of Attention
Debanjan Dutta, Anish Chakrabarty, Faizanuddin Ansari +1
Previous work on the learnability of transformers \textemdash\ focused primarily on examining their ability to approximate specific algorithmic patterns through training \textemdas…
Assessing the Limits of In-Context Learning beyond Functions using Partially Ordered Relation
Debanjan Dutta, Faizanuddin Ansari, Swagatam Das
Generating rational and generally accurate responses to tasks, often accompanied by example demonstrations, highlights Large Language Model's (LLM's) remarkable In-Context Learning…