4 papers
"The Whole Is Greater Than the Sum of Its Parts": A Compatibility-Aware Multi-Teacher CoT Distillation Framework
Jin Cui, Jiaqi Guo, Ruixuan Yang +6
Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) with remarkable capabilities but typically requires prohibitive parameter scales. CoT distillation has emerge…
ECHO: Continuous Hierarchical Memory for Vision-Language-Action Models
Yanbin Hu, Jin Cui, Jiayi Lu +6
Memory capacity is a critical factor determining the performance of Vision-Language-Action (VLA) models in long-horizon manipulation tasks. Existing memory-augmented architectures…
VecAttention: Vector-wise Sparse Attention for Accelerating Long Context Inference
Anmin Liu, Ruixuan Yang, Huiqiang Jiang +5
Long-context video understanding and generation pose a significant computational challenge for Transformer-based video models due to the quadratic complexity of self-attention. Whi…
MIND: From Passive Mimicry to Active Reasoning through Capability-Aware Multi-Perspective CoT Distillation
Jin Cui, Jiaqi Guo, Jiepeng Zhou +6
While Large Language Models (LLMs) have emerged with remarkable capabilities in complex tasks through Chain-of-Thought reasoning, practical resource constraints have sparked intere…