From the 1 of 28 linked papers with an AI index.
28 papers
CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents
Weizhong Huang, Jinchao Zhang, Xiawu Zheng
Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches grow accordingly, increasing…
YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
Xu Lin, WenJie Nie, Jinlong Peng +4
Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specif…
Toward Robust In-Context Segmentation via Concept Guidance
Zhigang Chen, Xiawu Zheng, Rongrong Ji
The paper proposes Concept-Guided In-Context Segmentation (CG-ICS), which improves the robustness of few-shot image segmentation by extracting high-level semantic concepts from ref…
AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
Wanqi Yang, Yuexiao Ma, Alexander Conzelmann +4
Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in me…
HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion
Xuzhe Zheng, Yuexiao Ma, Jing Xu +3
Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing methods already construct head-specific spars…
Motion-Aware Caching for Efficient Autoregressive Video Generation
Jing Xu, Yuexiao Ma, Xuzhe Zheng +7
Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential i…