works on

From the 1 of 28 linked papers with an AI index.

activity
20242026
collaborators

28 papers

cs.LG2026

CommitKV: Lifecycle-Aware KV Cache Compression via Commit Transitions for Multi-Turn Agents

Weizhong Huang, Jinchao Zhang, Xiawu Zheng

Multi-turn Reasoning-and-Acting (ReAct) agents accumulate growing trajectories of reasoning, tool calls, and observations. Their key-value (KV) caches grow accordingly, increasing…

cs.CV2026

YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family

Xu Lin, WenJie Nie, Jinlong Peng +4

Generic parameter-efficient fine-tuning (PEFT) methods transferred from language models can fail silently on real-time detectors, whose heterogeneous operators and detection-specif…

cs.CV2026

Toward Robust In-Context Segmentation via Concept Guidance

Zhigang Chen, Xiawu Zheng, Rongrong Ji

The paper proposes Concept-Guided In-Context Segmentation (CG-ICS), which improves the robustness of few-shot image segmentation by extracting high-level semantic concepts from ref…

cs.LG2026

AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization

Wanqi Yang, Yuexiao Ma, Alexander Conzelmann +4

Mixture-of-Experts (MoE) architectures scale model capacity through sparse expert activation, but their deployment remains memory-bound because all expert weights must reside in me…

cs.CV2026

HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion

Xuzhe Zheng, Yuexiao Ma, Jing Xu +3

Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing methods already construct head-specific spars…

cs.CV2026

Motion-Aware Caching for Efficient Autoregressive Video Generation

Jing Xu, Yuexiao Ma, Xuzhe Zheng +7

Autoregressive video generation paradigms offer theoretical promise for long video synthesis, yet their practical deployment is hindered by the computational burden of sequential i…