10 papers
Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
Jingbo Wen, Liang He, Mingyu Cao +4
Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty to maximize average accuracy under a fixed…
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
Liang He, Jingbo Wen, Hongyu Gu +5
Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has substantially improved skill retrieval du…
CRIP: Channel Level Representation Injection for Personalized One-Shot Federated Learning
Zijian Jiang, Chaoli Sun, Handing Wang +1
One-shot federated learning (OSFL) has emerged as a promising collaborative model learning framework with only a single round of communication, offering significant advantages in c…
TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents
Shaojie Zhuang, Lu Yin, Guangshun Wei +3
Automatic tooth segmentation and identification from intra-oral scanned 3D models are fundamental problems in digital dentistry, yet most existing approaches rely on task-specific…
BudgetDraft: Acceptance-Aware Multi-View Training for Sparse-KV Speculative Decoding
Liang He, Jingbo Wen, Qishi Zhan +4
Speculative decoding speeds up autoregressive decoding by using a drafter to propose multiple tokens that a verifier validates in parallel. In resource-constrained deployments, the…
ELAS: Efficient Pre-Training of Low-Rank Large Language Models via 2:4 Activation Sparsity
Jiaxi Li, Lu Yin, Li Shen +5
Large Language Models (LLMs) have achieved remarkable capabilities, but their immense computational demands during training remain a critical bottleneck for widespread adoption. Lo…