collaborators

5 papers

cs.AI2026

TrigReason: Trigger-Based Collaboration between Small and Large Reasoning Models

Yi Zhao, Yajuan Peng, Cam-Tu Nguyen +4

Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chains of thought but suffer from high inference latency due to autoregressive reasoning.…

cs.LG2026

SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference

Yi Zhao, Yajuan Peng, Cam-Tu Nguyen +4

KV cache eviction has emerged as an effective solution to alleviate resource constraints faced by LLMs in long-context scenarios. However, existing token-level eviction methods oft…

cs.CL2025

IAM: Efficient Inference through Attention Mapping between Different-scale LLMs

Yi Zhao, Zuchao Li, Hai Zhao

LLMs encounter significant challenges in resource consumption nowadays, especially with long contexts. Despite extensive efforts dedicate to enhancing inference efficiency, these m…

cs.CL2025

DAC: A Dynamic Attention-aware Approach for Task-Agnostic Prompt Compression

Yi Zhao, Zuchao Li, Hai Zhao +2

Task-agnostic prompt compression leverages the redundancy in natural language to reduce computational overhead and enhance information density within prompts, especially in long-co…

cs.MM2025

CDIO: Cross-Domain Inference Optimization with Resource Preference Prediction for Edge-Cloud Collaboration

Zheming Yang, Wen Ji, Qi Guo +7

Currently, massive video tasks are processed by edge-cloud collaboration. However, the diversity of task requirements and the dynamics of resources pose great challenges to efficie…