collaborators

6 papers

cs.LG2026

ReCal: Reward Calibration for RL-based LLM Routing

Qihang Yu, Hanwen Tong, Zhengqi Zhang +5

Large language model (LLM) routing has emerged as an effective paradigm for leveraging the complementary strengths of multiple LLMs through dynamic model and reasoning-strategy sel…

cs.CV2026

CIAR: Interval-based Collaborative Decoding for Image Generation Acceleration

Keming Ye, Zhou Zhao, Fan Wu +1

Auto-regressive (AR) models have recently made notable progress in image generation, achieving performance comparable to diffusion-based approaches. However, their computational in…

q-bio.BM2026

ZeroFold: Protein-RNA Binding Affinity Predictions from Pre-Structural Embeddings

Josef Hanke, Sebastian Pujalte Ojeda, Shengyu Zhang +3

The accurate prediction of protein-RNA binding affinity remains an unsolved problem in structural biology, limiting opportunities in understanding gene regulation and designing RNA…

cs.IR2026

MALLOC: Benchmarking the Memory-aware Long Sequence Compression for Large Sequential Recommendation

Qihang Yu, Kairui Fu, Zhaocheng Du +10

The scaling law, which indicates that model performance improves with increasing dataset and model capacity, has fueled a growing trend in expanding recommendation models in both i…

cs.LG2025

MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices

Zhaode Wang, Jingbang Yang, Xinyu Qian +4

Large language models (LLMs) have demonstrated exceptional performance across a variety of tasks. However, their substantial scale leads to significant computational resource consu…

cs.LG2025

MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference

Kunxi Li, Zhonghua Jiang, Zhouzhou Shen +5

This paper introduces MadaKV, a modality-adaptive key-value (KV) cache eviction strategy designed to enhance the efficiency of multimodal large language models (MLLMs) in long-cont…