collaborators

5 papers

cs.DC2025

OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving

Siyu Wu, Zihan Tang, Yuting Zeng +5

Large Language Models (LLMs) are increasingly deployed in both latency-sensitive online services and cost-sensitive offline workloads. Co-locating these workloads on shared serving…

cs.CV2025

PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning

Fengyuan Sun, Hui Chen, Xinhao Xu +5

While multi-modal large language models (MLLMs) have made significant progress in recent years, the issue of hallucinations remains a major challenge. To mitigate this phenomenon,…

cs.CV2025

AdaTP: Attention-Debiased Token Pruning for Video Large Language Models

Fengyuan Sun, Leqi Shen, Hui Chen +3

Video Large Language Models (Video LLMs) have achieved remarkable results in video understanding tasks. However, they often suffer from heavy computational overhead due to the larg…

cs.CV2025

Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning

Mengyao Lyu, Yan Li, Huasong Zhong +5

The hypothesis that pretrained large language models (LLMs) necessitate only minimal supervision during the fine-tuning (SFT) stage (Zhou et al., 2024) has been substantiated by re…

cs.CV2025

LLMI3D: MLLM-based 3D Perception from a Single 2D Image

Fan Yang, Sicheng Zhao, Yanhao Zhang +4

Recent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms. However, current 3D perception methods…