activity
20242026
collaborators

16 papers

cs.CV2026

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

Ben Wan, Yan Feng, Zihan Tang +4

DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural inf…

cs.AI2026

IFDNS: An Iterative Feedback-Driven Neuro-Symbolic Method for Faithful Logical Reasoning

Xiaoheng Wang, Tongxuan Liu, Zi Gong +5

Large language models (LLMs) have demonstrated impressive capabilities across a wide range of reasoning tasks, including logical and mathematical problem-solving. While prompt-base…

cs.AI2025

AgentBay: A Hybrid Interaction Sandbox for Seamless Human-AI Intervention in Agentic Systems

Yun Piao, Hongbo Min, Hang Su +28

The rapid advancement of Large Language Models (LLMs) is catalyzing a shift towards autonomous AI Agents capable of executing complex, multi-step tasks. However, these agents remai…

cs.CL2025

From Hypothesis to Premises: LLM-based Backward Logical Reasoning with Selective Symbolic Translation

Qingchuan Li, Mingyue Cheng, Zirui Liu +3

Logical reasoning is a core challenge in natural language understanding and a fundamental capability of artificial intelligence, underpinning scientific discovery, mathematical the…

cs.DC2025

ProServe: Unified Multi-Priority Request Scheduling for LLM Serving

Weizhe Huang, Tao Peng, Tongxuan Liu +4

The widespread deployment of large language models (LLMs) for interactive applications necessitates serving systems that can handle thousands of concurrent requests with diverse Se…

cs.DC2025

OOCO: Latency-disaggregated Architecture for Online-Offline Co-locate LLM Serving

Siyu Wu, Zihan Tang, Yuting Zeng +5

Large Language Models (LLMs) are increasingly deployed in both latency-sensitive online services and cost-sensitive offline workloads. Co-locating these workloads on shared serving…