collaborators

7 papers

cs.CV2026

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

Yin Wang, Haotian Hu, Jineng Han +4

Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen u…

cs.AI2026

SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models

Jinwei Kong, Runqi Meng, Fanyi Wang +4

Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited…

cs.CL2026

ANCHOR: Abductive Network Construction with Hierarchical Orchestration for Reliable Probability Inference in Large Language Models

Wentao Qiu, Guanran Luo, Zhongquan Jian +3

A central challenge in large-scale decision-making under incomplete information is estimating reliable probabilities. Recent approaches use Large Language Models (LLMs) to generate…

cs.CL2026

DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory

Wentao Qiu, Haotian Hu, Fanyi Wang +2

Large language model (LLM) agents require long-term memory to leverage information from past interactions. However, existing memory systems often face a fidelity--efficiency trade-…

cs.CL2026

AGSC: Adaptive Granularity and Semantic Clustering for Uncertainty Quantification in Long-text Generation

Guanran Luo, Wentao Qiu, Wanru Zhao +4

Large Language Models (LLMs) have demonstrated impressive capabilities in long-form generation, yet their application is hindered by the hallucination problem. While Uncertainty Qu…

cs.CL2026

DTCRS: Dynamic Tree Construction for Recursive Summarization

Guanran Luo, Zhongquan Jian, Wentao Qiu +2

Retrieval-Augmented Generation (RAG) mitigates the hallucination problem of Large Language Models (LLMs) by incorporating external knowledge. Recursive summarization constructs a h…