4 papers
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
Yin Wang, Haotian Hu, Jineng Han +4
Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen u…
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Jinwei Kong, Runqi Meng, Fanyi Wang +4
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited…
DimMem: Dimensional Structuring for Efficient Long-Term Agent Memory
Wentao Qiu, Haotian Hu, Fanyi Wang +2
Large language model (LLM) agents require long-term memory to leverage information from past interactions. However, existing memory systems often face a fidelity--efficiency trade-…
CLGRPO: Reasoning Ability Enhancement for Small VLMs
Fanyi Wang, Binzhi Dong, Haotian Hu +2
Small Vision Language Models (SVLMs) generally refer to models with parameter sizes less than or equal to 2B. Their low cost and power consumption characteristics confer high comme…