5 papers
StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
Yin Wang, Haotian Hu, Jineng Han +4
Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen u…
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Jinwei Kong, Runqi Meng, Fanyi Wang +4
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited…
CLGRPO: Reasoning Ability Enhancement for Small VLMs
Fanyi Wang, Binzhi Dong, Haotian Hu +2
Small Vision Language Models (SVLMs) generally refer to models with parameter sizes less than or equal to 2B. Their low cost and power consumption characteristics confer high comme…
FastMap: Fast Queries Initialization Based Vectorized HD Map Reconstruction Framework
Haotian Hu, Jingwei Xu, Fanyi Wang +4
Reconstruction of high-definition maps is a crucial task in perceiving the autonomous driving environment, as its accuracy directly impacts the reliability of prediction and planni…
LoopAnimate: Loopable Salient Object Animation
Fanyi Wang, Peng Liu, Haotian Hu +6
Research on diffusion model-based video generation has advanced rapidly. However, limitations in object fidelity and generation length hinder its practical applications. Additional…