activity
20242026
collaborators

5 papers

cs.CV2026

StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design

Yin Wang, Haotian Hu, Jineng Han +4

Deploying a vision-language model with full UI understanding on end devices has long been trapped between accuracy and efficiency: on one side is the accuracy bar for OCR, screen u…

cs.AI2026

SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models

Jinwei Kong, Runqi Meng, Fanyi Wang +4

Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools remain difficult to deploy under limited…

cs.CV2025

CLGRPO: Reasoning Ability Enhancement for Small VLMs

Fanyi Wang, Binzhi Dong, Haotian Hu +2

Small Vision Language Models (SVLMs) generally refer to models with parameter sizes less than or equal to 2B. Their low cost and power consumption characteristics confer high comme…

cs.CV2025

FastMap: Fast Queries Initialization Based Vectorized HD Map Reconstruction Framework

Haotian Hu, Jingwei Xu, Fanyi Wang +4

Reconstruction of high-definition maps is a crucial task in perceiving the autonomous driving environment, as its accuracy directly impacts the reliability of prediction and planni…

cs.CV2024

LoopAnimate: Loopable Salient Object Animation

Fanyi Wang, Peng Liu, Haotian Hu +6

Research on diffusion model-based video generation has advanced rapidly. However, limitations in object fidelity and generation length hinder its practical applications. Additional…