activity
20202026
most citedAnswer-Driven Visual State Estimator for Goal-Oriented Visual Dialogue

4 citations · 8 across the 16 of their papers we have counts for

collaborators

17 papers

cs.RO2026

What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency

Luoyang Sun, Guoyang Xia, Fengfa Li +9

Vision-Language-Action (VLA) models combine a pretrained vision encoder, a language backbone, and an action head, but their relative contribution has not been established under con…

cs.RO2026

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

Hongjin Ji, Guoyang Xia, Luoyang Sun +2

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in cl…

cs.CV2026

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

Guoyang Xia, Fengfa Li, Hongjin Ji +4

Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because…

cs.CV2025

FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning

Guoyang Xia, Yifeng Ding, Fengfa Li +4

Multimodal large language models (MLLMs) have achieved impressive performance, but high-resolution visual inputs result in long sequences of visual tokens and substantial inference…

cs.CL2025

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

Haojie Ouyang, Jianwei Lv, Lei Ren +3

Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to the self-attention's quadratic complexi…

cs.CL2025

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities

Guoyang Xia, Yifeng Ding, Fengfa Li +4

Mixture of Experts (MoE) architectures have become a key approach for scaling large language models, with growing interest in extending them to multimodal tasks. Existing methods t…