collaborators

14 papers

cs.RO2026

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

Hongjin Ji, Guoyang Xia, Luoyang Sun +2

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in cl…

cs.CV2026

VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment

Guoyang Xia, Fengfa Li, Hongjin Ji +4

Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because…

cs.LG2026

Dense2MoE: Pushing the Pareto Frontier of On-Device LLMs via Unified Pruning and Upcycling

Fengfa Li, Hongjin Ji, Yifeng Ding +2

The Mixture of Experts MoE architecture is highly promising for resource constrained on device deployments yet training these models from scratch incurs prohibitive costs Current m…

cs.CL2026

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

Haojie Ouyang, Jianwei Lv, Lei Ren +3

Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to the self-attention's quadratic complexi…

cs.CL2026

Lngram: N-gram Conditional Memory in Latent Space

Yunao Zheng, Guoyang Xia, Xiaojie Wang +1

Sequence modeling requires both compositional reasoning and local static knowledge retrieval, yet standard Transformers handle both through dense computation. Engram partially deco…

cs.CV2026

FreeGraftor: Training-Free Cross-Image Feature Grafting for Subject-Driven Text-to-Image Generation

Zebin Yao, Lei Ren, Huixing Jiang +4

Subject-driven image generation aims to synthesize novel scenes that faithfully preserve subject identity from reference images while adhering to textual guidance. However, existin…