4 papers
MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models
Zheyuan Zhou, Liang Du, Zixun Sun +5
Despite significant progress in Visual-Language-Action (VLA), in highly complex and dynamic environments that involve real-time unpredictable interactions (such as 3D open worlds a…
FURINA: Free from Unmergeable Router via LINear Aggregation of mixed experts
Jiayi Han, Liang Du, Yinda Chen +3
The Mixture of Experts (MoE) paradigm has been successfully integrated into Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning (PEFT), delivering performance gains with…
CAD-Judge: Toward Efficient Morphological Grading and Verification for Text-to-CAD Generation
Zheyuan Zhou, Jiayi Han, Liang Du +3
Computer-Aided Design (CAD) models are widely used across industrial design, simulation, and manufacturing processes. Text-to-CAD systems aim to generate editable, general-purpose…
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
Jiayi Han, Liang Du, Yiwen Wu +3
The success of VLMs often relies on the dynamic high-resolution schema that adaptively augments the input images to multiple crops, so that the details of the images can be retaine…