7 papers · 1 filter
3D Geometric Tooth Alignment Planning via Deep Reinforcement Learning
Yong Li, Jianwen Lou, Jiayue Ma +3
3D geometric tooth alignment planning, which determines sequential trajectories from initial malocclusion to the final target alignment, is a cornerstone of modern digital orthodon…
Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation
Boyuan Xiao, Bohong Chen, Yumeng Li +3
In embodied vision-language decision making tasks such as robotic manipulation and navigation, Vision-Language and Vision-Language-Action Models (VLMs & VLAs) are powerful tools wi…
FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
Yaoli Liu, Yao-Xiang Ding, Kun Zhou
This paper proposes FreeFuse, a training-free framework for multi-subject text-to-image generation through automatic fusion of multiple subject LoRAs. In contrast to prior studies…
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
Bohong Chen, Yumeng Li, Youyi Zheng +2
The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categ…
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
Jingming Liu, Yumeng Li, Boyuan Xiao +5
Under pure textual modality, Large Language Models (LLMs) have demonstrated remarkable success in complex reasoning tasks by decomposing them into simpler sub-problems. However, Mu…
Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation
Bohong Chen, Yumeng Li, Yao-Xiang Ding +2
Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic fu…