From the 1 of 11 linked papers with an AI index.
11 papers
3D Geometric Tooth Alignment Planning via Deep Reinforcement Learning
Yong Li, Jianwen Lou, Jiayue Ma +3
The paper introduces a deep reinforcement learning system that automatically generates 3D tooth alignment trajectories for orthodontic treatment, using a transformer-based agent an…
Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation
Boyuan Xiao, Bohong Chen, Yumeng Li +3
In embodied vision-language decision making tasks such as robotic manipulation and navigation, Vision-Language and Vision-Language-Action Models (VLMs & VLAs) are powerful tools wi…
OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis
Naaisha Agarwal, Yihan Wu, Yichang Jian +7
Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must understand the causal mechanisms and constraints…
FreeFuse: Multi-Subject LoRA Fusion via Adaptive Token-Level Routing at Test Time
Yaoli Liu, Yao-Xiang Ding, Kun Zhou
This paper proposes FreeFuse, a training-free framework for multi-subject text-to-image generation through automatic fusion of multiple subject LoRAs. In contrast to prior studies…
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
Jingming Liu, Yumeng Li, Boyuan Xiao +5
Under pure textual modality, Large Language Models (LLMs) have demonstrated remarkable success in complex reasoning tasks by decomposing them into simpler sub-problems. However, Mu…
Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
Yifei Peng, Yaoli Liu, Enbo Xia +5
We propose ILP-CoT, a method that bridges Inductive Logic Programming (ILP) and Multimodal Large Language Models (MLLMs) for abductive logical rule induction. The task involves bot…