From the 1 of 7 linked papers with an AI index.
7 papers
Tactile Modality Fusion for Vision-Language-Action Models
Charlotte Morissette, Amin Abyaneh, Wei-Di Chang +5
The paper introduces TacFiLM, a lightweight method that fuses tactile data with visual features in vision‑language‑action models to improve robot manipulation tasks that involve co…
EMoE: Training-Free Expert Disagreement for Uncertainty-Aware Text-to-Image Diffusion
Lucas Berry, Axel Brando, Wei-Di Chang +2
Large text-to-image diffusion models rarely expose reliable signals of when a prompt is likely to produce a poorly aligned generation, especially when training data is undisclosed.…
The Surprising Difficulty of Search in Model-Based Reinforcement Learning
Wei-Di Chang, Mikael Henaff, Brandon Amos +2
This paper investigates search in model-based reinforcement learning (RL). Conventional wisdom holds that long-term predictions and compounding errors are the primary obstacles for…
Large Pre-Trained Models for Bimanual Manipulation in 3D
Hanna Yurchyk, Wei-Di Chang, Gregory Dudek +1
We investigate the integration of attention maps from a pre-trained Vision Transformer into voxel representations to enhance bimanual robotic manipulation. Specifically, we extract…
Generalizable Imitation Learning Through Pre-Trained Representations
Wei-Di Chang, Francois Hogan, Scott Fujimoto +2
In this paper, we leverage self-supervised vision transformer models and their emergent semantic abilities to improve the generalization abilities of imitation learning policies. W…
Learning Capacity: A Measure of the Effective Dimensionality of a Model
Daiwei Chen, Wei-Kai Chang, Pratik Chaudhari
We use a formal correspondence between thermodynamics and inference, where the number of samples can be thought of as the inverse temperature, to study a quantity called ``learning…