6 papers
HACMan++: Spatially-Grounded Motion Primitives for Manipulation
Bowen Jiang, Yilin Wu, Wenxuan Zhou +2
Although end-to-end robot learning has shown some success for robot manipulation, the learned policies are often not sufficiently robust to variations in object pose or geometry. T…
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Fei Wang, Xingyu Fu, James Y. Huang +18
We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tas…
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
Sathish Reddy Indurthi, Wenxuan Zhou, Shamil Chollampatt +4
Advancements in Large Language Models (LLMs) have significantly enhanced instruction-following capabilities. However, most Instruction Fine-Tuning (IFT) datasets are predominantly…
Sim2Real Manipulation on Unknown Objects with Tactile-based Reinforcement Learning
Entong Su, Chengzhe Jia, Yuzhe Qin +4
Using tactile sensors for manipulation remains one of the most challenging problems in robotics. At the heart of these challenges is generalization: How can we train a tactile-base…
GeoLM: Empowering Language Models for Geospatially Grounded Language Understanding
Zekun Li, Wenxuan Zhou, Yao-Yi Chiang +1
Humans subconsciously engage in geospatial reasoning when reading articles. We recognize place names and their spatial relations in text and mentally associate them with their phys…
Robust Natural Language Understanding with Residual Attention Debiasing
Fei Wang, James Y. Huang, Tianyi Yan +2
Natural language understanding (NLU) models often suffer from unintended dataset biases. Among bias mitigation methods, ensemble-based debiasing methods, especially product-of-expe…