2 papers
cs.RO2026
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
Jiahang Tu, Fengyu Yang, Chenyang Ma +8
Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, existing research rarely extends this para…
cs.CV2025
Iris: Integrating Language into Diffusion-based Monocular Depth Estimation
Ziyao Zeng, Jingcheng Ni, Daniel Wang +5
Traditional monocular depth estimation suffers from inherent ambiguity and visual nuisances. We demonstrate that language can enhance monocular depth estimation by providing an add…