3 papers
cs.RO2024
Visuo-Tactile Zero-Shot Object Recognition with Vision-Language Model
Shiori Ueda, Atsushi Hashimoto, Masashi Hamaya +2
Tactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visu…
cs.CV2024
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
Koki Maeda, Tosho Hirasawa, Atsushi Hashimoto +4
Procedural video understanding is gaining attention in the vision and language community. Deep learning-based video analysis requires extensive data. Consequently, existing works o…
cs.CV2023
A Critical Look at the Current Usage of Foundation Model for Dense Recognition Task
Shiqi Yang, Atsushi Hashimoto, Yoshitaka Ushiku
In recent years large model trained on huge amount of cross-modality data, which is usually be termed as foundation model, achieves conspicuous accomplishment in many fields, such…