6 papers
Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda +2
Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts b…
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
Kuniaki Saito, Risa Shinoda, Shohei Tanaka +3
Hallucination detection in captions (HalDec) assesses a vision-language model's ability to correctly align image content with text by identifying errors in captions that misreprese…
HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
Kuniaki Saito, Risa Shinoda, Shohei Tanaka +3
Hallucination detection in captions (HalDec) assesses a vision-language model's ability to correctly align image content with text by identifying errors in captions that misreprese…
Zero-shot Hierarchical Plant Segmentation via Foundation Segmentation Models and Text-to-image Attention
Junhao Xing, Ryohei Miyakawa, Yang Yang +5
Foundation segmentation models achieve reasonable leaf instance extraction from top-view crop images without training (i.e., zero-shot). However, segmenting entire plant individual…
Interaction-via-Actions: Cattle Interaction Detection with Joint Learning of Action-Interaction Latent Space
Ren Nakagawa, Yang Yang, Risa Shinoda +4
This paper introduces a method and application for automatically detecting behavioral interactions between grazing cattle from a single image, which is essential for smart livestoc…
GaussianPlant: Structure-aligned Gaussian Splatting for 3D Reconstruction of Plants
Yang Yang, Risa Shinoda, Hiroaki Santo +1
We present a method for jointly recovering the appearance and internal structure of botanical plants from multi-view images based on 3D Gaussian Splatting (3DGS). While 3DGS exhibi…