3 papers
cs.CV2026
360° Image Perception with MLLMs: A Comprehensive Benchmark and a Training-Free Method
Huyen T. T. Tran, Van-Quang Nguyen, Farros Alferro +2
Multimodal Large Language Models (MLLMs) have shown impressive abilities in understanding and reasoning over conventional images. However, their perception of 360° images remains l…
cs.CV2025
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
Naoya Sogi, Takashi Shibata, Makoto Terao +2
Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasi…
cs.CV2024
Action-Agnostic Point-Level Supervision for Temporal Action Detection
Shuhei M. Yoshida, Takashi Shibata, Makoto Terao +2
We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the propo…