3 papers
cs.CV2025
MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
Naoya Sogi, Takashi Shibata, Makoto Terao +2
Result diversification (RD) is a crucial technique in Text-to-Image Retrieval for enhancing the efficiency of a practical application. Conventional methods focus solely on increasi…
cs.CV2024
Action-Agnostic Point-Level Supervision for Temporal Action Detection
Shuhei M. Yoshida, Takashi Shibata, Makoto Terao +2
We propose action-agnostic point-level (AAPL) supervision for temporal action detection to achieve accurate action instance detection with a lightly annotated dataset. In the propo…
cs.CV2024
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
Naoya Sogi, Takashi Shibata, Makoto Terao
The pre-trained vision and language (V\&L) models have substantially improved the performance of cross-modal image-text retrieval. In general, however, V\&L models have limited ret…