4 papers
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
Tao Yu, Haopeng Jin, Hao Wang +18
In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Ope…
CODA: Repurposing Continuous VAEs for Discrete Tokenization
Zeyu Liu, Zanlin Ni, Yeguo Hua +4
Discrete visual tokenizers transform images into a sequence of tokens, enabling token-based visual generation akin to language models. However, this process is inherently challengi…
MECD: Unlocking Multi-Event Causal Discovery in Video Reasoning
Tieyuan Chen, Huabin Liu, Tianyao He +8
Video causal reasoning aims to achieve a high-level understanding of video content from a causal perspective. However, current video reasoning tasks are limited in scope, primarily…
Dermoscopic Image Analysis for ISIC Challenge 2018
Jinyi Zou, Xiao Ma, Cheng Zhong +1
This short paper reports the algorithms we used and the evaluation performances for ISIC Challenge 2018. Our team participates in all the tasks in this challenge. In lesion segment…