103 citations · 120 across the 63 of their papers we have counts for
6 papers · 1 filter
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
Yue Xun, Junyu Liu, Qian Niu +10
We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF materials released by the Japanese M…
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
Koya Sakamoto, Taiki Miyanishi, Daichi Azuma +6
Visual search in 3D environments requires embodied agents to actively explore their surroundings and acquire task-relevant evidence. However, existing visual search and embodied AI…
Towards High-resolution and Disentangled Reference-based Sketch Colorization
Dingkun Yan, Xinrui Wang, Ru Wang +5
Sketch colorization is a critical task for automating and assisting in the creation of animations and digital illustrations. Previous research identified the primary difficulty as…
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
Bum Jun Kim, Makoto Kawano, Yusuke Iwasawa +1
While the robustness of vision models is often measured, their dependence on specific architectural design choices is rarely dissected. We investigate why certain vision architectu…
ColorizeDiffusion v2: Enhancing Reference-based Sketch Colorization Through Separating Utilities
Dingkun Yan, Xinrui Wang, Yusuke Iwasawa +3
Reference-based sketch colorization methods have garnered significant attention due to their potential applications in the animation production industry. However, most existing met…
Image Referenced Sketch Colorization Based on Animation Creation Workflow
Dingkun Yan, Xinrui Wang, Zhuoru Li +4
Sketch colorization plays an important role in animation and digital illustration production tasks. However, existing methods still meet problems in that text-guided methods fail t…