3 papers
cs.CV2026
Reference-Free Image Quality Assessment for Virtual Try-On via Human Feedback
Yuki Hirakawa, Takashi Wada, Ryotaro Shimizu +6
As virtual try-on (VTON) systems become increasingly important in fashion e-commerce, there is a growing need for reliable reference-free evaluation methods, since ground-truth ima…
cs.RO2026
ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
Zeyuan He, Bowen Yang, Zhirui Fang +9
Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observati…
cs.CV2026
MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models
Tianwei Chen, Takuya Furusawa, Yuki Hirakawa +3
This paper introduces a multi-label visual emotion analysis benchmark dataset for comprehensively evaluating the ability of multimodal large language models (MLLMs) to predict the…