2 papers
cs.CV2025
HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
Junwen Chen, Peilin Xiong, Keiji Yanai
Recent human-object interaction detection (HOID) methods highly require prior knowledge from vision-language models (VLMs) to enhance the interaction recognition capabilities. The…
cs.CV2025
PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models
Junwen Chen, Heyang Jiang, Yanbin Wang +6
Generating high-quality, multi-layer transparent images from text prompts can unlock a new level of creative control, allowing users to edit each layer as effortlessly as editing t…