4 papers · 1 filter
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
Guosheng Zhang, Keyao Wang, Haixiao Yue +5
Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tas…
UniDet: Unified and Universal Framework for Prompt-Guided Multi-dataset 3D Detection
Yubin Wang, Zhikang Zou, Xiaoqing Ye +3
We present UniDet, a brand new framework for unified and universal multi-dataset training on 3D detection, enabling robust performance across diverse domains and generalization…
MonoFormer: One Transformer for Both Diffusion and Autoregression
Chuyang Zhao, Yuxing Song, Wenhao Wang +5
Most existing multimodality methods use separate backbones for autoregression-based discrete text generation and diffusion-based continuous visual generation, or the same backbone…
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
Jing Hao, Yuxiang Zhao, Song Chen +6
Multimodal Large Language Models (MLLMs) have shown promise in a broad range of vision-language tasks with their strong reasoning and generalization capabilities. However, they hea…