14 papers
ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection
Ningning Han, Lei Fan, Jia Guo +5
The deployment of Industrial Anomaly Detection (IAD) in real-world manufacturing frequently encounters a challenging cold-start bottleneck, in which limited normal samples fail to…
CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning
He Feng, Yongjia Ma, Donglin Di +2
Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input g…
GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning
Kun Wang, Yiming Li, Mingcheng Qu +3
Implicit spatial relations and deep semantic structures encoded in object attributes are crucial for procedural planning in embodied AI systems. However, existing approaches often…
Chain of World: World Model Thinking in Latent Motion
Fuxiang Yang, Donglin Di, Lulu Tang +6
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…
FOCA: Frequency-Oriented Cross-Domain Forgery Detection, Localization and Explanation via Multi-Modal Large Language Model
Zhou Liu, Tonghua Su, Hongshi Zhang +4
Advances in image tampering techniques, particularly generative models, pose significant challenges to media verification, digital forensics, and public trust. Existing image forge…
HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models
Kun Wang, Xiao Feng, Mingcheng Qu +1
Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-…