6 papers
TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
Yijin Wang, Shuyi Wang, Wenhan Zhang +1
Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation models become increasingly capable of synth…
Few-Shot Prediction for Pulsar Noise with Long Short-Term Memory Network
Qingye Tang, Dechao An, Haoran Peng +1
This work proposes a novel solution to predict pulsar timing residuals with limited data, addressing the critical challenge of data scarcity across spin-frequency subgroups of mill…
AgentMV: A State-Guided Multi-Agent Framework for Budget-Aware Music Video Generation
Huimin Wang, Leilei Ouyang, Chang Xia +3
Generating a complete music video from a song requires more than synthesizing visually plausible clips for individual lyric prompts. A practical system must maintain long-range vis…
Spiking Layer-Adaptive Magnitude-based Pruning
Junqiao Wang, Zhehang Ye, Yuqi Ouyang
Spiking Neural Networks (SNNs) provide energy-efficient computation but their deployment is constrained by dense connectivity and high spiking operation costs. Existing magnitude-b…
Adaptive Detector-Verifier Framework for Zero-Shot Polyp Detection in Open-World Settings
Shengkai Xu, Hsiang Lun Kao, Tianxiang Xu +9
Polyp detectors trained on clean datasets often underperform in real-world endoscopy, where illumination changes, motion blur, and occlusions degrade image quality. Existing approa…
Advancing Video Anomaly Detection: A Bi-Directional Hybrid Framework for Enhanced Single- and Multi-Task Approaches
Guodong Shen, Yuqi Ouyang, Junru Lu +2
Despite the prevailing transition from single-task to multi-task approaches in video anomaly detection, we observe that many adopt sub-optimal frameworks for individual proxy tasks…