4 papers
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou +6
AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimoda…
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
Yapeng Mi, Yanpeng Zhao, Hengli Li +6
Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image gen…
NEP: Autoregressive Image Editing via Next Editing Token Prediction
Huimin Wu, Xiaojian Ma, Haozhe Zhao +2
Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approach…
AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents
Haitao Hu, Peng Chen, Yanpeng Zhao +1
Large Language Models (LLMs) have been increasingly integrated into computer-use agents, which can autonomously operate tools on a user's computer to accomplish complex tasks. Howe…