activity
20242026
collaborators

5 papers

cs.CV2026

v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound

Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou +6

AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimoda…

cs.CV2026

MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning

Yapeng Mi, Yanpeng Zhao, Hengli Li +6

Reasoning-augmented machine learning systems have shown improved performance in various domains, including image generation. However, existing reasoning-based methods for image gen…

cs.CV2025

NEP: Autoregressive Image Editing via Next Editing Token Prediction

Huimin Wu, Xiaojian Ma, Haozhe Zhao +2

Text-guided image editing involves modifying a source image based on a language instruction and, typically, requires changes to only small local regions. However, existing approach…

cs.CR2025

AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents

Haitao Hu, Peng Chen, Yanpeng Zhao +1

Large Language Models (LLMs) have been increasingly integrated into computer-use agents, which can autonomously operate tools on a user's computer to accomplish complex tasks. Howe…

cs.CL2024

An Efficient Recipe for Long Context Extension via Middle-Focused Positional Encoding

Tong Wu, Yanpeng Zhao, Zilong Zheng

Recently, many methods have been developed to extend the context length of pre-trained large language models (LLMs), but they often require fine-tuning at the target length ($\gg4K…