4 papers
Cross-Domain Transfer and Few-Shot Learning for Personal Identifiable Information Recognition
Junhong Ye, Xu Yuan, Xinying Qiu
Accurate recognition of personally identifiable information (PII) is central to automated text anonymization. This paper investigates the effectiveness of cross-domain model transf…
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
Yiran Meng, Junhong Ye, Wei Zhou +4
Cross-video question answering presents significant challenges beyond traditional single-video understanding, particularly in establishing meaningful connections across video strea…
WebSynthesis: World-Model-Guided MCTS for Efficient WebUI-Trajectory Synthesis
Yifei Gao, Junhong Ye, Jiaqi Wang +1
Recent advancements in large language models (LLMs) have significantly improved the capabilities of web agents. However, effectively navigating complex and dynamic web environments…
AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models
Jiaming Zhang, Junhong Ye, Xingjun Ma +5
Due to their multimodal capabilities, Vision-Language Models (VLMs) have found numerous impactful applications in real-world scenarios. However, recent studies have revealed that V…