4 papers
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
Yu Ran, Wentao Zhao, Xin Zhang +1
Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executable clicks. The security impl…
Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality
Junliang Li, Yucheng Wang, Yan Chen +5
Hallucination in large language models (LLMs) during long-form generation remains difficult to address under existing reinforcement learning from human feedback (RLHF) frameworks,…
ERNIE 5.0 Technical Report
Haifeng Wang, Hua Wu, Tian Wu +432
In this report, we introduce ERNIE 5.0, a natively autoregressive foundation model desinged for unified multimodal understanding and generation across text, image, video, and audio…
Secure Video Quality Assessment Resisting Adversarial Attacks
Ao-Xiang Zhang, Yuan-Gen Wang, Yu Ran +3
The exponential surge in video traffic has intensified the imperative for Video Quality Assessment (VQA). Leveraging cutting-edge architectures, current VQA models have achieved hu…