3 papers
cs.CV2026
UI-Venus-1.5 Technical Report
Venus Team, Changlong Gao, Zhangxuan Gu +24
GUI agents have emerged as a powerful paradigm for automating interactions in digital environments, yet achieving both broad generality and consistently strong task performance rem…
cs.AI2025
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Zhen Zeng, Leijiang Gu, Zhangling Duan +4
Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearning methods can remove such content, they o…
cs.CV2025
SAM 3: Segment Anything with Concepts
Nicolas Carion, Laura Gustafson, Yuan-Ting Hu +35
We present Segment Anything Model (SAM) 3, a unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short…