10 papers · 1 filter
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang +7
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either s…
Adversarial Attacks for Good: A Survey of Proactive Protection across the Visual Content Lifecycle
Jiaming Zhang, Boyang Chen, Zherui Li +14
Once visual content enters an AI pipeline, its owner often retains little technical control over how it is used. Legal and regulatory remedies can address misuse, but many technica…
ShutterMuse: Capture-Time Photography Guidance with MLLMs
Jiayu Li, Yixiao Fang, Tianyu Hu +5
Real-world photography requires capture-time guidance for both camera framing and subject pose. Yet existing aesthetic cropping benchmarks mainly evaluate post-hoc crop prediction…
AudioMosaic: Contrastive Masked Audio Representation Learning
Hanxun Huang, Qizhou Wang, Xingjun Ma +3
Audio self-supervised learning (SSL) aims to learn general-purpose representations from large-scale unlabeled audio data. While recent advances have been driven mainly by generativ…
ImageAttributionBench: How Far Are We from Generalizable Attribution?
Tingshu Mou, Zhipeng Wei, Chao Gong +2
The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image provenance and misinformation…
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
Tingshu Mou, Jiabo He, Renying Wang +5
Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training on curated benchmarks, leavin…