2 papers
cs.AI2026
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Zhengbo Zhang, Changtao Miao, Jinbo Su +10
Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex…
cs.CV2026
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
Fanwei Zeng, Changtao Miao, Jing Huang +9
The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document safety. Existing forensic methods mainly re…