From the 1 of 34 linked papers with an AI index.
9 papers · 1 filter
A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports
Jia Yu, Yan Zhu, Yili He +12
The paper presents EndoCLIP, a vision‑language foundation model for colonoscopy that learns from lesion‑level image‑text pairs extracted from routine colonoscopy reports, achieving…
A Tale of LLMs and Induced Small Proxies: Scalable Small Language Models for Knowledge Mining
Sipeng Zhang, Longfei Yun, Shuhuai Lin +8
At the core of Deep Research is knowledge mining, the task of extracting structured information from massive unstructured text in response to user instructions. Large language mode…
MedCUA-Bench: A Screenshot-Only Benchmark for Clinical Computer-Use Agents
Jia Yu, Zilong Wang, Xinyang Jiang +2
Computer-use agents could automate repetitive screen-based clinical work, but their reliability in medical graphical user interfaces remains largely unvalidated. Existing benchmark…
Reasoning-Driven Multimodal LLM for Domain Generalization
Zhipeng Xu, Zilong Wang, Xinyang Jiang +3
This paper addresses the domain generalization (DG) problem in deep learning. While most DG methods focus on enforcing visual feature invariance, we leverage the reasoning capabili…
Mirror: A Multi-Agent System for AI-Assisted Ethics Review
Yifan Ding, Yuhui Shi, Zhiyan Li +10
Ethics review is a foundational mechanism of modern research governance, yet contemporary systems face increasing strain as ethical risks arise as structural consequences of large-…
A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
Xingjun Ma, Yixu Wang, Hengyuan Xu +18
The rapid evolution of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) has driven major gains in reasoning, perception, and generation across language and…