most citedSA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks

14 citations · 14 across the 1 of their papers we have counts for

collaborators

6 papers

cs.RO2026

LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction

Jin Lou, Zhiyuan Jing, Xupeng Wang +21

Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: actions are exposed, but their explanato…

cs.AI2026

DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents

Yu Liu, Zhilin Liu, Zhiwei Yang +7

As large language models evolve from question-answering systems into general-purpose agents, evaluation must move beyond static answer correctness to assess multimodal perception,…

cs.AI2026

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

Xingming Long, Yu Liu, Zhiwei Yang +7

Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or extern…

cs.RO2026

ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation

Zeyuan He, Bowen Yang, Zhirui Fang +9

Vision-Language-Action (VLA) models have shown promise for robotic manipulation, yet most existing policies operate reactively by directly regressing actions from current observati…

cs.RO2026

InternVLA-A1: Unifying Understanding, Generation and Action for Robotic Manipulation

Junhao Cai, Zetao Cai, Jiafei Cao +39

Prevalent Vision-Language-Action (VLA) models are typically built upon Multimodal Large Language Models (MLLMs) and demonstrate exceptional proficiency in semantic understanding, b…

eess.IV202314 cited

SA-Med2D-20M Dataset: Segment Anything in 2D Medical Imaging with 20 Million masks

Jin Ye, Junlong Cheng, Jianpin Chen +12

Segment Anything Model (SAM) has achieved impressive results for natural image segmentation with input prompts such as points and bounding boxes. Its success largely owes to massiv…