collaborators

6 papers

cs.RO2026

SAIN: Structure-Aware Interactive Navigation with Active Dialogue Grounding for Mobile Robot

Yuhao Cao, Xiao Liu, Yang Xie +2

Most existing vision-language navigation tasks assume that instructions are complete and unambiguous. However, real-world robots often encounter natural human instructions that are…

cs.AI2026

Deep reflective reasoning in interdependence constrained structured data extraction from clinical notes for digital health

Jingwei Huang, Kuroush Nezafati, Zhikai Chi +9

Extracting structured information from clinical notes requires navigating a dense web of interdependent variables where the value of one attribute logically constrains others. Exis…

cs.HC2026

Software as Content: Dynamic Applications as the Human-Agent Interaction Layer

Mulong Xie, Yang Xie

Chat-based natural language interfaces have emerged as the dominant paradigm for human-agent interaction, yet they fundamentally constrain engagement with structured information an…

cs.CV2026

MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning

Meng Lu, Yuxing Lu, Yuchen Zhuang +6

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning t…

cs.AI2025

Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs

Meng Lu, Ran Xu, Yi Fang +14

While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images", i.e., to reason through multi-step visual interactions, rem…

cs.CL2025

MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science

Ran Xu, Yuchen Zhuang, Yishan Zhong +13

We introduce MedAgentGym, a scalable and interactive training environment designed to enhance coding-based biomedical reasoning capabilities in large language model (LLM) agents. M…