works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
most citedRepresentation-Centric Survey of Supervised Skeletal Action Recognition and the New Benchmark

1 citations · 1 across the 2 of their papers we have counts for

collaborators

7 papers

cs.CV2026

Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space

Quoc-Huy Trinh, Xi Ding, Yang Liu +7

The paper introduces SpatialMed, a benchmark and an automated pipeline that generates 3D spatial visual question‑answer pairs for medical imaging, and shows that current multimodal…

cs.CV20261 cited

Representation-Centric Survey of Supervised Skeletal Action Recognition and the New Benchmark

Yang Liu, Jiyao Yang, Madhawa Perera +8

3D skeletal action recognition has emerged as a powerful alternative to traditional RGB and depth-based approaches, offering robustness to environmental variations, computational e…

cs.CV2026

SonoSelect: Efficient Ultrasound Perception via Active Probe Exploration

Yixin Zhang, Yunzhong Hou, Longqi Li +3

Ultrasound perception typically requires multiple scan views through probe movement to reduce diagnostic ambiguity, mitigate acoustic occlusions, and improve anatomical coverage. H…

cs.CV2026

Mind the Rarities: Can Rare Skin Diseases Be Reliably Diagnosed via Diagnostic Reasoning?

Yang Liu, Jiyao Yang, Hongjin Zhao +10

Large vision-language models (LVLMs) demonstrate strong performance in dermatology; however, evaluating diagnostic reasoning for rare conditions remains largely unexplored. Existin…

cs.CV2025

GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder

Seunghyuk Cho, Zhenyue Qin, Yang Liu +3

We introduce GeoDANO, a geometric vision-language model (VLM) with a domain-agnostic vision encoder, for solving plane geometry problems. Although VLMs have been employed for solvi…

cs.CV2025

Plane Geometry Problem Solving with Multi-modal Reasoning: A Survey

Seunghyuk Cho, Zhenyue Qin, Yang Liu +3

Plane geometry problem solving (PGPS) has recently gained significant attention as a benchmark to assess the multi-modal reasoning capabilities of large vision-language models. Des…