most citedDIVESPOT: Depth Integrated Volume Estimation of Pile of Things Based on Point Cloud

2 citations · 2 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CV2026

CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding

Jing Jiang, Yiran Ling, Ruonan Li +2

Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods i…

cs.RO2026

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

Yiran Ling, Qing Lian, Jinghang Li +6

In this paper, we propose GTA-VLA(Guide, Think, Act), an interactive Vision-Language-Action (VLA) framework that enables spatially steerable embodied reasoning by allowing users to…

cs.RO2026

CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping

Yiran Ling, Wenxuan Li, Siying Dong +5

Robot grasping of desktop object is widely used in intelligent manufacturing, logistics, and agriculture.Although vision-language models (VLMs) show strong potential for robotic ma…

cs.CV20242 cited

DIVESPOT: Depth Integrated Volume Estimation of Pile of Things Based on Point Cloud

Yiran Ling, Rongqiang Zhao, Yixuan Shen +3

Non-contact volume estimation of pile-type objects has considerable potential in industrial scenarios, including grain, coal, mining, and stone materials. However, using existing m…

cs.CV2024

Poetry2Image: An Iterative Correction Framework for Images Generated from Chinese Classical Poetry

Jing Jiang, Yiran Ling, Binzhu Li +3

Text-to-image generation models often struggle with key element loss or semantic confusion in tasks involving Chinese classical poetry.Addressing this issue through fine-tuning mod…