2 citations · 2 across the 5 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
CoFiE: Coarse-to-Fine Evidence Selection for Efficient Streaming Video Understanding
Jing Jiang, Yiran Ling, Ruonan Li +2
Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods i…
cs.CV2024★ 2 cited
DIVESPOT: Depth Integrated Volume Estimation of Pile of Things Based on Point Cloud
Yiran Ling, Rongqiang Zhao, Yixuan Shen +3
Non-contact volume estimation of pile-type objects has considerable potential in industrial scenarios, including grain, coal, mining, and stone materials. However, using existing m…
cs.CV2024
Poetry2Image: An Iterative Correction Framework for Images Generated from Chinese Classical Poetry
Jing Jiang, Yiran Ling, Binzhu Li +3
Text-to-image generation models often struggle with key element loss or semantic confusion in tasks involving Chinese classical poetry.Addressing this issue through fine-tuning mod…