collaborators

7 papers

cs.CV2026

Benchmarking Vision-Language Models for Microscopic Plant Image Understanding

Tianqi Wei, Xin Yu, Zhi Chen +2

Microscopic imaging provides essential visual evidence for studying plant biology and pathology at the cellular and subcellular levels. However, existing benchmarks on vision-langu…

cs.CV2026

StomataSeg: Semi-Supervised Instance Segmentation for Sorghum Stomatal Components

Zhongtian Huang, Zhi Chen, Zi Huang +8

Sorghum is a globally important cereal grown widely in water-limited and stress-prone regions. Its strong drought tolerance makes it a priority crop for climate-resilient agricultu…

cs.CV2026

Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval

Zecheng Zhao, Zhi Chen, Zi Huang +2

Text-to-Video Retrieval (TVR) is essential in video platforms. Dense retrieval with dual-modality encoders leads in accuracy, but its computation and storage scale poorly with corp…

cs.CV2025

Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation

Zhi Chen, Xin Yu, Xiaohui Tao +2

Vision-language models (VLMs) such as CLIP achieve zero-shot transfer across various tasks by pre-training on numerous image-text pairs. These models often benefit from using an en…

cs.CV2025

Augment to Segment: Tackling Pixel-Level Imbalance in Wheat Disease and Pest Segmentation

Tianqi Wei, Xin Yu, Zhi Chen +2

Accurate segmentation of foliar diseases and insect damage in wheat is crucial for effective crop management and disease control. However, the insect damage typically occupies only…

cs.CV2025

Are Synthetic Videos Useful? A Benchmark for Retrieval-Centric Evaluation of Synthetic Videos

Zecheng Zhao, Selena Song, Tong Chen +3

Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synt…