collaborators

5 papers

cs.CV2026

UMamba: A Two-level Nested U-structure Mamba for Salient Object Detection

Junhui Li, Jialu Li, Youshan Zhang

Mamba-based models have emerged as a promising alternative for salient object detection (SOD), offering significant advantages in modeling long sequences. However, existing models…

cs.CV2026

FruitEnsemble: MLLM-Guided Arbitration for Heterogeneous ensemble in Fine-Grained Fruit Recognition

Enhui Yu, Junhui Li, Ruitong Lu +2

Fine-grained fruit classification is a critical yet challenging task in agricultural computer vision, primarily hindered by a severe shortage of high-quality datasets and the high…

cs.SD2026

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

Sahil Kumar, Namrataben Patel, Honggang Wang +1

MambaVoiceCloning (MVC) asks whether the conditioning path of diffusion-based TTS can be made fully SSM-only at inference, removing all attention and explicit RNN-style recurrence…

cs.CV2025

Automatic Teaching Platform on Vision Language Retrieval Augmented Generation

Ruslan Gokhman, Jialu Li, Youshan Zhang

Automating teaching presents unique challenges, as replicating human interaction and adaptability is complex. Automated systems cannot often provide nuanced, real-time feedback tha…

cs.CV2025

SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing

Varun Biyyala, Bharat Chanderprakash Kathuria, Jialu Li +1

Video editing models have advanced significantly, but evaluating their performance remains challenging. Traditional metrics, such as CLIP text and image scores, often fall short: t…