activity
20232026
most citedBDC-Adapter: Brownian Distance Covariance for Better Vision-Language Reasoning

2 citations · 5 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

11 papers · 1 filter

cs.CV2026

StreamScout: Learning When to Look Deeper for Streaming Video Understanding

Ce Zhang, Jing Bi, Jinxi He +9

Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a…

cs.CV2026

LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos

Ce Zhang, Jinxi He, Katia Sycara +1

Despite rapid progress in Multi-modal Large Language Models (MLLMs), understanding long-form videos is still bottlenecked by limited context windows. While recent keyframe sampling…

cs.CV2025

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Ce Zhang, Zifu Wan, Zhehan Kan +7

While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text responses that do not alig…

cs.CV2024

Dual Prototype Evolving for Test-Time Generalization of Vision-Language Models

Ce Zhang, Simon Stepputtis, Katia Sycara +1

Test-time adaptation, which enables models to generalize to diverse data with unlabeled test samples, holds significant value in real-world scenarios. Recently, researchers have ap…

cs.CV20241 cited

HiKER-SGG: Hierarchical Knowledge Enhanced Robust Scene Graph Generation

Ce Zhang, Simon Stepputtis, Joseph Campbell +2

Being able to understand visual scenes is a precursor for many downstream tasks, including autonomous driving, robotics, and other vision-based approaches. A common approach enabli…

cs.CV2024

Test-time Distribution Learning Adapter for Cross-modal Visual Reasoning

Yi Zhang, Ce Zhang

Vision-Language Pre-Trained (VLP) models, such as CLIP, have demonstrated remarkable effectiveness in learning generic visual representations. Several approaches aim to efficiently…