activity
20222026
most citedEnhancing the Protein Tertiary Structure Prediction by Multiple Sequence Alignment Generation

6 citations · 9 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

Le Zhang, Hao Chen, Vlad Roznyatovskiy +2

Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems:…

cs.CV2026

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

Qian Yang, Ankur Sikarwar, Huy Le +4

Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they reason in language and discard the fine-grained geometry the task requires. Thinking with i…

cs.CV2026

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding

Peng Zhang, Guanghao Zhang, Wanggui He +10

Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…

cs.CV2026

RiT: Vanilla Diffusion Transformers Suffice in Representation Space

Le Zhang, Ning Mang, Aishwarya Agrawal

Flow matching with -prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold structure effectively in pixel…

cs.CV20241 cited

Assessing and Learning Alignment of Unimodal Vision and Language Models

Le Zhang, Qian Yang, Aishwarya Agrawal

How well are unimodal vision and language models aligned? Although prior work have approached answering this question, their assessment methods do not directly translate to how the…

cs.CV2024

VisMin: Visual Minimal-Change Understanding

Rabiul Awal, Saba Ahmadi, Le Zhang +1

Fine-grained understanding of objects, attributes, and relationships between objects is crucial for visual-language models (VLMs). Existing benchmarks primarily focus on evaluating…