collaborators

8 papers

cs.CL2026

LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning

Yi Wu, Zhimin Hu

Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, finite attention resources preven…

cs.CV2026

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

Qian Yang, Ankur Sikarwar, Huy Le +4

Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry needed for the task. Thinking w…

cs.CV2026

DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding

Peng Zhang, Guanghao Zhang, Wanggui He +10

Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…

cs.CV2026

RiT: Vanilla Diffusion Transformers Suffice in Representation Space

Le Zhang, Ning Mang, Aishwarya Agrawal

Flow matching with -prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold structure effectively in pixel…

cs.AI2026

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition

Le Zhang, Shengming Zhang, Rui Zha +3

Accurate Point of Interest (POI) attribute acquisition is essential for location-based services, yet traditional modular Interactive Voice Response (IVR) systems suffer from error…

cs.CV2025

Assessing and Learning Alignment of Unimodal Vision and Language Models

Le Zhang, Qian Yang, Aishwarya Agrawal

How well are unimodal vision and language models aligned? Although prior work have approached answering this question, their assessment methods do not directly translate to how the…