collaborators

7 papers

cs.AI2025

Latent Chain-of-Thought for Visual Reasoning

Guohao Sun, Hang Hua, Jian Wang +5

Chain-of-thought (CoT) reasoning is critical for improving the interpretability and reliability of Large Vision-Language Models (LVLMs). However, existing training algorithms such…

cs.CV2025

Visual Self-Refinement for Autoregressive Models

Jiamian Wang, Ziqi Zhou, Chaithanya Kumar Mummadi +5

Autoregressive models excel in sequential modeling and have proven to be effective for vision-language data. However, the spatial nature of visual signals conflicts with the sequen…

cs.CV2025

X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning

Prasanna Reddy Pulakurthi, Jiamian Wang, Majid Rabbani +3

Prevalent text-to-video retrieval systems mainly adopt embedding models for feature extraction and compute cosine similarities for ranking. However, this design presents two limita…

cs.LG2025

MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper

Runjia Zeng, Guangyan Sun, Qifan Wang +8

Considering deep neural networks as manifold mappers, the pretrain-then-fine-tune paradigm can be interpreted as a two-stage process: pretrain establishes a broad knowledge base, a…

cs.CV2025

Effective Dual-Region Augmentation for Reduced Reliance on Large Amounts of Labeled Data

Prasanna Reddy Pulakurthi, Majid Rabbani, Celso M. de Melo +2

This paper introduces a novel dual-region augmentation approach designed to reduce reliance on large-scale labeled datasets while improving model robustness and adaptability across…

cs.CV2025

Shuffle PatchMix Augmentation with Confidence-Margin Weighted Pseudo-Labels for Enhanced Source-Free Domain Adaptation

Prasanna Reddy Pulakurthi, Majid Rabbani, Jamison Heard +3

This work investigates Source-Free Domain Adaptation (SFDA), where a model adapts to a target domain without access to source data. A new augmentation technique, Shuffle PatchMix (…