5 citations · 5 across the 6 of their papers we have counts for
5 papers · 1 filter
ReGuLaR: Relation-Grounded Latent Reasoning for Large Vision-Language Models
Zihu Wang, Karthik Somayaji N. S, Peng Li
Chain-of-thought (CoT) reasoning has significantly improved the reasoning ability of large vision-language models (LVLMs) by verbalizing intermediate reasoning steps in natural lan…
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
Zihu Wang, Boxun Xu, Yuxuan Xia +1
Large vision-language models (LVLMs) exhibit impressive ability to jointly reason over visual and textual inputs. However, they often produce outputs that are linguistically fluent…
AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive Transformers
Boxun Xu, Yu Wang, Zihu Wang +1
Visual autoregressive modeling (VAR) via next-scale prediction has emerged as a scalable image generation paradigm. While Key and Value (KV) caching in large language models (LLMs)…
On Learning Discriminative Features from Synthesized Data for Self-Supervised Fine-Grained Visual Recognition
Zihu Wang, Lingqiao Liu, Scott Ricardo Figueroa Weston +2
Self-Supervised Learning (SSL) has become a prominent approach for acquiring visual representations across various tasks, yet its application in fine-grained visual recognition (FG…
Contrastive Learning with Consistent Representations
Zihu Wang, Yu Wang, Zhuotong Chen +2
Contrastive learning demonstrates great promise for representation learning. Data augmentations play a critical role in contrastive learning by providing informative views of the d…