4 papers
On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs
Rosie Zhao, Anshul Shah, Xiaoyu Zhu +5
Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-langua…
MetricGold: Leveraging Text-To-Image Latent Diffusion Models for Metric Depth Estimation
Ansh Shah, K Madhava Krishna
Recovering metric depth from a single image remains a fundamental challenge in computer vision, requiring both scene understanding and accurate scaling. While deep learning has adv…
GaitContour: Efficient Gait Recognition based on a Contour-Pose Representation
Yuxiang Guo, Anshul Shah, Jiang Liu +3
Gait recognition holds the promise to robustly identify subjects based on walking patterns instead of appearance information. In recent years, this field has been dominated by lear…
Learning to Prompt Your Domain for Vision-Language Models
Guoyizhe Wei, Feng Wang, Anshul Shah +1
Prompt learning has recently become a very efficient transfer learning paradigm for Contrastive Language Image Pretraining (CLIP) models. Compared with fine-tuning the entire encod…