4 citations · 5 across the 3 of their papers we have counts for
6 papers · 1 filter
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
Chenrui Fan, Yijun Liang, Shweta Bhardwaj +3
While many vision-language models (VLMs) are developed to answer well-defined, straightforward questions with highly specified targets, as in most benchmarks, they often struggle i…
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
Yijun Liang, Ming Li, Chenrui Fan +7
Color plays an important role in human perception and usually provides critical clues in visual reasoning. However, it is unclear whether and how vision-language models (VLMs) can…
Diffusion Curriculum: Synthetic-to-Real Data Curriculum via Image-Guided Diffusion
Yijun Liang, Shweta Bhardwaj, Tianyi Zhou
Low-quality or scarce data has posed significant challenges for training deep neural networks in practice. While classical data augmentation cannot contribute very different new da…
Efficient Video Classification Using Fewer Frames
Shweta Bhardwaj, Mukundhan Srinivasan, Mitesh M. Khapra
Recently,there has been a lot of interest in building compact models for video classification which have a small memory footprint (<1 GB). While these models are compact, they typi…
I Have Seen Enough: A Teacher Student Network for Video Classification Using Fewer Frames
Shweta Bhardwaj, Mitesh M. Khapra
Over the past few years, various tasks involving videos such as classification, description, summarization and question answering have received a lot of attention. Current models f…
Recovering from Random Pruning: On the Plasticity of Deep Convolutional Neural Networks
Deepak Mittal, Shweta Bhardwaj, Mitesh M. Khapra +1
Recently there has been a lot of work on pruning filters from deep convolutional neural networks (CNNs) with the intention of reducing computations. The key idea is to rank the fil…