1 citations · 3 across the 6 of their papers we have counts for
8 papers
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara +1
Multimodal Large Language Models (MLLMs) have experienced rapid progress in visual recognition tasks in recent years. Given their potential integration into many critical applicati…
A Critical Review of Predominant Bias in Neural Networks
Jiazhi Li, Mahyar Khayatkhoei, Jiageng Zhu +3
Bias issues of neural networks garner significant attention along with its promising advancement. Among various bias issues, mitigating two predominant biases is crucial in advanci…
Look, Learn and Leverage (L): Mitigating Visual-Domain Shift and Discovering Intrinsic Relations via Symbolic Alignment
Hanchen Xie, Jiageng Zhu, Mahyar Khayatkhoei +2
Modern deep learning models have demonstrated outstanding performance on discovering the underlying mechanisms when both visual appearance and intrinsic relations (e.g., causal str…
An Investigation on The Position Encoding in Vision-Based Dynamics Prediction
Jiageng Zhu, Hanchen Xie, Jiazhi Li +2
Despite the success of vision-based dynamics prediction models, which predict object states by utilizing RGB images and simple object descriptions, they were challenged by environm…
A Critical View of Vision-Based Long-Term Dynamics Prediction Under Environment Misalignment
Hanchen Xie, Jiageng Zhu, Mahyar Khayatkhoei +3
Dynamics prediction, which is the problem of predicting future states of scene objects based on current and prior states, is drawing increasing attention as an instance of learning…
Spatial Frequency Bias in Convolutional Generative Adversarial Networks
Mahyar Khayatkhoei, Ahmed Elgammal
As the success of Generative Adversarial Networks (GANs) on natural images quickly propels them into various real-life applications across different domains, it becomes more and mo…