9 citations · 9 across the 6 of their papers we have counts for
13 papers
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
Muhammad Kamran Janjua, Hugo Silva, Di Niu +1
Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, despite access to these tool-…
Panoptic Pairwise Distortion Graph
Muhammad Kamran Janjua, Abdul Wahab, Bahador Rashidi
In this work, we introduce a new perspective on comparative image assessment by representing an image pair as a structured composition of its regions. In contrast, existing methods…
Grounding Degradations in Natural Language for All-In-One Video Restoration
Muhammad Kamran Janjua, Amirhosein Ghasemabadi, Kunlin Zhang +3
In this work, we propose an all-in-one video restoration framework that grounds degradation-aware semantic context of video frames in natural language via foundation models, offeri…
Fantastic Multi-Task Gradient Updates and How to Find Them In a Cone
Negar Hassanpour, Muhammad Kamran Janjua, Kunlin Zhang +4
Balancing competing objectives remains a fundamental challenge in multi-task learning (MTL), primarily due to conflicting gradients across individual tasks. A common solution relie…
Learning Truncated Causal History Model for Video Restoration
Amirhosein Ghasemabadi, Muhammad Kamran Janjua, Mohammad Salameh +1
One key challenge to video restoration is to model the transition dynamics of video frames governed by motion. In this work, we propose TURTLE to learn the truncated causal history…
Movement-induced Priors for Deep Stereo
Yuxin Hou, Muhammad Kamran Janjua, Juho Kannala +1
We propose a method for fusing stereo disparity estimation with movement-induced prior information. Instead of independent inference frame-by-frame, we formulate the problem as a n…