activity
20242026
collaborators

5 papers

cs.CV2026

CodePercept: Code-Grounded Visual STEM Perception for MLLMs

Tongkun Guan, Zhibo Yang, Jianqiang Wan +10

When MLLMs fail at Science, Technology, Engineering, and Mathematics (STEM) visual reasoning, a fundamental question arises: is it due to perceptual deficiencies or reasoning limit…

cs.CV2025

Video-CoM: Interactive Video Reasoning via Chain of Manipulations

Hanoona Rasheed, Mohammed Zumri, Muhammad Maaz +3

Recent multimodal large language models (MLLMs) have advanced video understanding, yet most still "think about videos" ie once a video is encoded, reasoning unfolds entirely in tex…

cs.CV2025

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos

Hanoona Rasheed, Abdelrahman Shaker, Anqi Tang +4

Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual informa…

cs.CL2025

LLM Post-Training: A Deep Dive into Reasoning Large Language Models

Komal Kumar, Tajamul Ashraf, Omkar Thawakar +7

Large Language Models (LLMs) have transformed the natural language processing landscape and brought to life diverse applications. Pretraining on vast web-scale data has laid the fo…

cs.CV2024

Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration

Akshay Dudhane, Omkar Thawakar, Syed Waqas Zamir +3

All-in-one image restoration tackles different types of degradations with a unified model instead of having task-specific, non-generic models for each degradation. The requirement…