activity
20242026
most citedFault Diagnosis in Power Grids with Large Language Model

3 citations · 3 across the 6 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Capability-Routed Visual Retrieval and Evidence Threading for Long-Context Document Question Answering

Amirul Rahman, Aisha Karim, Kenji Nakamura +1

Annual reports, diligence packs, and infographic dashboards bury numbers in page images: axes, cell grids, and footnotes that OCR pipelines flatten and that page-level visual retri…

cs.CV2026

UniMark: Unified Adaptive Multi-bit Watermarking for Autoregressive Image Generators

Yigit Yilmaz, Elena Petrova, Mehmet Kaya +2

Invisible watermarking for autoregressive (AR) image generation has recently gained attention as a means of protecting image ownership and tracing AI-generated content. However, ex…

cs.CV2025

MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models

Amirul Rahman, Qiang Xu, Xueying Huang

Despite significant advancements, Large Vision-Language Models (LVLMs) continue to face challenges in complex visual reasoning tasks that demand deep contextual understanding, mult…

cs.CV2025

Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs

Liu Jing, Amirul Rahman

Large Vision-Language Models (LVLMs) have shown remarkable progress in various multimodal tasks, yet they often struggle with complex visual reasoning that requires multi-step infe…

cs.CV2024

Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction

Liu Jing, Amirul Rahman

Semantic location prediction from multimodal social media posts is a critical task with applications in personalized services and human mobility analysis. This paper introduces \te…