6 papers
VeriBound: PAC-Bayesian Generalization Bounds for Process Reward Models Trained with Formal Verification Tools
Amirul Rahman, Mohammed Sabih Alsharari
Process Reward Models (PRMs) provide step-level verification for Large Language Model (LLM) reasoning, yet their training data acquisition remains a bottleneck: human annotation is…
UniMark: Unified Adaptive Multi-bit Watermarking for Autoregressive Image Generators
Yigit Yilmaz, Elena Petrova, Mehmet Kaya +2
Invisible watermarking for autoregressive (AR) image generation has recently gained attention as a means of protecting image ownership and tracing AI-generated content. However, ex…
Avoiding Overthinking and Underthinking: Curriculum-Aware Budget Scheduling for LLMs
Amirul Rahman, Aisha Karim, Kenji Nakamura +1
Scaling test-time compute via extended reasoning has become a key paradigm for improving the capabilities of large language models (LLMs). However, existing approaches optimize rea…
MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
Amirul Rahman, Qiang Xu, Xueying Huang
Despite significant advancements, Large Vision-Language Models (LVLMs) continue to face challenges in complex visual reasoning tasks that demand deep contextual understanding, mult…
Elevating Visual Question Answering through Implicitly Learned Reasoning Pathways in LVLMs
Liu Jing, Amirul Rahman
Large Vision-Language Models (LVLMs) have shown remarkable progress in various multimodal tasks, yet they often struggle with complex visual reasoning that requires multi-step infe…
Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction
Liu Jing, Amirul Rahman
Semantic location prediction from multimodal social media posts is a critical task with applications in personalized services and human mobility analysis. This paper introduces \te…