collaborators

5 papers

cs.CV2026

Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning

Qinwu Xu

Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive representation learning. Reconstruction-base…

cs.LG2026

Miller-Index-Based Latent Crystallographic Fracture Plane Reasoning and generation with Vision-Language Models

Qinwu Xu, Xiaofu Ma, Yifan Jiang

We study whether multimodal large language models (MLLMs) can leverage crystallographic plane indices (Miller indices) as a structured latent representation for reasoning about fra…

cs.CV2026

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

Qinwu Xu, Yifan Jiang, Haoyu Ren

Optical character recognition (OCR) and multilingual scene-text understanding remain challenging for multimodal large language models (MLLMs), particularly in real-world images con…

cs.LG2026

Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking

Qinwu Xu, Zhuoheng Li, Jessie Salas

Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and downstream evaluation signals are noisy.…

cs.CV2026

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

Qinwu Xu

Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or…