activity
20242026
collaborators

13 papers

cs.CV2026

GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing

Qigan Sun, Chaoning Zhang, Jianwei Zhang +8

In recent years, Multimodal Large Language Models (MLLMs) have made significant progress in visual question answering tasks. However, directly applying existing fine-tuning methods…

cs.IR2026

Unleashing the Potential of Neighbors: Diffusion-based Latent Neighbor Generation for Session-based Recommendation

Yuhan Yang, Jie Zou, Guojia An +3

Session-based recommendation aims to predict the next item that anonymous users may be interested in, based on their current session interactions. Recent studies have demonstrated…

cs.CV2025

Fast SAM2 with Text-Driven Token Pruning

Avilasha Mandal, Chaoning Zhang, Fachrina Dewi Puspitasari +6

Segment Anything Model 2 (SAM2), a vision foundation model has significantly advanced in prompt-driven video object segmentation, yet their practical deployment remains limited by…

cs.CV2025

HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models

Haoxi Zeng, Haoxuan Li, Yi Bin +4

Contrastive Language-Image Pre-training (CLIP) has demonstrated remarkable generalization ability and strong performance across a wide range of vision-language tasks. However, due…

cs.LG2025

GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions

Bing Liu, Wenqiang Yv, Xuzheng Yang +6

AI-driven geometric problem solving is a complex vision-language task that requires accurate diagram interpretation, mathematical reasoning, and robust cross-modal grounding. A fou…

cs.AI2025

Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models

Jun Ling, Yao Qi, Tao Huang +8

In this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inpu…