collaborators

15 papers

cs.CV2026

REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding

Boyang Li, Chenhui Gou, Jianfei Cai

Video temporal grounding (VTG) refers to the task of identifying the time interval in a video that corresponds to a given natural-language query. A common zero-shot strategy asks a…

cs.CV2026

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding

Jiayun Luo, Mir Rayat Imtiaz Hossain, Pritam Sarkar +2

Vision-Language Models (VLMs) have achieved strong performance on implicit and explicit visual grounding and related tasks. However, such abilities are generally tested on simple,…

cs.LG2025

Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge Integration

Wenju Sun, Qingyong Li, Wen Wang +3

Multi-task model merging aims to consolidate knowledge from multiple fine-tuned task-specific experts into a unified model while minimizing performance degradation. Existing method…

cs.MM2025

Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction

Qin Chao, Eunsoo Kim, Boyang Li

The movie industry is associated with an elevated level of risk, which necessitates the use of automated tools to predict box-office revenue and facilitate human decision-making. I…

cs.AI2025

GELD: A Unified Neural Model for Efficiently Solving Traveling Salesman Problems Across Different Scales

Yubin Xiao, Di Wang, Rui Cao +3

The Traveling Salesman Problem (TSP) is a well-known combinatorial optimization problem with broad real-world applications. Recent advancements in neural network-based TSP solvers…

cs.CV2025

SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation

Wenyu Zhang, Wei En Ng, Lixin Ma +5

Current vision-language models may grasp basic spatial cues and simple directions (e.g. left, right, front, back), but struggle with the multi-dimensional spatial reasoning necessa…