4 papers
Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
Yi Xu, Yesheng Zhang, Jiajia Liu +1
Multimodal large language models (MLLMs) have emerged as pivotal tools in enhancing human-computer interaction. In this paper we focus on the application of MLLMs in the field of g…
Trajectory Entropy: Modeling Game State Stability from Multimodality Trajectory Prediction
Yesheng Zhang, Wenjian Sun, Yuheng Chen +4
Complex interactions among agents present a significant challenge for autonomous driving in real-world scenarios. Recently, a promising approach has emerged, which formulates the i…
MESA: Effective Matching Redundancy Reduction by Semantic Area Segmentation
Yesheng Zhang, Shuhan Shen, Xu Zhao
We propose MESA and DMESA as novel feature matching methods, which utilize Segment Anything Model (SAM) to effectively mitigate matching redundancy. The key insight of our methods…
Searching from Area to Point: A Hierarchical Framework for Semantic-Geometric Combined Feature Matching
Yesheng Zhang, Xu Zhao
Feature matching is a crucial technique in computer vision. A unified perspective for this task is to treat it as a searching problem, aiming at an efficient search strategy to nar…