4 papers
M^3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
Yang Zhou, Mingyu Zhao, Zhenting Wang +6
We present M^3-Bench, the first benchmark for evaluating multimodal tool use under the Model Context Protocol. The benchmark targets realistic, multi-hop and multi-threaded workflo…
MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition
Mingyu Zhao, Zhanfu Yang, Yang Zhou +4
This paper employs a multimodal approach for continuous sign recognition by first using ML for detecting the start and end frames of signs in videos of American Sign Language (ASL)…
Large Sign Language Models: Toward 3D American Sign Language Translation
Sen Zhang, Xiaoxiao He, Di Liu +7
We present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which ca…
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
Can Jin, Ying Li, Mingyu Zhao +6
Visual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing vis…