6 papers
KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model
Jie Yang, Wang Zeng, Sheng Jin +5
The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle wit…
NADER: Neural Architecture Design via Multi-Agent Collaboration
Zekang Yang, Wang Zeng, Sheng Jin +3
Designing effective neural architectures poses a significant challenge in deep learning. While Neural Architecture Search (NAS) automates the search for optimal architectures, exis…
AutoMMLab: Automatically Generating Deployable Models from Language Instructions for Computer Vision Tasks
Zekang Yang, Wang Zeng, Sheng Jin +3
Automated machine learning (AutoML) is a collection of techniques designed to automate the machine learning development process. While traditional AutoML approaches have been succe…
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
Wangyu Xue, Chen Qian, Jiayi Wu +5
Existing works on human-centric video understanding typically focus on analyzing specific moment or entire videos. However, many applications require higher precision at the frame…
CAS-ViT: Convolutional Additive Self-attention Vision Transformers for Efficient Mobile Applications
Tianfang Zhang, Lei Li, Yang Zhou +4
Vision Transformers (ViTs) mark a revolutionary advance in neural networks with their token mixer's powerful global context capability. However, the pairwise token affinity and com…
KptLLM: Unveiling the Power of Large Language Model for Keypoint Comprehension
Jie Yang, Wang Zeng, Sheng Jin +4
Recent advancements in Multimodal Large Language Models (MLLMs) have greatly improved their abilities in image understanding. However, these models often struggle with grasping pix…