6 papers
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
Weide Liu, Wei Zhou, Jun Liu +4
Feature matching is a cornerstone task in computer vision, essential for applications such as image retrieval, stereo matching, 3D reconstruction, and SLAM. This survey comprehensi…
On Efficient Variants of Segment Anything Model: A Survey
Xiaorui Sun, Jun Liu, Heng Tao Shen +2
The Segment Anything Model (SAM) is a foundational model for image segmentation tasks, known for its strong generalization across diverse applications. However, its impressive perf…
Automatic Method Illustration Generation for AI Scientific Papers via Drawing Middleware Creation, Evolution, and Orchestration
Zhuoling Li, Jiarui Zhang, Ping Hu +4
Method illustrations (MIs) play a crucial role in conveying the core ideas of scientific papers, yet their generation remains a labor-intensive process. Here, we take inspiration f…
LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios
Bingxi Zhao, Lin Geng Foo, Ping Hu +3
Recent advances in the intrinsic reasoning capabilities of large language models (LLMs) have given rise to LLM-based agent systems that exhibit near-human performance on a variety…
TSTMotion: Training-free Scene-aware Text-to-motion Generation
Ziyan Guo, Haoxuan Qu, Hossein Rahmani +4
Text-to-motion generation has recently garnered significant research interest, primarily focusing on generating human motion sequences in blank backgrounds. However, human motions…
Unified Prompt Attack Against Text-to-Image Generation Models
Duo Peng, Qiuhong Ke, Mark He Huang +2
Text-to-Image (T2I) models have advanced significantly, but their growing popularity raises security concerns due to their potential to generate harmful images. To address these is…