44 citations · 106 across the 14 of their papers we have counts for
12 papers · 1 filter
Visual Text Compression as Measure Transport
Lv Tang, Tianyi Zheng, Yang Liu +2
Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing --$20\t…
FractalMamba++: Scaling Vision Mamba Across Resolutions via Hilbert Fractal Geometry
Bo Li, Haoke Xiao, Lv Tang
Vision Mamba offers linear complexity for long visual sequences, yet its performance depends critically on how a two-dimensional patch grid is serialized into a one-dimensional sta…
Evaluating SAM2's Role in Camouflaged Object Detection: From SAM to SAM2
Lv Tang, Bo Li
The Segment Anything Model (SAM), introduced by Meta AI Research as a generic object segmentation model, quickly garnered widespread attention and significantly influenced the acad…
Scalable Visual State Space Model with Fractal Scanning
Lv Tang, HaoKe Xiao, Peng-Tao Jiang +3
Foundational models have significantly advanced in natural language processing (NLP) and computer vision (CV), with the Transformer architecture becoming a standard backbone. Howev…
ASAM: Boosting Segment Anything Model with Adversarial Tuning
Bo Li, Haoke Xiao, Lv Tang
In the evolving landscape of computer vision, foundation models have emerged as pivotal tools, exhibiting exceptional adaptability to a myriad of tasks. Among these, the Segment An…
Towards Training-free Open-world Segmentation via Image Prompt Foundation Models
Lv Tang, Peng-Tao Jiang, Hao-Ke Xiao +1
The realm of computer vision has witnessed a paradigm shift with the advent of foundational models, mirroring the transformative influence of large language models in the domain of…