1 citations · 1 across the 4 of their papers we have counts for
5 papers · 1 filter
CBDiff:Conditional Bernoulli Diffusion Models for Image Forgery Localization
Zhou Lei, Pan Gang, Wang Jiahao +1
Image Forgery Localization (IFL) is a crucial task in image forensics, aimed at accurately identifying manipulated or tampered regions within an image at the pixel level. Existing…
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
Hongze Mi, Jinyuan Li, Xuying Zhang +4
Multimodal entity linking (MEL), a task aimed at linking mentions within multimodal contexts to their corresponding entities in a knowledge base (KB), has attracted much attention…
Sewer Image Super-Resolution with Depth Priors and Its Lightweight Network
Gang Pan, Chen Wang, Zhijie Sui +5
The Quick-view (QV) technique serves as a primary method for detecting defects within sewerage systems. However, the effectiveness of QV is impeded by the limited visual range of i…
LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model
Hongen Liu, Di Sun, Jiahao Wang +2
Video text spotting (VTS) aims to simultaneously localize, recognize and track text instances in videos. To address the limited recognition capability of end-to-end methods, recent…
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
Jinyuan Li, Han Li, Di Sun +4
Grounded Multimodal Named Entity Recognition (GMNER) is a nascent multimodal task that aims to identify named entities, entity types and their corresponding visual regions. GMNER t…