1 citations · 1 across the 2 of their papers we have counts for
4 papers · 1 filter
CBDiff:Conditional Bernoulli Diffusion Models for Image Forgery Localization
Zhou Lei, Pan Gang, Wang Jiahao +1
Image Forgery Localization (IFL) is a crucial task in image forensics, aimed at accurately identifying manipulated or tampered regions within an image at the pixel level. Existing…
VP-MEL: Visual Prompts Guided Multimodal Entity Linking
Hongze Mi, Jinyuan Li, Xuying Zhang +4
Multimodal entity linking (MEL), a task aimed at linking mentions within multimodal contexts to their corresponding entities in a knowledge base (KB), has attracted much attention…
LOGO: Video Text Spotting with Language Collaboration and Glyph Perception Model
Hongen Liu, Di Sun, Jiahao Wang +2
Video text spotting (VTS) aims to simultaneously localize, recognize and track text instances in videos. To address the limited recognition capability of end-to-end methods, recent…
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
Jinyuan Li, Han Li, Di Sun +4
Grounded Multimodal Named Entity Recognition (GMNER) is a nascent multimodal task that aims to identify named entities, entity types and their corresponding visual regions. GMNER t…