From the 1 of 15 linked papers with an AI index.
15 papers
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…
Introspective Attention Modulation for Safe Text-to-Image Generation
Basim Azam, Hossein Rahmani, Naveed Akhtar
The paper proposes a method that monitors and adjusts the attention mechanisms of text‑to‑image diffusion models at inference time to prevent the generation of unsafe content while…
LBTCap: A Lightweight Bilateral Transformer for Real-Time Remote Sensing Image Change Captioning
Licheng Zhang, Siew-Kei Lam, Naveed Akhtar
Remote sensing image change captioning (RSICC) generates natural-language descriptions of semantic changes between paired remote sensing images (RSIs), supporting applications such…
Beyond Pixel Diffs: Benchmarking Image Change Captioning for Web UI Visual Regression Testing
Licheng Zhang, Bach Le, Pengtao Zhao +1
Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interface (UI) screenshots, compares…
Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment
Soyoun Won, Aryan Yazdan Parast, Basim Azam +2
Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors due to a phenomenon referre…
HSFM: Hard-Set-Guided Feature-Space Meta-Learning for Robust Classification under Spurious Correlations
Aryan Yazdan Parast, Khawar Islam, Soyoun Won +2
Deep neural networks often rely on spurious features to make predictions, which makes them brittle under distribution shift and on samples where the spurious correlation does not h…