4 papers
Generating Concept Lexicalizations via Dictionary-Based Cross-Lingual Sense Projection
David Basil, Chirooth Girigowda, Bradley Hauer +3
We study the task of automatically expanding WordNet-style lexical resources to new languages through sense generation. We generate senses by associating target-language lemmas wit…
MIO: A Foundation Model on Multimodal Tokens
Zekun Wang, King Zhu, Chunpu Xu +14
In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, aut…
Cross-Modal Consistency in Multimodal Large Language Models
Xiang Zhang, Senyu Li, Ning Shi +5
Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual…
Optimizing Negative Prompts for Enhanced Aesthetics and Fidelity in Text-To-Image Generation
Michael Ogezi, Ning Shi
In text-to-image generation, using negative prompts, which describe undesirable image characteristics, can significantly boost image quality. However, producing good negative promp…