1 citations · 1 across the 2 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
Yipeng Zhang, Yifan Liu, Zonghao Guo +10
Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-g…
cs.CV2024
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
Neale Ratzlaff, Man Luo, Xin Su +2
Multimodal models typically combine a powerful large language model (LLM) with a vision encoder and are then trained on multimodal data via instruction tuning. While this process a…
cs.CV2024
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
Sanyam Lakhanpal, Shivang Chopra, Vinija Jain +2
Over the past few years, Text-to-Image (T2I) generation approaches based on diffusion models have gained significant attention. However, vanilla diffusion models often suffer from…