1 citations · 1 across the 4 of their papers we have counts for
4 papers
MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
Bo Li, Ningyuan Deng, Tianyu Dong +3
Multimodal large language models (MLLMs) have shown impressive capabilities, yet they often struggle to effectively capture the fine-grained textual information within images cruci…
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning in Unified Multimodal Models
Lit Sin Tan, Junzhe Chen, Xiaolong Fu +5
Existing test-time scaling (TTS) methods for unified multimodal models (UMMs) in text-to-image (T2I) generation primarily rely on search or sampling strategies that produce only in…
REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
Bo Li, Guanzhi Deng, Ronghao Chen +5
Understanding how Large Language Models (LLMs) perform complex reasoning and their failure mechanisms is a challenge in interpretability research. To provide a measurable geometric…
MIT-10M: A Large Scale Parallel Corpus of Multilingual Image Translation
Bo Li, Shaolin Zhu, Lijie Wen
Image Translation (IT) holds immense potential across diverse domains, enabling the translation of textual content within images into various languages. However, existing datasets…