From the 1 of 7 linked papers with an AI index.
8 papers
MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning
Zirui Cheng, Xun Xu, Tiankai Chen +7
Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the…
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
Guoxuan Chen, Chufeng Xiao, Haoran Yang +30
Boogu-Image-0.1 is an open-source multimodal model family that supports high-quality text-to-image generation, fast inference, instruction-based image editing, and bilingual (Chine…
SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Token
Ming Ma, Bowen Zheng, Zhongqiao Lin +1
Intermediate-layer predictions in large language models (LLMs) are informative but hard to decode accurately, especially at early layers. Existing lens-style methods typically rely…
VII: Visual Instruction Injection for Jailbreaking Image-to-Video Generation Models
Bowen Zheng, Yongli Xiang, Ziming Hong +4
Image-to-Video (I2V) generation models, which condition video generation on reference images, have shown emerging visual instruction-following capability, allowing certain visual c…
Label Words as Local Task Vectors in In-Context Learning
Bowen Zheng, Ming Ma, Zhongqiao Lin +1
Large Language Models (LLMs) have demonstrated remarkable abilities, one of the most important being in-context learning (ICL). With ICL, LLMs can derive the underlying rule from a…
Integrated Pipeline for Monocular 3D Reconstruction and Finite Element Simulation in Industrial Applications
Bowen Zheng
To address the challenges of 3D modeling and structural simulation in industrial environment, such as the difficulty of equipment deployment, and the difficulty of balancing accura…