papers

Publications (6)

cs.CV2026

MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals

Junyu Shen, Zhendong She, Chenghanyu Zhang +13

The paradigm of Multimodal Large Language Models (MLLMs) offers a promising blueprint for advancing the electromagnetic (EM) domain. However, prevailing approaches often deviate fr…

cs.CV2026

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

Lichen Bai, Tianhao Zhang, Shitong Shao +14

As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but…

cs.CV2026

MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

Yinhan Zhang, Dingwei Tan, Dinwei Tan +4

The paper introduces MagicPrompt, a lightweight method that uses attention-embedded soft prompts and dual-space reward feedback to fine‑tune large video diffusion models with less…

#video generation#diffusion models#prompt tuning#parameter-efficient fine-tuning
cs.AI2025

Gazing at Rewards: Eye Movements as a Lens into Human and AI Decision-Making in Hybrid Visual Foraging

Bo Wang, Dingwei Tan, Yen-Ling Kuo +4

Imagine searching a collection of coins for quarters (), dimes (), nickels (), and pennies ()-a hybrid foraging task where observers look for multiple insta…

cs.AI2024

DreamFactory: Pioneering Multi-Scene Long Video Generation with a Multi-Agent Framework

Zhifei Xie, Daniel Tang, Dingwei Tan +3

Current video generation models excel at creating short, realistic clips, but struggle with longer, multi-scene videos. We introduce \texttt{DreamFactory}, an LLM-based framework t…

cs.AI2026

PReD: An LLM-based Foundation Multimodal Model for Electromagnetic Perception, Recognition, and Decision

Zehua Han, Jing Xiao, Yiqi Duan +13

Multimodal Large Language Models have demonstrated powerful cross-modal understanding and reasoning capabilities in general domains. However, in the electromagnetic (EM) domain, th…