15 papers
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu +1
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associatio…
Instant Personalized Large Language Model Adaptation via Hypernetwork
Zhaoxuan Tan, Zixuan Zhang, Haoyang Wen +8
Personalized large language models (LLMs) tailor content to individual preferences using user profiles or histories. However, existing parameter-efficient fine-tuning (PEFT) method…
3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation
Nicole Meng, Zheyuan Liu, Meng Jiang +1
Recent advances in 3D generative editing, particularly pipelines based on 3D Gaussian Splatting (3DGS), have achieved high-fidelity, multi-view-consistent scene manipulation from t…
Teaching Large Language Models When Not to Know: Learning Temporal Critique for Ex-Ante Reasoning
Chenlu Ding, Jiancan Wu, Yanchen Luo +3
Large language models (LLMs) often fail to reason under temporal cutoffs: when prompted to answer from the standpoint of an earlier time, they exploit knowledge that became availab…
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
Diancheng Kang, Zheyuan Liu, Ningshan Ma +3
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful…
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
Mengzhao Jia, Zhihan Zhang, Ignacio Cases +3
Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) ofte…