3 papers
cs.CV2026
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
Yecheng Zhang, Rong Zhao, Zhizhou Sha +10
Vision-language models (VLMs) can describe urban scenes in rich detail, yet consistently fail to produce reliable human preference labels in domain-specific tasks such as safety as…
cs.CV2025
Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
Jianyu Qi, Ding Zou, Wenrui Yan +5
Recent advances in Multimodal Large Language Models (MLLMs) have spurred significant progress in Chain-of-Thought (CoT) reasoning. Building on the success of Deepseek-R1, researche…
cs.LG2025
CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair Disentanglement
Chenrui Ma, Xi Xiao, Tianyang Wang +2
While deep generative models have significantly advanced representation learning, they may inherit or amplify biases and fairness issues by encoding sensitive attributes alongside…