6 papers
Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents
XiuYu Zhang, Junfeng Fang, Zhenkai Liang
Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses alignment between these lat…
Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
XiuYu Zhang, Yi Shan, Junfeng Fang +1
Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is…
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
Yi Zhang, An Zhang, XiuYu Zhang +4
Large language models (LLMs), despite possessing latent safety understanding from their vast pretraining data, remain vulnerable to generating harmful content and exhibit issues su…
CLAS: A Machine Learning Enhanced Framework for Exploring Large 3D Design Datasets
XiuYu Zhang, Xiaolei Ye, Jui-Che Chang +1
Three-dimensional (3D) objects have wide applications. Despite the growing interest in 3D modeling in academia and industries, designing and/or creating 3D objects from scratch rem…
Advancing Conversational Psychotherapy: Integrating Privacy, Dual-Memory, and Domain Expertise with Large Language Models
XiuYu Zhang, Zening Luo
Mental health has increasingly become a global issue that reveals the limitations of traditional conversational psychotherapy, constrained by location, time, expense, and privacy c…
Partially Conditioned Patch Parallelism for Accelerated Diffusion Model Inference
XiuYu Zhang, Zening Luo, Michelle E. Lu
Diffusion models have exhibited exciting capabilities in generating images and are also very promising for video creation. However, the inference speed of diffusion models is limit…