4 papers
BalancedDPO: Adaptive Multi-Metric Alignment
Dipesh Tamboli, Souradip Chakraborty, Aditya Malusare +3
Diffusion models have achieved remarkable progress in text-to-image generation, yet aligning them with human preference remains challenging due to the presence of multiple, sometim…
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
Arushi Rai, Qiang Zhang, Hanqing Zeng +5
Large language models (LLMs) exhibit strong reasoning capabilities but typically require expensive post-training to reach high performance. Recent test-time alignment methods offer…
PRIVATEEDIT: A Privacy-Preserving Pipeline for Face-Centric Generative Image Editing
Dipesh Tamboli, Vineet Punyamoorty, Atharv Pawar +1
Recent advances in generative image editing have enabled transformative applications, from professional head shot generation to avatar stylization. However, these systems often req…
Exploring System 1 and 2 communication for latent reasoning in LLMs
Julian Coda-Forno, Zhuokai Zhao, Qiang Zhang +6
Should LLM reasoning live in a separate module, or within a single model's forward pass and representational space? We study dual-architecture latent reasoning, where a fluent Base…