8 papers
ELDiff: When Evidential Learning Meets Text-to-Image Diffusion
Qingtao Pan, Kai Ye, Zhihao Dou +2
In multi-object text-to-image (T2I) diffusion, ensuring semantic consistency between textual prompts and generated visual content is crucial for image synthesis. However, such cons…
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning
Qinjian Zhao, Zhihao Dou, Dinggen Zhang +10
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities of large language models. However, existi…
CPS4: Class Prompt driven Semi-Supervised Spine Segmentation with Class-specific Consistency Constraint
Qingtao Pan, Hongzan Sun, Bing Ji +1
Vision Language Model (VLM) has great potential to enhance the quality of pseudo labels in semi-supervised spine segmentation by leveraging textual class prompts to generate segmen…
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
Zhihao Dou, Qinjian Zhao, Zhongwei Wan +10
Large language models (LLMs) demonstrate strong reasoning abilities via Chain-of-Thought (CoT), but their token-level generation encourages local decisions and lacks global plannin…
Conditional Factuality Controlled LLMs with Generalization Certificates via Conformal Sampling
Kai Ye, Qingtao Pan, Shuo Li
Large language models (LLMs) need reliable test-time control of hallucinations. Existing conformal methods for LLMs typically provide only \emph{marginal} guarantees and rely on a…
Frequency-Modulated Visual Restoration for Matryoshka Large Multimodal Models
Qingtao Pan, Zhihao Dou, Shuo Li
Large Multimodal Models (LMMs) struggle to adapt varying computational budgets due to numerous visual tokens. Previous methods attempted to reduce the number of visual tokens befor…