7 papers · 1 filter
Why Sample What You Can Enumerate? Exact Policy Optimization for Genomic Tool Selection
Haoyue Liu, Xiaoyu Ma, Ye Chen +2
Reinforcement learning over a frozen reasoner has become a common recipe for teaching a policy which external tools to invoke. We show that this recipe becomes structurally mismatc…
SEPO: Evidence-Grounded Prompt Optimization via Structural Editing
Xiaoyu Ma, Haoyue Liu, Yiwen Li +4
Existing API-only prompt optimisers are often described as interpretable, but in practice, this usually means only post-hoc inspectability: each iteration still rewrites the prompt…
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
Haoyue Liu, Ye Chen, Zhichao Wang +1
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, the…
Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
Haoyue Liu, Xiaoyu Ma, Ye Chen +2
Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updates, yielding promising results. However,…
One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization
Haoyue Liu, Xiaoyu Ma, Ye Chen +2
Text-to-image (T2I) generators often fail to follow their prompts faithfully, producing wrong counts, swapped attributes, ambiguous relations, and illegible text. Prompt optimizati…
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
Miao Wang, Yuling Shi, Yijiang Li +8
Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for immersive applications such as V…