3 papers
cs.CV2026
Listener-Rewarded Thinking in VLMs for Image Preferences
Alexander Gambashidze, Li Pengyi, Matvey Skripkin +5
Training robust and generalizable reward models for human visual preferences is essential for aligning text-to-image and text-to-video generative models with human intent. However,…
cs.CV2025
MaxInfo: A Training-Free Key-Frame Selection Method Using Maximum Volume for Enhanced Video Understanding
Pengyi Li, Irina Abdullaeva, Alexander Gambashidze +2
Modern Video Large Language Models (VLLMs) often rely on uniform frame sampling for video understanding, but this approach frequently fails to capture critical information due to f…
cs.CL2025
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Pengyi Li, Matvey Skripkin, Alexander Zubrey +2
Large language models (LLMs) excel at reasoning, yet post-training remains critical for aligning their behavior with task goals. Existing reinforcement learning (RL) methods often…