activity
20242026
most citedUncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2026

Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM

SooHwan Eom, Jay Shim, Gwanhyeong Koo +4

The Transformer's quadratic complexity with input length imposes an unsustainable computational load on large language models (LLMs). In contrast, the Selective Scan Structured Sta…

cs.LG20261 cited

Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning

Thanh Nguyen, Tung Luu, Tri Ton +2

Offline reinforcement learning (RL) has garnered significant interest due to its safe and easily scalable paradigm. However, training under this paradigm presents its own challenge…

cs.CL2025

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization

Hee Suk Yoon, Eunseop Yoon, Mark Hasegawa-Johnson +2

We introduce ConfPO, a method for preference learning in Large Language Models (LLMs) that identifies and optimizes preference-critical tokens based solely on the training policy's…

cs.CL2024

TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback

Eunseop Yoon, Hee Suk Yoon, SooHwan Eom +7

Reinforcement Learning from Human Feedback (RLHF) leverages human preference data to train language models to align more closely with human essence. These human preference data, ho…

cs.LG2024

Mitigating Adversarial Perturbations for Deep Reinforcement Learning via Vector Quantization

Tung M. Luu, Thanh Nguyen, Tee Joshua Tian Jin +2

Recent studies reveal that well-performing reinforcement learning (RL) agents in training often lack resilience against adversarial perturbations during deployment. This highlights…