collaborators

7 papers

cs.LG2026

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

Minglai Yang, Xinyu Guo, Utkarsh Tyagi +6

Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer. The…

cs.CL2026

The Answer Lies Within: Self-Derived Rewards Enable Explainable Relation Extraction

Xinyu Guo, Zhengliang Shi, Minglai Yang +1

Despite the remarkable reasoning capabilities of large language models, they still struggle with one-shot relation extraction without predefined relation labels. We identify two pi…

cs.CV2026

AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

Shuhan Xia, Peipei Li, Xuannan Liu +3

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing ben…

cs.CV2026

TrackNetV5: Residual-Driven Spatio-Temporal Refinement and Motion Direction Decoupling for Fast Object Tracking

Haonan Tang, Yanjun Chen, Lezhi Jiang +2

The TrackNet series has established a strong baseline for fast-moving small object tracking in sports. However, existing iterations face significant limitations: V1-V3 struggle wit…

cs.LG2026

AlignSAE: Concept-Aligned Sparse Autoencoders

Minglai Yang, Xinyu Guo, Zhengliang Shi +4

Large Language Models (LLMs) encode factual knowledge within hidden parametric spaces that are difficult to inspect or control. While Sparse Autoencoders (SAEs) can decompose hidde…

cs.CL2025

Deep Research: A Systematic Survey

Zhengliang Shi, Yiqun Chen, Haitao Li +23

Large language models (LLMs) have rapidly evolved from text generators into powerful problem solvers. Yet, many open tasks demand critical thinking, multi-source, and verifiable ou…