11 papers
EviRerank: Adaptive Evidence Construction for Long-Document LLM Reranking
Minghan Li, Eric Gaussier, Juntao Li +1
Decoder-only LLM rerankers struggle with long documents: inference is costly and relevance signals can be diluted by irrelevant context. Motivated by a diagnostic attention analysi…
Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive Survey
Minghan Li, Xinxuan Lv, Junjie Zou +5
Modern information retrieval must reconcile short, ambiguous queries with increasingly diverse and dynamic corpora. Query expansion (QE) remains a core technique for mitigating voc…
S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA
Minghan Li, Junjie Zou, Xinxuan Lv +2
Retrieval-Augmented Generation (RAG) grounds language models in external evidence, but multi-hop question answering remains difficult because iterative pipelines must control what…
GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval
Minghan Li, Tianrui Lv, Chao Zhang +1
The semantic gap between colloquial user queries and professional legal documents presents a fundamental challenge in Legal Case Retrieval (LCR). Existing dense retrieval methods t…
NavGSim: High-Fidelity Gaussian Splatting Simulator for Large-Scale Navigation
Jiahang Liu, Yuanxing Duan, Jiazhao Zhang +4
Simulating realistic environments for robots is widely recognized as a critical challenge in robot learning, particularly in terms of rendering and physical simulation. This challe…
GenState-AI: State-Aware Dataset for Text-to-Video Retrieval on AI-Generated Videos
Minghan Li, Tongna Chen, Tianrui Lv +3
Existing text-to-video retrieval benchmarks are dominated by real-world footage where much of the semantics can be inferred from a single frame, leaving temporal reasoning and expl…