6 papers
PRISMR: Overcoming Parse Collapse in Multimodal Listwise Ranking via Parameterized Representation Internalization
Hao Jiang, Xin Li, Annan Wang +4
Generative listwise ranking with Large Multimodal Models (LMMs) aims to capture global list context in a single forward pass, but its effectiveness degrades in long-context multimo…
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
Xudong Lu, Xueying Li, Annan Wang +8
We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline v…
The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputs
Xin Li, Hao Jiang, Annan Wang +2
On-policy distillation (OPD) is widely used for LLM post-training. When pushed with a reward-extrapolation coefficient lambda > 1, the student can lift past the teacher in domain,…
RLPO: Residual Listwise Preference Optimization for Long-Context Review Ranking
Hao Jiang, Zhi Yang, Annan Wang +2
Review ranking is pivotal in e-commerce for prioritizing diagnostic and authentic feedback from the deluge of user-generated content. While large language models have improved sema…
MRSE: An Efficient Multi-modality Retrieval System for Large Scale E-commerce
Hao Jiang, Haoxiang Zhang, Qingshan Hou +4
Providing high-quality item recall for text queries is crucial in large-scale e-commerce search systems. Current Embedding-based Retrieval Systems (ERS) embed queries and items int…
Q-Ground: Image Quality Grounding with Large Multi-modality Models
Chaofeng Chen, Sensen Yang, Haoning Wu +6
Recent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual conten…