activity
20232026
most citedjina-embeddings-v3: Multilingual Embeddings With Task LoRA

17 citations · 43 across the 22 of their papers we have counts for

collaborators

22 papers

cs.LG2026

TV-Regulated OPD: Direction Matters in On-Policy Distillation

Han Xiao, Yifan Niu, Dongyi Liu +2

On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervisio…

cs.CL2026

Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards

Alejandro Barón García, Feng Wang, Emilia Garcia Casademont +1

We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of D…

cs.IR2026

omni-macos: On-Device Omni-Modal Search on Apple Silicon

Han Xiao

A search engine that embeds text, code, documents, images, audio and video into the same representation space has to run its encoder and keep its index somewhere, and almost every…

cs.LG2026

Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization

Hao Wang, Kun Yuan, Wenlin Zhong +4

Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD).…

cs.IR2026

jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation

Christina Nasika, Feng Wang, Antonis Krasakis +1

Listwise rerankers are the discriminative core of agentic retrieval pipelines, yet production deployment demands efficiency, domain robustness, and fluency on semi-structured data…

cs.LG2026

Test-Time Compute for Frozen Embedding Models through Agentic Program Search

Han Xiao

Test-time compute is widely believed to benefit only large reasoning models, leaving small models with nothing to gain. We argue the opposite for dense retrieval, since modern smal…