◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jiajun Shen

4 papers hereh-index 4171 citations8 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CL3
  • cs.LG1
same name
  • Jiajun Shen — 4 papers, h 3
  • Jiajun Shen — 2 papers, h 1
  • Jiajun Shen — 2 papers, h 3
  • Jiajun Shen — 1 paper, h 6
  • Jiajun Shen — 1 paper, h 1
  • Jiajun Shen — 1 paper, h 1

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

works on
instruction following 1multimodal reasoning 1parametric skill learning 1prefix-tuning 1skill composition 1

From the 1 of 4 linked papers with an AI index.

collaborators

4 papers

cs.CL2026

SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge

Lucio M. Dery, Benedict Aaron Tjandra, Siavash Samiei +4

SkillSmith is a method that lets a large language model reason over both textual knowledge and prefix‑tuned model weights, enabling it to generate new parametric skill prefixes for…

cs.CL2026

Context Training with Active Information Seeking

Zeyu Huang, Adhiguna Kuncoro, Qixuan Feng +4

Most existing large language models (LLMs) are expensive to adapt after deployment, especially when a task requires newly produced information or niche domain knowledge. Recent wor…

cs.LG2026

Latent Space Communication via K-V Cache Alignment

Lucio M. Dery, Zohar Yahav, Henry Prior +3

Solving increasingly complex problems with large language models (LLMs) necessitates a move beyond individual models and towards multi-model systems that can effectively collaborat…

cs.CL2025

Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch

Arthur Douillard, Yanislav Donchev, Keith Rush +11

Training of large language models (LLMs) is typically distributed across a large number of accelerators to reduce training time. Since internal states and parameter gradients need…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.