◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Shimin Yao

4 papers hereh-index 293 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author3

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CV4

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.CV2026

HumanOmni-Speaker: Identifying Who said What and When

Detao Bai, Zhiheng Ma, Xihan Wei

While Omni-modal Large Language Models have made strides in joint sensory processing, they fundamentally struggle with a cornerstone of human interaction: deciphering complex, mult…

cs.CV2026

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

Detao Bai, Shimin Yao, Weixuan Chen +4

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specifi…

cs.CV2025

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Qize Yang, Shimin Yao, Weixuan Chen +7

With the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands d…

cs.CV2025

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding

Jiaxing Zhao, Qize Yang, Yixing Peng +8

In human-centric scenes, the ability to simultaneously understand visual and auditory information is crucial. While recent omni models can process multiple modalities, they general…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.