◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ming-Hsuan Yang

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author2
  • last author1

Across the 3 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CV4
same name
  • Ming-Hsuan Yang — 90 papers, h 72
  • Ming-Hsuan Yang — 55 papers, h 127
  • Ming-Hsuan Yang — 40 papers, h 14
  • Ming-Hsuan Yang — 18 papers, h 5
  • Ming-Hsuan Yang — 12 papers, h 4
  • Ming-Hsuan Yang — 12 papers, h 5

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedModule-wise Adaptive Distillation for Multimodality Foundation Models

4 citations · 4 across the 1 of their papers we have counts for

collaborators

4 papers

cs.CV2024

Cropper: Vision-Language Model for Image Cropping through In-Context Learning

Seung Hyun Lee, Jijun Jiang, Yiran Xu +10

The goal of image cropping is to identify visually appealing crops in an image. Conventional methods are trained on specific datasets and fail to adapt to new requirements. Recent…

cs.CV2023

Text-Driven Image Editing via Learnable Regions

Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai +2

Language has emerged as a natural interface for image editing. In this paper, we introduce a method for region-based image editing driven by textual prompts, without the need for u…

cs.CV2023★ 4 cited

Module-wise Adaptive Distillation for Multimodality Foundation Models

Chen Liang, Jiahui Yu, Ming-Hsuan Yang +5

Pre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reduci…

cs.CV2023

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Lijun Yu, José Lezama, Nitesh B. Gundavarapu +13

While Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effec…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.