◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Ming-Tao Chen

6 papers hereh-index 324 citations10 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author5

Across the 5 of 6 papers where every author was matched, so the position is known.

fields
  • cs.CV4
  • cs.AI1
  • cs.LG1

identity via Semantic Scholar / OpenAlex

activity
20242026
collaborators
Showing cs.CVShow all

4 papers · 1 filter

cs.CV2026

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

Jingru Chen, Yiming Liu, Mingtao Chen +5

Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imp…

cs.CV2026

TAGRPO: Boosting GRPO on Image-to-Video Generation with Direct Trajectory Alignment

Jin Wang, Jianxiang Lu, Guangzheng Xu +10

Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video g…

cs.CV2025

Q-Save: Towards Scoring and Attribution for Generated Video Evaluation

Xiele Wu, Zicheng Zhang, Mingtao Chen +7

Evaluating AI-generated video (AIGV) quality hinges on three crucial dimensions: visual quality, dynamic quality, and text-video alignment. While numerous evaluation datasets and a…

cs.CV2024

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Zhimin Li, Jianwei Zhang, Qin Lin +42

We present Hunyuan-DiT, a text-to-image diffusion transformer with fine-grained understanding of both English and Chinese. To construct Hunyuan-DiT, we carefully design the transfo…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.