◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Yu Tian

4 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author4

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.CV4
ORCID 0000-0002-2810-1482
same name
  • Yu Tian — 30 papers, h 24
  • Yu Tian — 7 papers
  • Yu Tian — 7 papers, h 13
  • Yu Tian — 5 papers, h 15
  • Yu Tian — 4 papers, h 6
  • Yu Tian — 3 papers

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedWorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

4 citations · 8 across the 4 of their papers we have counts for

collaborators

4 papers

cs.CV2024★ 1 cited

VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding

Chris Kelly, Luhui Hu, Jiayin Hu +7

The evolution of text to visual components facilitates people's daily lives, such as generating image, videos from text and identifying the desired elements within the images. Comp…

cs.CV2024★ 3 cited

VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Bang Yang +7

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achie…

cs.CV2024★ 4 cited

WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs

Deshun Yang, Luhui Hu, Yu Tian +5

Several text-to-video diffusion models have demonstrated commendable capabilities in synthesizing high-quality video content. However, it remains a formidable challenge pertaining…

cs.CV2023

UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework

Chris Kelly, Luhui Hu, Cindy Yang +6

In the current landscape of artificial intelligence, foundation models serve as the bedrock for advancements in both language and vision domains. OpenAI GPT-4 has emerged as the pi…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.