Publications (11)
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
Jun Zhan, Mingyang Han, Yuxuan Xie +11
Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress h…
On the super-Liouville equations on the sphere
Mingyang Han, Chunqin Zhou
In this paper, we investigate the existence of nontrivial least-energy solutions for the super-Liouville equation with positive coefficient functions on the two-dimensional sphere.…
Visual-CoG: Stage-Aware Reinforcement Learning with Chain of Guidance for Text-to-Image Generation
Yaqi Li, Peng Chen, Mingyang Han +7
Despite the promising progress of recent autoregressive models in text-to-image (T2I) generation, their ability to handle multi-attribute and ambiguous prompts remains limited. To…
High energy solutions of quadratic coupling Schrodinger equation with nonconstant potential
Mingyang Han, Kai Zhang
In this paper, we use the variational method, especially the perturbation method, to find the perturbed high energy solutions of the quadratic coupled Schrodinger system with asymm…
Normalized solution to coupled nonhomogeneous nonlinear elliptic system with three wave interaction under unbounded potentials
Mingyang Han
In this paper, we use the variational method to find the normalized solutions of the quadratic coupled three wave Schrodinger equation with asymmetric coercive potential. We prove…
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Tianhong Zhou, Mingyang Han, Boyu Li +8
Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…