5 papers
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Jiajun Li, Mingshu Cai, Yixuan Li +5
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to perform realistic operations research (OR…
KGEdit: Ambiguity-Aware Knowledge Graphs for Training-Free Precise Video Generation and Editing
Mingshu Cai, Miao Zhang, Chenghe Yang +3
In recent years, training-free video generation has progressed remarkably. However, when handling complex textual instructions, existing methods still suffer from semantic ambiguit…
Where, What, Why: Toward Explainable 3D-GS Watermarking
Mingshu Cai, Jiajun Li, Osamu Yoshie +2
As 3D Gaussian Splatting becomes the de facto representation for interactive 3D assets, robust yet imperceptible watermarking is critical. We present a representation-native framew…
FluencyVE: Marrying Temporal-Aware Mamba with Bypass Attention for Video Editing
Mingshu Cai, Yixuan Li, Osamu Yoshie +1
Large-scale text-to-image diffusion models have achieved unprecedented success in image generation and editing. However, extending this success to video editing remains challenging…
Multi-Attribute guided Thermal Face Image Translation based on Latent Diffusion Model
Mingshu Cai, Osamu Yoshie, Yuya Ieiri
Modern surveillance systems increasingly rely on multi-wavelength sensors and deep neural networks to recognize faces in infrared images captured at night. However, most facial rec…