2 papers
cs.CV2026
T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation
Yan Wang, Xinyi Hou, Weiguo Lin +2
Recent text-to-image models have become increasingly capable of rendering explicit text, but reliable localized text control requires more than generating the correct string. In ap…
cs.CV2026
MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
Yingying Fan, Penghui Du, Leyan Zhu +10
Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the quest…