1 citations · 1 across the 5 of their papers we have counts for
5 papers
SANEval: Open-Vocabulary Compositional Benchmarks with Failure-mode Diagnosis
Rishav Pramanik, Ian E. Nielsen, Jeff Smith +3
The rapid progress of text-to-image (T2I) models has unlocked unprecedented creative potential, yet their ability to faithfully render complex prompts involving multiple objects, a…
ASBA: A-line State Space Model and B-line Attention for Sparse Optical Doppler Tomography Reconstruction
Zhenghong Li, Wensheng Cheng, Congwu Du +3
Optical Doppler Tomography (ODT) is an emerging blood flow analysis technique. A 2D ODT image (B-scan) is generated by sequentially acquiring 1D depth-resolved raw A-scans (A-line)…
Beyond Words: Multimodal LLM Knows When to Speak
Zikai Liao, Yi Ouyang, Yi-Lun Lee +3
Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue.…
Distilling Specialized Orders for Visual Generation
Rishav Pramanik, Amin Sghaier, Masih Aminbeidokhti +6
Autoregressive (AR) image generators are becoming increasingly popular due to their ability to produce high-quality images and their scalability. Typical AR models are locked onto…
CUPre: Cross-domain Unsupervised Pre-training for Few-Shot Cell Segmentation
Weibin Liao, Xuhong Li, Qingzhong Wang +3
While pre-training on object detection tasks, such as Common Objects in Contexts (COCO) [1], could significantly boost the performance of cell segmentation, it still consumes on ma…