8 papers
Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance
Yu Song, Hao Sun, Ikuko Nishikawa +1
Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq data is becoming increasingly important.…
SAM+D: Parameter-Efficient Dimensional Lifting of SAM-Family Models via Depth-Routed LoRA and Depth Shifting
Yu Song, Hao Sun, Shiyu Teng +2
Existing methods for adapting 2D foundation models such as SAM to 3D volumes either process slices independently---ignoring inter-slice context---or require substantial architectur…
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
Hao Sun, Yu Song, Shiyu Teng +2
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suff…
Drag within Prior Distribution: Text-Conditioned Point-Based Image Editing within Distribution Constraints
Haoyang Hu, Masataka Seo, Yen-Wei Chen
Diffusion-based point editing methods have gained significant traction in image editing tasks due to their ability to manipulate image semantics and fine details by applying locali…
One Framework to Rule Them All: Unifying Multimodal Tasks with LLM Neural-Tuning
Hao Sun, Yu Song, Jiaqing Liu +3
Large-scale models have exhibited remarkable capabilities across diverse domains, including automated medical services and intelligent customer support. However, as most large mode…
EPIC: Efficient Prompt Interaction for Text-Image Classification
Xinyao Yu, Hao Sun, Zeyu Ling +5
In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tas…