Showing cs.CVShow all
3 papers · 1 filter
cs.CV2024
PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
Shufan Li, Harkanwar Singh, Aditya Grover
Text-to-image (T2I) models achieve high-fidelity generation through extensive training on large datasets. However, these models may unintentionally pick up undesirable biases of th…
cs.CV2024
Mamba-ND: Selective State Space Modeling for Multi-Dimensional Data
Shufan Li, Harkanwar Singh, Aditya Grover
In recent years, Transformers have become the de-facto architecture for sequence modeling on text and a variety of multi-dimensional data, such as images and video. However, the us…
cs.CV2023
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following
Shufan Li, Harkanwar Singh, Aditya Grover
The ability to provide fine-grained control for generating and editing visual imagery has profound implications for computer vision and its applications. Previous works have explor…