1 paper
Binh M. Le, Shaoyuan Xu, Jinmiao Fu +6
In Visual Document Understanding (VDU) tasks, fine-tuning a pre-trained Vision-Language Model (VLM) with new datasets often falls short in optimizing the vision encoder to identify…