collaborators

5 papers

cs.IR2026

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

Amirhossein Yousefiramandi, Ciaran Cooney

Two questions regarding practitioners' use of patent embeddings arise: (i) Does one fine-tuning recipe suffice for all downstream applications? (ii) Is fine-tuning on one patent la…

cs.AI2026

When Does Synthetic Patent Data Help? Volume-Fidelity Trade-offs in Low-Resource Multi-Label Classification

Amirhossein Yousefiramandi, Ciaran Cooney

The issues that must be considered regarding the utilization of synthetic data generated through LLMs for multilabel patent classification include (i) when the use of such data may…

cs.CL2025

Fine-Tuning Causal LLMs for Text Classification: Embedding-Based vs. Instruction-Based Approaches

Amirhossein Yousefiramandi, Ciaran Cooney

We explore efficient strategies to fine-tune decoder-only Large Language Models (LLMs) for downstream text classification under resource constraints. Two approaches are investigate…

cs.CL2025

Patent Language Model Pretraining with ModernBERT

Amirhossein Yousefiramandi, Ciaran Cooney

Transformer-based language models such as BERT have become foundational in NLP, yet their performance degrades in specialized domains like patents, which contain long, technical, a…

cs.CL2022

Unimodal and Multimodal Representation Training for Relation Extraction

Ciaran Cooney, Rachel Heyburn, Liam Madigan +3

Multimodal integration of text, layout and visual information has achieved SOTA results in visually rich document understanding (VrDU) tasks, including relation extraction (RE). Ho…