CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification
arXiv:2306.10649 · doi:10.1109/TBDATA.2024.3407573
Abstract
In the investment industry, it is often essential to carry out fine-grained company similarity quantification for a range of purposes, including market mapping, competitor analysis, and mergers and acquisitions. We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations. Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings; and 15 different inter-company relations result in 51.06 million weighted edges. To enable a comprehensive assessment of methods for company similarity quantification, we have devised and compiled three evaluation tasks with annotated test sets: similarity prediction, competitor retrieval and similarity ranking. We present extensive benchmarking results for 11 reproducible predictive methods categorized into three groups: node-only, edge-only, and node+edge. To the best of our knowledge, CompanyKG is the first large-scale heterogeneous graph dataset originating from a real-world investment platform, tailored for quantifying inter-company similarity.
CompanyKG (version 1.x). Published by IEEE Transactions on Big Data (12 pages, 10 figures and 2 tables) + Appendix (9 pages, 1 figures and 4 tables). Code: https://github.com/EQTPartners/CompanyKG ; Data: https://zenodo.org/record/8010239
References in corpus (8)
- Training language models to follow instructions with human feedback
- LLaMA: Open and Efficient Foundation Language Models
- LaMDA: Language Models for Dialog Applications
- Deep Graph Contrastive Representation Learning
- Contrastive Multi-View Representation Learning on Graphs
- BloombergGPT: A Large Language Model for Finance
- Using Deep Learning to Find the Next Unicorn: A Practical Synthesis
- A Scalable and Adaptive System to Infer the Industry Sectors of Companies: Prompt + Model Tuning of Generative Language Models