Company2Vec -- German Company Embeddings based on Corporate Websites
- URL: http://arxiv.org/abs/2307.09332v1
- Date: Tue, 18 Jul 2023 15:14:09 GMT
- Title: Company2Vec -- German Company Embeddings based on Corporate Websites
- Authors: Christopher Gerling
- Abstract summary: The paper proposes a novel application in representation learning with Company2Vec.
The model analyzes business activities from unstructured company website data using Word2Vec and dimensionality reduction.
Company2Vec maintains semantic language structures and thus creates efficient company embeddings in fine-granular industries.
- Score: 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: With Company2Vec, the paper proposes a novel application in representation
learning. The model analyzes business activities from unstructured company
website data using Word2Vec and dimensionality reduction. Company2Vec maintains
semantic language structures and thus creates efficient company embeddings in
fine-granular industries. These semantic embeddings can be used for various
applications in banking. Direct relations between companies and words allow
semantic business analytics (e.g. top-n words for a company). Furthermore,
industry prediction is presented as a supervised learning application and
evaluation method. The vectorized structure of the embeddings allows measuring
companies similarities with the cosine distance. Company2Vec hence offers a
more fine-grained comparison of companies than the standard industry labels
(NACE). This property is relevant for unsupervised learning tasks, such as
clustering. An alternative industry segmentation is shown with k-means
clustering on the company embeddings. Finally, this paper proposes three
algorithms for (1) firm-centric, (2) industry-centric and (3) portfolio-centric
peer-firm identification.
Related papers
- Interpretable Company Similarity with Sparse Autoencoders [0.0]
We show howparse Autoencoders (SAEs) can enhance the interpretability of Large Language Models (LLMs)
We benchmark SAE features against SIC-codes, Major Group codes, and Embeddings.
Our results demonstrate that SAE features not only replicate but often surpass sector classifications and embeddings in capturing fundamental company characteristics.
arXiv Detail & Related papers (2024-12-03T17:34:50Z) - Contextualization Distillation from Large Language Model for Knowledge
Graph Completion [51.126166442122546]
We introduce the Contextualization Distillation strategy, a plug-in-and-play approach compatible with both discriminative and generative KGC frameworks.
Our method begins by instructing large language models to transform compact, structural triplets into context-rich segments.
Comprehensive evaluations across diverse datasets and KGC techniques highlight the efficacy and adaptability of our approach.
arXiv Detail & Related papers (2024-01-28T08:56:49Z) - Coherent Entity Disambiguation via Modeling Topic and Categorical
Dependency [87.16283281290053]
Previous entity disambiguation (ED) methods adopt a discriminative paradigm, where prediction is made based on matching scores between mention context and candidate entities.
We propose CoherentED, an ED system equipped with novel designs aimed at enhancing the coherence of entity predictions.
We achieve new state-of-the-art results on popular ED benchmarks, with an average improvement of 1.3 F1 points.
arXiv Detail & Related papers (2023-11-06T16:40:13Z) - CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification [1.7156312157033258]
We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations.
Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings.
15 different inter-company relations result in 51.06 million weighted edges.
arXiv Detail & Related papers (2023-06-18T23:45:15Z) - Description-Enhanced Label Embedding Contrastive Learning for Text
Classification [65.01077813330559]
Self-Supervised Learning (SSL) in model learning process and design a novel self-supervised Relation of Relation (R2) classification task.
Relation of Relation Learning Network (R2-Net) for text classification, in which text classification and R2 classification are treated as optimization targets.
external knowledge from WordNet to obtain multi-aspect descriptions for label semantic learning.
arXiv Detail & Related papers (2023-06-15T02:19:34Z) - Company classification using zero-shot learning [0.0]
We propose an approach for company classification using NLP and zero-shot learning.
We evaluate our approach on a dataset obtained through the Wharton Research Data Services (WRDS)
arXiv Detail & Related papers (2023-05-01T18:36:06Z) - Investigating Graph Structure Information for Entity Alignment with
Dangling Cases [31.779386064600956]
Entity alignment aims to discover the equivalent entities in different knowledge graphs (KGs)
We propose a novel entity alignment framework called Weakly-optimal Graph Contrastive Learning (WOGCL)
We show that WOGCL outperforms the current state-of-the-art methods with pure structural information in both traditional (relaxed) and dangling settings.
arXiv Detail & Related papers (2023-04-10T17:24:43Z) - InfoCSE: Information-aggregated Contrastive Learning of Sentence
Embeddings [61.77760317554826]
This paper proposes an information-d contrastive learning framework for learning unsupervised sentence embeddings, termed InfoCSE.
We evaluate the proposed InfoCSE on several benchmark datasets w.r.t the semantic text similarity (STS) task.
Experimental results show that InfoCSE outperforms SimCSE by an average Spearman correlation of 2.60% on BERT-base, and 1.77% on BERT-large.
arXiv Detail & Related papers (2022-10-08T15:53:19Z) - Stock2Vec: An Embedding to Improve Predictive Models for Companies [0.5872014229110215]
We create an embedding of company stocks, Stock2Vec, which can be easily added to any prediction model.
We then conduct comprehensive experiments to evaluate this embedding in applied machine learning problems.
Our experiment results demonstrate that the four features in the Stock2Vec embedding can readily augment existing cross-company models.
arXiv Detail & Related papers (2022-01-27T02:57:01Z) - R$^2$-Net: Relation of Relation Learning Network for Sentence Semantic
Matching [58.72111690643359]
We propose a Relation of Relation Learning Network (R2-Net) for sentence semantic matching.
We first employ BERT to encode the input sentences from a global perspective.
Then a CNN-based encoder is designed to capture keywords and phrase information from a local perspective.
To fully leverage labels for better relation information extraction, we introduce a self-supervised relation of relation classification task.
arXiv Detail & Related papers (2020-12-16T13:11:30Z) - Cascaded Semantic and Positional Self-Attention Network for Document
Classification [9.292885582770092]
We propose a new architecture to aggregate the two sources of information using cascaded semantic and positional self-attention network (CSPAN)
The CSPAN uses a semantic self-attention layer cascaded with Bi-LSTM to process the semantic and positional information in a sequential manner, and then adaptively combine them together through a residue connection.
We evaluate the CSPAN model on several benchmark data sets for document classification with careful ablation studies, and demonstrate the encouraging results compared with state of the art.
arXiv Detail & Related papers (2020-09-15T15:02:28Z) - A Corpus Study and Annotation Schema for Named Entity Recognition and
Relation Extraction of Business Products [68.26059718611914]
We present a corpus study, an annotation schema and associated guidelines, for the annotation of product entity and company-product relation mentions.
We find that although product mentions are often realized as noun phrases, defining their exact extent is difficult due to high boundary ambiguity.
We present a preliminary corpus of English web and social media documents annotated according to the proposed guidelines.
arXiv Detail & Related papers (2020-04-07T11:45:22Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.