Named entity recognition using GPT for identifying comparable companies
- URL: http://arxiv.org/abs/2307.07420v2
- Date: Sat, 23 Sep 2023 18:25:24 GMT
- Title: Named entity recognition using GPT for identifying comparable companies
- Authors: Eurico Covas
- Abstract summary: We show that using large language models (LLMs), such as GPT from OpenAI, has a much higher precision and success rate than using the standard named entity recognition (NER) methods.
We demonstrate quantitatively a higher precision rate, and show that, qualitatively, it can be used to create appropriate comparable companies peer groups which could then be used for equity valuation.
- Score: 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: For both public and private firms, comparable companies' analysis is widely
used as a method for company valuation. In particular, the method is of great
value for valuation of private equity companies. The several approaches to the
comparable companies' method usually rely on a qualitative approach to
identifying similar peer companies, which tend to use established industry
classification schemes and/or analyst intuition and knowledge. However, more
quantitative methods have started being used in the literature and in the
private equity industry, in particular, machine learning clustering, and
natural language processing (NLP). For NLP methods, the process consists of
extracting product entities from e.g., the company's website or company
descriptions from some financial database system and then to perform similarity
analysis. Here, using companies' descriptions/summaries from publicly available
companies' Wikipedia websites, we show that using large language models (LLMs),
such as GPT from OpenAI, has a much higher precision and success rate than
using the standard named entity recognition (NER) methods which use manual
annotation. We demonstrate quantitatively a higher precision rate, and show
that, qualitatively, it can be used to create appropriate comparable companies
peer groups which could then be used for equity valuation.
Related papers
- Autobidding Arena: unified evaluation of the classical and RL-based autobidding algorithms [71.47275796833235]
We present a standardized and transparent evaluation protocol for comparing classical and reinforcement learning autobidding algorithms.<n>We utilize the most recent open-source environment developed in the industry, which accurately emulates the bidding process.
arXiv Detail & Related papers (2025-10-22T08:27:56Z) - FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering [57.18367828883773]
FinAgentBench is a benchmark for evaluating agentic retrieval with multi-step reasoning in finance.<n>The benchmark consists of 26K expert-annotated examples on S&P-500 listed firms.<n>We evaluate a suite of state-of-the-art models and demonstrate how targeted fine-tuning can significantly improve agentic retrieval performance.
arXiv Detail & Related papers (2025-08-07T22:15:22Z) - Evaluating Large Language Models (LLMs) in Financial NLP: A Comparative Study on Financial Report Analysis [0.0]
Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide variety of Financial Natural Language Processing (FinNLP) tasks.<n>This study conducts a thorough comparative evaluation of five leading LLMs, GPT, Claude, Perplexity, Gemini and DeepSeek, using 10-K filings from the 'Magnificent Seven' technology companies.
arXiv Detail & Related papers (2025-07-24T20:10:27Z) - Interpretable Company Similarity with Sparse Autoencoders [0.0]
We show howparse Autoencoders (SAEs) can enhance the interpretability of Large Language Models (LLMs)
We benchmark SAE features against SIC-codes, Major Group codes, and Embeddings.
Our results demonstrate that SAE features not only replicate but often surpass sector classifications and embeddings in capturing fundamental company characteristics.
arXiv Detail & Related papers (2024-12-03T17:34:50Z) - POGEMA: A Benchmark Platform for Cooperative Multi-Agent Pathfinding [76.67608003501479]
We introduce POGEMA, a comprehensive set of tools that includes a fast environment for learning, a problem instance generator, and a visualization toolkit.
We also introduce and define an evaluation protocol that specifies a range of domain-related metrics, computed based on primary evaluation indicators.
The results of this comparison, which involves a variety of state-of-the-art MARL, search-based, and hybrid methods, are presented.
arXiv Detail & Related papers (2024-07-20T16:37:21Z) - An energy-based comparative analysis of common approaches to text
classification in the Legal domain [0.856335408411906]
Large Language Models (LLMs) are extensively adopted to address NLP problems in academia and industry.
In this work, we present a detailed comparison of LLM and traditional approaches (e.g. SVM) on the LexGLUE benchmark.
The results indicate that very often, the simplest algorithms achieve performance very close to that of large LLMs.
arXiv Detail & Related papers (2023-11-02T14:16:48Z) - Empowering Many, Biasing a Few: Generalist Credit Scoring through Large
Language Models [53.620827459684094]
Large Language Models (LLMs) have great potential for credit scoring tasks, with strong generalization ability across multiple tasks.
We propose the first open-source comprehensive framework for exploring LLMs for credit scoring.
We then propose the first Credit and Risk Assessment Large Language Model (CALM) by instruction tuning, tailored to the nuanced demands of various financial risk assessment tasks.
arXiv Detail & Related papers (2023-10-01T03:50:34Z) - LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise
Comparisons using Large Language Models [55.60306377044225]
Large language models (LLMs) have enabled impressive zero-shot capabilities across various natural language tasks.
This paper explores two options for exploiting the emergent abilities of LLMs for zero-shot NLG assessment.
For moderate-sized open-source LLMs, such as FlanT5 and Llama2-chat, comparative assessment is superior to prompt scoring.
arXiv Detail & Related papers (2023-07-15T22:02:12Z) - CompanyKG: A Large-Scale Heterogeneous Graph for Company Similarity Quantification [1.7156312157033258]
We propose and publish a knowledge graph, named CompanyKG, to represent and learn diverse company features and relations.
Specifically, 1.17 million companies are represented as nodes enriched with company description embeddings.
15 different inter-company relations result in 51.06 million weighted edges.
arXiv Detail & Related papers (2023-06-18T23:45:15Z) - Company classification using zero-shot learning [0.0]
We propose an approach for company classification using NLP and zero-shot learning.
We evaluate our approach on a dataset obtained through the Wharton Research Data Services (WRDS)
arXiv Detail & Related papers (2023-05-01T18:36:06Z) - Evaluating Machine Unlearning via Epistemic Uncertainty [78.27542864367821]
This work presents an evaluation of Machine Unlearning algorithms based on uncertainty.
This is the first definition of a general evaluation of our best knowledge.
arXiv Detail & Related papers (2022-08-23T09:37:31Z) - A Data-Driven Framework for Identifying Investment Opportunities in
Private Equity [0.0]
This paper proposes a framework for automated data-driven screening of investment opportunities.
The framework draws on data from several sources to assess the financial and managerial position of a company.
It then uses an explainable artificial intelligence (XAI) engine to suggest investment recommendations.
arXiv Detail & Related papers (2022-04-04T21:28:34Z) - Knowledge-Rich Self-Supervised Entity Linking [58.838404666183656]
Knowledge-RIch Self-Supervision ($tt KRISSBERT$) is a universal entity linker for four million UMLS entities.
Our approach subsumes zero-shot and few-shot methods, and can easily incorporate entity descriptions and gold mention labels if available.
Without using any labeled information, our method produces $tt KRISSBERT$, a universal entity linker for four million UMLS entities.
arXiv Detail & Related papers (2021-12-15T05:05:12Z) - The Benchmark Lottery [114.43978017484893]
"A benchmark lottery" describes the overall fragility of the machine learning benchmarking process.
We show that the relative performance of algorithms may be altered significantly simply by choosing different benchmark tasks.
arXiv Detail & Related papers (2021-07-14T21:08:30Z) - Few-Shot Named Entity Recognition: A Comprehensive Study [92.40991050806544]
We investigate three schemes to improve the model generalization ability for few-shot settings.
We perform empirical comparisons on 10 public NER datasets with various proportions of labeled data.
We create new state-of-the-art results on both few-shot and training-free settings.
arXiv Detail & Related papers (2020-12-29T23:43:16Z) - A Novel Classification Approach for Credit Scoring based on Gaussian
Mixture Models [0.0]
This paper introduces a new method for credit scoring based on Gaussian Mixture Models.
Our algorithm classifies consumers into groups which are labeled as positive or negative.
We apply our model with real world databases from Australia, Japan, and Germany.
arXiv Detail & Related papers (2020-10-26T07:34:27Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.