A New Approach to Multilabel Stratified Cross Validation with
Application to Large and Sparse Gene Ontology Datasets
- URL: http://arxiv.org/abs/2109.01425v1
- Date: Fri, 3 Sep 2021 10:34:22 GMT
- Title: A New Approach to Multilabel Stratified Cross Validation with
Application to Large and Sparse Gene Ontology Datasets
- Authors: Henri Tiittanen, Liisa Holm and Petri T\"or\"onen
- Abstract summary: We show a weakness in an evaluation metric widely used in literature.
We present improved versions of this metric and a general method, optisplit, for optimising cross validations splits.
We show that optisplit produces better cross validation splits than the existing methods and that it is fast enough to be used on big Gene Ontology datasets.
- Score: 0.0
- License: http://creativecommons.org/licenses/by/4.0/
- Abstract: Multilabel learning is an important topic in machine learning research.
Evaluating models in multilabel settings requires specific cross validation
methods designed for multilabel data. In this article, we show a weakness in an
evaluation metric widely used in literature and we present improved versions of
this metric and a general method, optisplit, for optimising cross validations
splits. We present an extensive comparison of various types of cross validation
methods in which we show that optisplit produces better cross validation splits
than the existing methods and that it is fast enough to be used on big Gene
Ontology (GO) datasets
Related papers
- LC-Protonets: Multi-label Few-shot learning for world music audio tagging [65.72891334156706]
We introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification.
LC-Protonets generate one prototype per label combination, derived from the power set of labels present in the limited training items.
Our method is applied to automatic audio tagging across diverse music datasets, covering various cultures and including both modern and traditional music.
arXiv Detail & Related papers (2024-09-17T15:13:07Z) - Dual-Decoupling Learning and Metric-Adaptive Thresholding for Semi-Supervised Multi-Label Learning [81.83013974171364]
Semi-supervised multi-label learning (SSMLL) is a powerful framework for leveraging unlabeled data to reduce the expensive cost of collecting precise multi-label annotations.
Unlike semi-supervised learning, one cannot select the most probable label as the pseudo-label in SSMLL due to multiple semantics contained in an instance.
We propose a dual-perspective method to generate high-quality pseudo-labels.
arXiv Detail & Related papers (2024-07-26T09:33:53Z) - CrossMatch: Enhance Semi-Supervised Medical Image Segmentation with Perturbation Strategies and Knowledge Distillation [7.6057981800052845]
CrossMatch is a novel framework that integrates knowledge distillation with dual strategies-image-level and feature-level to improve the model's learning from both labeled and unlabeled data.
Our method significantly surpasses other state-of-the-art techniques in standard benchmarks by effectively minimizing the gap between training on labeled and unlabeled data.
arXiv Detail & Related papers (2024-05-01T07:16:03Z) - Multi-Label Feature Selection Using Adaptive and Transformed Relevance [0.0]
This paper presents a novel information-theoretical filter-based multi-label feature selection, called ATR, with a new function.
ATR ranks features considering individual labels as well as abstract label space discriminative powers.
Our experiments affirm the scalability of ATR for benchmarks characterized by extensive feature and label spaces.
arXiv Detail & Related papers (2023-09-26T09:01:38Z) - Convolutional autoencoder-based multimodal one-class classification [80.52334952912808]
One-class classification refers to approaches of learning using data from a single class only.
We propose a deep learning one-class classification method suitable for multimodal data.
arXiv Detail & Related papers (2023-09-25T12:31:18Z) - Retrieval-augmented Multi-label Text Classification [20.100081284294973]
Multi-label text classification is a challenging task in settings of large label sets.
Retrieval augmentation aims to improve the sample efficiency of classification models.
We evaluate this approach on four datasets from the legal and biomedical domains.
arXiv Detail & Related papers (2023-05-22T14:16:23Z) - Reliable Representations Learning for Incomplete Multi-View Partial Multi-Label Classification [78.15629210659516]
In this paper, we propose an incomplete multi-view partial multi-label classification network named RANK.
We break through the view-level weights inherent in existing methods and propose a quality-aware sub-network to dynamically assign quality scores to each view of each sample.
Our model is not only able to handle complete multi-view multi-label datasets, but also works on datasets with missing instances and labels.
arXiv Detail & Related papers (2023-03-30T03:09:25Z) - Efficient Classification of Long Documents Using Transformers [13.927622630633344]
We evaluate the relative efficacy measured against various baselines and diverse datasets.
Results show that more complex models often fail to outperform simple baselines and yield inconsistent performance across datasets.
arXiv Detail & Related papers (2022-03-21T18:36:18Z) - Multiway sparse distance weighted discrimination [3.574492630046327]
Distance weighted discrimination (DWD) is a popular high-dimensional classification method that has been extended to the multiway context.
We develop a general framework for multiway classification which is applicable to any number of dimensions and any degree of sparsity.
arXiv Detail & Related papers (2021-10-11T16:11:04Z) - Gated recurrent units and temporal convolutional network for multilabel
classification [122.84638446560663]
This work proposes a new ensemble method for managing multilabel classification.
The core of the proposed approach combines a set of gated recurrent units and temporal convolutional neural networks trained with variants of the Adam gradients optimization approach.
arXiv Detail & Related papers (2021-10-09T00:00:16Z) - Interaction Matching for Long-Tail Multi-Label Classification [57.262792333593644]
We present an elegant and effective approach for addressing limitations in existing multi-label classification models.
By performing soft n-gram interaction matching, we match labels with natural language descriptions.
arXiv Detail & Related papers (2020-05-18T15:27:55Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.