EfficientNet-Absolute Zero for Continuous Speech Keyword Spotting
- URL: http://arxiv.org/abs/2012.15695v1
- Date: Thu, 31 Dec 2020 16:21:27 GMT
- Title: EfficientNet-Absolute Zero for Continuous Speech Keyword Spotting
- Authors: Amir Mohammad Rostami, Ali Karimi, Mohammad Ali Akhaee
- Abstract summary: Football keyword dataset (FKD) is a new keyword spotting dataset in Persian.
This dataset contains nearly 31000 samples in 18 classes.
It is realized that EfficientNet-A0 and Resnet models outperform other models on this dataset.
- Score: 7.313613282363873
- License: http://creativecommons.org/licenses/by-nc-sa/4.0/
- Abstract: Keyword spotting is a process of finding some specific words or phrases in
recorded speeches by computers. Deep neural network algorithms, as a powerful
engine, can handle this problem if they are trained over an appropriate
dataset. To this end, the football keyword dataset (FKD), as a new keyword
spotting dataset in Persian, is collected with crowdsourcing. This dataset
contains nearly 31000 samples in 18 classes. The continuous speech synthesis
method proposed to made FKD usable in the practical application which works
with continuous speeches. Besides, we proposed a lightweight architecture
called EfficientNet-A0 (absolute zero) by applying the compound scaling method
on EfficientNet-B0 for keyword spotting task. Finally, the proposed
architecture is evaluated with various models. It is realized that
EfficientNet-A0 and Resnet models outperform other models on this dataset.
Related papers
- Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning [11.182456667123835]
We present a novel Query-by-Example (QbyE) KWS system that employs spectral-temporal attentive graph pooling and multi-task learning.
This framework aims to effectively learn speaker-invariant and linguistic-informative embeddings for QbyE KWS tasks.
arXiv Detail & Related papers (2024-08-27T03:44:57Z) - Improving Small Footprint Few-shot Keyword Spotting with Supervision on
Auxiliary Data [19.075820340282934]
We propose a framework that uses easily collectible, unlabeled reading speech data as an auxiliary source.
We then adopt multi-task learning that helps the model to enhance the representation power from out-of-domain auxiliary data.
arXiv Detail & Related papers (2023-08-31T07:29:42Z) - Towards Realistic Zero-Shot Classification via Self Structural Semantic
Alignment [53.2701026843921]
Large-scale pre-trained Vision Language Models (VLMs) have proven effective for zero-shot classification.
In this paper, we aim at a more challenging setting, Realistic Zero-Shot Classification, which assumes no annotation but instead a broad vocabulary.
We propose the Self Structural Semantic Alignment (S3A) framework, which extracts structural semantic information from unlabeled data while simultaneously self-learning.
arXiv Detail & Related papers (2023-08-24T17:56:46Z) - CompoundPiece: Evaluating and Improving Decompounding Performance of
Language Models [77.45934004406283]
We systematically study decompounding, the task of splitting compound words into their constituents.
We introduce a dataset of 255k compound and non-compound words across 56 diverse languages obtained from Wiktionary.
We introduce a novel methodology to train dedicated models for decompounding.
arXiv Detail & Related papers (2023-05-23T16:32:27Z) - XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented
Languages [105.54207724678767]
Data scarcity is a crucial issue for the development of highly multilingual NLP systems.
We propose XTREME-UP, a benchmark defined by its focus on the scarce-data scenario rather than zero-shot.
XTREME-UP evaluates the capabilities of language models across 88 under-represented languages over 9 key user-centric technologies.
arXiv Detail & Related papers (2023-05-19T18:00:03Z) - DORIC : Domain Robust Fine-Tuning for Open Intent Clustering through
Dependency Parsing [14.709084509818474]
DSTC11-Track2 aims to provide a benchmark for zero-shot, cross-domain, intent-set induction.
We leveraged a multi-domain dialogue dataset to fine-tune the language model and proposed extracting Verb-Object pairs.
Our approach achieved 3rd place in the precision score and showed superior accuracy and normalized mutual information (NMI) score than the baseline model.
arXiv Detail & Related papers (2023-03-17T08:12:36Z) - Ensemble Transfer Learning for Multilingual Coreference Resolution [60.409789753164944]
A problem that frequently occurs when working with a non-English language is the scarcity of annotated training data.
We design a simple but effective ensemble-based framework that combines various transfer learning techniques.
We also propose a low-cost TL method that bootstraps coreference resolution models by utilizing Wikipedia anchor texts.
arXiv Detail & Related papers (2023-01-22T18:22:55Z) - Model Composition: Can Multiple Neural Networks Be Combined into a
Single Network Using Only Unlabeled Data? [6.0945220518329855]
This paper investigates the idea of combining multiple trained neural networks using unlabeled data.
To this end, the proposed method makes use of generation, filtering, and aggregation of reliable pseudo-labels collected from unlabeled data.
Our method supports using an arbitrary number of input models with arbitrary architectures and categories.
arXiv Detail & Related papers (2021-10-20T04:17:25Z) - Generative Conversational Networks [67.13144697969501]
We propose a framework called Generative Conversational Networks, in which conversational agents learn to generate their own labelled training data.
We show an average improvement of 35% in intent detection and 21% in slot tagging over a baseline model trained from the seed data.
arXiv Detail & Related papers (2021-06-15T23:19:37Z) - MTOP: A Comprehensive Multilingual Task-Oriented Semantic Parsing
Benchmark [31.91964553419665]
We present a new multilingual dataset, called MTOP, comprising of 100k annotated utterances in 6 languages across 11 domains.
We achieve an average improvement of +6.3 points on Slot F1 for the two existing multilingual datasets, over best results reported in their experiments.
We demonstrate strong zero-shot performance using pre-trained models combined with automatic translation and alignment, and a proposed distant supervision method to reduce the noise in slot label projection.
arXiv Detail & Related papers (2020-08-21T07:02:11Z) - ContextNet: Improving Convolutional Neural Networks for Automatic Speech
Recognition with Global Context [58.40112382877868]
We propose a novel CNN-RNN-transducer architecture, which we call ContextNet.
ContextNet features a fully convolutional encoder that incorporates global context information into convolution layers by adding squeeze-and-excitation modules.
We demonstrate that ContextNet achieves a word error rate (WER) of 2.1%/4.6% without external language model (LM), 1.9%/4.1% with LM and 2.9%/7.0% with only 10M parameters on the clean/noisy LibriSpeech test sets.
arXiv Detail & Related papers (2020-05-07T01:03:18Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.