Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
- URL: http://arxiv.org/abs/2409.00099v2
- Date: Sat, 23 Nov 2024 20:55:13 GMT
- Title: Query-by-Example Keyword Spotting Using Spectral-Temporal Graph Attentive Pooling and Multi-Task Learning
- Authors: Zhenyu Wang, Shuyu Kong, Li Wan, Biqiao Zhang, Yiteng Huang, Mumin Jin, Ming Sun, Xin Lei, Zhaojun Yang,
- Abstract summary: We present a novel Query-by-Example (QbyE) KWS system that employs spectral-temporal attentive graph pooling and multi-task learning.
This framework aims to effectively learn speaker-invariant and linguistic-informative embeddings for QbyE KWS tasks.
- Score: 11.182456667123835
- License: http://creativecommons.org/licenses/by-nc-nd/4.0/
- Abstract: Existing keyword spotting (KWS) systems primarily rely on predefined keyword phrases. However, the ability to recognize customized keywords is crucial for tailoring interactions with intelligent devices. In this paper, we present a novel Query-by-Example (QbyE) KWS system that employs spectral-temporal graph attentive pooling and multi-task learning. This framework aims to effectively learn speaker-invariant and linguistic-informative embeddings for QbyE KWS tasks. Within this framework, we investigate three distinct network architectures for encoder modeling: LiCoNet, Conformer and ECAPA_TDNN. The experimental results on a substantial internal dataset of $629$ speakers have demonstrated the effectiveness of the proposed QbyE framework in maximizing the potential of simpler models such as LiCoNet. Particularly, LiCoNet, which is 13x more efficient, achieves comparable performance to the computationally intensive Conformer model (1.98% vs. 1.63\% FRR at 0.3 FAs/Hr).
Related papers
- RaCo: Ranking and Covariance for Practical Learned Keypoints [51.38393049306958]
RaCo is designed to learn robust and versatile keypoints suitable for a variety of 3D computer vision tasks.<n>RaCo operates without the need for covisible image pairs.
arXiv Detail & Related papers (2026-02-17T17:39:52Z) - Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification [0.0]
This research evaluates the feasibility of a modular architecture inspired by core functionalities of perceptive systems like 'Smart Eye'<n>We propose and benchmark three independent sensing modules: a Convolutional Neural Network (CNN) for eye state detection (drowsiness/attention), a deep CNN for facial expression recognition, and a Long Short-Term Memory (LSTM) network for voice-based speaker identification.
arXiv Detail & Related papers (2025-11-25T16:35:42Z) - EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models [64.18350535770357]
We propose an automatic pruning method for large vision-language models to enhance the efficiency of multimodal reasoning.
Our approach only leverages a small number of samples to search for the desired pruning policy.
We conduct extensive experiments on the ScienceQA, Vizwiz, MM-vet, and LLaVA-Bench datasets for the task of visual question answering.
arXiv Detail & Related papers (2025-03-19T16:07:04Z) - A Point-Based Approach to Efficient LiDAR Multi-Task Perception [49.91741677556553]
PAttFormer is an efficient multi-task architecture for joint semantic segmentation and object detection in point clouds.
Unlike other LiDAR-based multi-task architectures, our proposed PAttFormer does not require separate feature encoders for task-specific point cloud representations.
Our evaluations show substantial gains from multi-task learning, improving LiDAR semantic segmentation by +1.7% in mIou and 3D object detection by +1.7% in mAP.
arXiv Detail & Related papers (2024-04-19T11:24:34Z) - Unveiling Backbone Effects in CLIP: Exploring Representational Synergies
and Variances [49.631908848868505]
Contrastive Language-Image Pretraining (CLIP) stands out as a prominent method for image representation learning.
We investigate the differences in CLIP performance among various neural architectures.
We propose a simple, yet effective approach to combine predictions from multiple backbones, leading to a notable performance boost of up to 6.34%.
arXiv Detail & Related papers (2023-12-22T03:01:41Z) - Towards Realistic Zero-Shot Classification via Self Structural Semantic
Alignment [53.2701026843921]
Large-scale pre-trained Vision Language Models (VLMs) have proven effective for zero-shot classification.
In this paper, we aim at a more challenging setting, Realistic Zero-Shot Classification, which assumes no annotation but instead a broad vocabulary.
We propose the Self Structural Semantic Alignment (S3A) framework, which extracts structural semantic information from unlabeled data while simultaneously self-learning.
arXiv Detail & Related papers (2023-08-24T17:56:46Z) - On-Device Constrained Self-Supervised Speech Representation Learning for
Keyword Spotting via Knowledge Distillation [13.08005728839078]
We propose a knowledge distillation-based self-supervised speech representation learning architecture for on-device keyword spotting.
Our approach used a teacher-student framework to transfer knowledge from a larger, more complex model to a smaller, light-weight model.
We evaluated our model's performance on an Alexa keyword spotting detection task using a 16.6k-hour in-house dataset.
arXiv Detail & Related papers (2023-07-06T02:03:31Z) - Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural
Networks [49.808194368781095]
We show that three-layer neural networks have provably richer feature learning capabilities than two-layer networks.
This work makes progress towards understanding the provable benefit of three-layer neural networks over two-layer networks in the feature learning regime.
arXiv Detail & Related papers (2023-05-11T17:19:30Z) - On the Efficiency of Integrating Self-supervised Learning and
Meta-learning for User-defined Few-shot Keyword Spotting [51.41426141283203]
User-defined keyword spotting is a task to detect new spoken terms defined by users.
Previous works try to incorporate self-supervised learning models or apply meta-learning algorithms.
Our result shows that HuBERT combined with Matching network achieves the best result.
arXiv Detail & Related papers (2022-04-01T10:59:39Z) - Learning Decoupling Features Through Orthogonality Regularization [55.79910376189138]
Keywords spotting (KWS) and speaker verification (SV) are two important tasks in speech applications.
We develop a two-branch deep network (KWS branch and SV branch) with the same network structure.
A novel decoupling feature learning method is proposed to push up the performance of KWS and SV simultaneously.
arXiv Detail & Related papers (2022-03-31T03:18:13Z) - EfficientTDNN: Efficient Architecture Search for Speaker Recognition in
the Wild [29.59228560095565]
We propose a neural architecture search-based efficient time-delay neural network (EfficientTDNN) to improve inference efficiency while maintaining recognition accuracy.
Experiments on the VoxCeleb dataset show EfficientTDNN provides a huge search space including approximately $1013$s and achieves 1.66% EER and 0.156 DCF$_0.01$ with 565M MACs.
arXiv Detail & Related papers (2021-03-25T03:28:07Z) - Query-by-Example Keyword Spotting system using Multi-head Attention and
Softtriple Loss [1.179778723980276]
This paper proposes a neural network architecture for tackling the query-by-example user-defined keyword spotting task.
A multi-head attention module is added on top of a multi-layered GRU for effective feature extraction.
We also adopt the softtriple loss - a combination of triplet loss and softmax loss - and showcase its effectiveness.
arXiv Detail & Related papers (2021-02-14T03:37:37Z) - EfficientNet-Absolute Zero for Continuous Speech Keyword Spotting [7.313613282363873]
Football keyword dataset (FKD) is a new keyword spotting dataset in Persian.
This dataset contains nearly 31000 samples in 18 classes.
It is realized that EfficientNet-A0 and Resnet models outperform other models on this dataset.
arXiv Detail & Related papers (2020-12-31T16:21:27Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.