Related papers: 3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces

3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces

URL: http://arxiv.org/abs/2104.11532v1
Date: Fri, 23 Apr 2021 10:56:34 GMT
Title: 3D Convolutional Neural Networks for Ultrasound-Based Silent Speech Interfaces
Authors: L\'aszl\'o T\'oth, Amin Honarmandi Shandiz
Abstract summary: Silent speech interfaces (SSI) aim to reconstruct the speech signal from a recording of the articulatory movement, such as an ultrasound video of the tongue. Deep neural networks are the most successful technology for this task. One option for this is to apply recurrent neural structures such as the long short-term memory network (LSTM) in combination with 2D convolutional neural networks (CNNs)
Score: 0.0
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Abstract: Silent speech interfaces (SSI) aim to reconstruct the speech signal from a recording of the articulatory movement, such as an ultrasound video of the tongue. Currently, deep neural networks are the most successful technology for this task. The efficient solution requires methods that do not simply process single images, but are able to extract the tongue movement information from a sequence of video frames. One option for this is to apply recurrent neural structures such as the long short-term memory network (LSTM) in combination with 2D convolutional neural networks (CNNs). Here, we experiment with another approach that extends the CNN to perform 3D convolution, where the extra dimension corresponds to time. In particular, we apply the spatial and temporal convolutions in a decomposed form, which proved very successful recently in video action recognition. We find experimentally that our 3D network outperforms the CNN+LSTM model, indicating that 3D CNNs may be a feasible alternative to CNN+LSTM networks in SSI systems.

Related papers

3DPyranet Features Fusion for Spatio-temporal Feature Learning [2.327279581393927]
3D pyramidal neural pyramid called 3DPyraNet and a discriminative approach for classifier-temporal feature learning called 3DPyraNet-F are proposed. 3DPyraNet-F extract the features maps of the highest layer of the learned network, fuse them in a single vector, and provide it as input in a way to a linear-SVM. Results are reported with 3DPyraNet in real-world environments, especially in the presence of camera induced motion.
arXiv Detail & Related papers (2025-04-26T17:32:37Z)
Efficient 3D Recognition with Event-driven Spike Sparse Convolution [15.20476631850388]
Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D-temporal features. We introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space. We propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features.
arXiv Detail & Related papers (2024-12-10T09:55:15Z)
Maximizing Spatio-Temporal Entropy of Deep 3D CNNs for Efficient Video Recognition [25.364148451584356]
3D convolution neural networks (CNNs) have been the prevailing option for video recognition. We propose to automatically design efficient 3D CNN architectures via a novel training-free neural architecture search approach. Experiments on Something-Something V1&V2 and Kinetics400 demonstrate that the E3D family achieves state-of-the-art performance.
arXiv Detail & Related papers (2023-03-05T15:11:53Z)
Continual 3D Convolutional Neural Networks for Real-time Processing of Videos [93.73198973454944]
We introduce Continual 3D Contemporalal Neural Networks (Co3D CNNs) Co3D CNNs process videos frame-by-frame rather than by clip by clip. We show that Co3D CNNs initialised on the weights from preexisting state-of-the-art video recognition models reduce floating point operations for frame-wise computations by 10.0-12.4x while improving accuracy on Kinetics-400 by 2.3-3.8x.
arXiv Detail & Related papers (2021-05-31T18:30:52Z)
MoViNets: Mobile Video Networks for Efficient Video Recognition [52.49314494202433]
3D convolutional neural networks (CNNs) are accurate at video recognition but require large computation and memory budgets. We propose a three-step approach to improve computational efficiency while substantially reducing the peak memory usage of 3D CNNs.
arXiv Detail & Related papers (2021-03-21T23:06:38Z)
Learning Hybrid Representations for Automatic 3D Vessel Centerline Extraction [57.74609918453932]
Automatic blood vessel extraction from 3D medical images is crucial for vascular disease diagnoses. Existing methods may suffer from discontinuities of extracted vessels when segmenting such thin tubular structures from 3D images. We argue that preserving the continuity of extracted vessels requires to take into account the global geometry. We propose a hybrid representation learning approach to address this challenge.
arXiv Detail & Related papers (2020-12-14T05:22:49Z)
3D CNNs with Adaptive Temporal Feature Resolutions [83.43776851586351]
Similarity Guided Sampling (SGS) module can be plugged into any existing 3D CNN architecture. SGS empowers 3D CNNs by learning the similarity of temporal features and grouping similar features together. Our evaluations show that the proposed module improves the state-of-the-art by reducing the computational cost (GFLOPs) by half while preserving or even improving the accuracy.
arXiv Detail & Related papers (2020-11-17T14:34:05Z)
Efficient Arabic emotion recognition using deep neural networks [21.379338888447602]
We implement two neural architectures to address the problem of emotion recognition from speech signal. The first is an attention-based CNN-LSTM-DNN model; the second is a deep CNN model. The results on an Arabic speech emotion recognition task show that our innovative approach can lead to significant improvements.
arXiv Detail & Related papers (2020-10-31T19:39:37Z)
Human Activity Recognition using Multi-Head CNN followed by LSTM [1.8830374973687412]
This study presents a novel method to recognize human physical activities using CNN followed by LSTM. By using the proposed method, we achieve state-of-the-art accuracy, which is comparable to traditional machine learning algorithms and other deep neural network algorithms.
arXiv Detail & Related papers (2020-02-21T14:29:59Z)
V4D:4D Convolutional Neural Networks for Video-level Representation Learning [58.548331848942865]
Most 3D CNNs for video representation learning are clip-based, and thus do not consider video-temporal evolution of features. We propose Video-level 4D Conal Neural Networks, or V4D, to model long-range representation with 4D convolutions. V4D achieves excellent results, surpassing recent 3D CNNs by a large margin.
arXiv Detail & Related papers (2020-02-18T09:27:41Z)
An Information-rich Sampling Technique over Spatio-Temporal CNN for Classification of Human Actions in Videos [5.414308305392762]
We propose a novel scheme for human action recognition in videos, using a 3-dimensional Convolutional Neural Network (3D CNN) based classifier. In this paper, a 3D CNN architecture is proposed to extract featuresweighted and follows Long Short-Term Memory (LSTM) to recognize human actions. Experiments are performed with KTH and WEIZMANN human actions datasets, whereby it is shown to produce comparable results with the state-of-the-art techniques.
arXiv Detail & Related papers (2020-02-06T05:07:41Z)
PV-RCNN: Point-Voxel Feature Set Abstraction for 3D Object Detection [76.30585706811993]
We present a novel and high-performance 3D object detection framework, named PointVoxel-RCNN (PV-RCNN) Our proposed method deeply integrates both 3D voxel Convolutional Neural Network (CNN) and PointNet-based set abstraction. It takes advantages of efficient learning and high-quality proposals of the 3D voxel CNN and the flexible receptive fields of the PointNet-based networks.
arXiv Detail & Related papers (2019-12-31T06:34:10Z)

This list is automatically generated from the titles and abstracts of the papers in this site.