Towards Synthetic Multivariate Time Series Generation for Flare
Forecasting
- URL: http://arxiv.org/abs/2105.07532v1
- Date: Sun, 16 May 2021 22:23:23 GMT
- Title: Towards Synthetic Multivariate Time Series Generation for Flare
Forecasting
- Authors: Yang Chen, Dustin J. Kempton, Azim Ahmadzadeh and Rafal A. Angryk
- Abstract summary: One of the limiting factors in training data-driven, rare-event prediction algorithms is the scarcity of the events of interest.
In this study, we explore the usefulness of the conditional generative adversarial network (CGAN) as a means to perform data-informed oversampling.
- Score: 5.098461305284216
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: One of the limiting factors in training data-driven, rare-event prediction
algorithms is the scarcity of the events of interest resulting in an extreme
imbalance in the data. There have been many methods introduced in the
literature for overcoming this issue; simple data manipulation through
undersampling and oversampling, utilizing cost-sensitive learning algorithms,
or by generating synthetic data points following the distribution of the
existing data. While synthetic data generation has recently received a great
deal of attention, there are real challenges involved in doing so for
high-dimensional data such as multivariate time series. In this study, we
explore the usefulness of the conditional generative adversarial network (CGAN)
as a means to perform data-informed oversampling in order to balance a large
dataset of multivariate time series. We utilize a flare forecasting benchmark
dataset, named SWAN-SF, and design two verification methods to both
quantitatively and qualitatively evaluate the similarity between the generated
minority and the ground-truth samples. We further assess the quality of the
generated samples by training a classical, supervised machine learning
algorithm on synthetic data, and testing the trained model on the unseen, real
data. The results show that the classifier trained on the data augmented with
the synthetic multivariate time series achieves a significant improvement
compared with the case where no augmentation is used. The popular flare
forecasting evaluation metrics, TSS and HSS, report 20-fold and 5-fold
improvements, respectively, indicating the remarkable statistical similarities,
and the usefulness of CGAN-based data generation for complicated tasks such as
flare forecasting.
Related papers
- Tackling Data Heterogeneity in Federated Time Series Forecasting [61.021413959988216]
Time series forecasting plays a critical role in various real-world applications, including energy consumption prediction, disease transmission monitoring, and weather forecasting.
Most existing methods rely on a centralized training paradigm, where large amounts of data are collected from distributed devices to a central cloud server.
We propose a novel framework, Fed-TREND, to address data heterogeneity by generating informative synthetic data as auxiliary knowledge carriers.
arXiv Detail & Related papers (2024-11-24T04:56:45Z) - Generating Realistic Tabular Data with Large Language Models [49.03536886067729]
Large language models (LLM) have been used for diverse tasks, but do not capture the correct correlation between the features and the target variable.
We propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data.
Our experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks.
arXiv Detail & Related papers (2024-10-29T04:14:32Z) - Synthetic Oversampling: Theory and A Practical Approach Using LLMs to Address Data Imbalance [16.047084318753377]
Imbalanced data and spurious correlations are common challenges in machine learning and data science.
Oversampling, which artificially increases the number of instances in the underrepresented classes, has been widely adopted to tackle these challenges.
We introduce OPAL, a systematic oversampling approach that leverages the capabilities of large language models to generate high-quality synthetic data for minority groups.
arXiv Detail & Related papers (2024-06-05T21:24:26Z) - Class-Based Time Series Data Augmentation to Mitigate Extreme Class Imbalance for Solar Flare Prediction [1.4272411349249625]
Time series data plays a crucial role across various domains, making it valuable for decision-making and predictive modeling.
Machine learning (ML) and deep learning (DL) have shown promise in this regard, yet their performance hinges on data quality and quantity.
Data augmentation techniques offer a potential solution to address these challenges, yet their effectiveness on multivariate time series datasets remains underexplored.
arXiv Detail & Related papers (2024-05-31T03:03:19Z) - Towards Theoretical Understandings of Self-Consuming Generative Models [56.84592466204185]
This paper tackles the emerging challenge of training generative models within a self-consuming loop.
We construct a theoretical framework to rigorously evaluate how this training procedure impacts the data distributions learned by future models.
We present results for kernel density estimation, delivering nuanced insights such as the impact of mixed data training on error propagation.
arXiv Detail & Related papers (2024-02-19T02:08:09Z) - Synthetic data, real errors: how (not) to publish and use synthetic data [86.65594304109567]
We show how the generative process affects the downstream ML task.
We introduce Deep Generative Ensemble (DGE) to approximate the posterior distribution over the generative process model parameters.
arXiv Detail & Related papers (2023-05-16T07:30:29Z) - TimeVAE: A Variational Auto-Encoder for Multivariate Time Series
Generation [6.824692201913679]
We propose a novel architecture for synthetically generating time-series data with the use of Variversaational Auto-Encoders (VAEs)
The proposed architecture has several distinct properties: interpretability, ability to encode domain knowledge, and reduced training times.
arXiv Detail & Related papers (2021-11-15T21:42:14Z) - Convolutional generative adversarial imputation networks for
spatio-temporal missing data in storm surge simulations [86.5302150777089]
Generative Adversarial Imputation Nets (GANs) and GAN-based techniques have attracted attention as unsupervised machine learning methods.
We name our proposed method as Con Conval Generative Adversarial Imputation Nets (Conv-GAIN)
arXiv Detail & Related papers (2021-11-03T03:50:48Z) - Global Models for Time Series Forecasting: A Simulation Study [2.580765958706854]
We simulate time series from simple data generating processes (DGP), such as Auto Regressive (AR) and Seasonal AR, to complex DGPs, such as Chaotic Logistic Map, Self-Exciting Threshold Auto-Regressive, and Mackey-Glass equations.
The lengths and the number of series in the dataset are varied in different scenarios.
We perform experiments on these datasets using global forecasting models including Recurrent Neural Networks (RNN), Feed-Forward Neural Networks, Pooled Regression (PR) models, and Light Gradient Boosting Models (LGBM)
arXiv Detail & Related papers (2020-12-23T04:45:52Z) - Learning summary features of time series for likelihood free inference [93.08098361687722]
We present a data-driven strategy for automatically learning summary features from time series data.
Our results indicate that learning summary features from data can compete and even outperform LFI methods based on hand-crafted values.
arXiv Detail & Related papers (2020-12-04T19:21:37Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.