BUILDA: A Thermal Building Data Generation Framework for Transfer Learning
- URL: http://arxiv.org/abs/2508.12703v1
- Date: Mon, 18 Aug 2025 08:01:37 GMT
- Title: BUILDA: A Thermal Building Data Generation Framework for Transfer Learning
- Authors: Thomas Krug, Fabian Raisch, Dominik Aimer, Markus Wirnsberger, Ferdinand Sigg, Benjamin Schäfer, Benjamin Tischler,
- Abstract summary: Transfer learning can improve data-driven modeling of building thermal dynamics.<n>We present BuilDa, a framework for producing synthetic data of adequate quality and quantity for TL research.
- Score: 26.47874938214435
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Transfer learning (TL) can improve data-driven modeling of building thermal dynamics. Therefore, many new TL research areas emerge in the field, such as selecting the right source model for TL. However, these research directions require massive amounts of thermal building data which is lacking presently. Neither public datasets nor existing data generators meet the needs of TL research in terms of data quality and quantity. Moreover, existing data generation approaches typically require expert knowledge in building simulation. We present BuilDa, a thermal building data generation framework for producing synthetic data of adequate quality and quantity for TL research. The framework does not require profound building simulation knowledge to generate large volumes of data. BuilDa uses a single-zone Modelica model that is exported as a Functional Mock-up Unit (FMU) and simulated in Python. We demonstrate BuilDa by generating data and utilizing it for pretraining and fine-tuning TL models.
Related papers
- A Highly Configurable Framework for Large-Scale Thermal Building Data Generation to drive Machine Learning Research [22.54521342959957]
BuilDa is designed to produce synthetic data of adequate quality and quantity for machine learning (ML) research.<n>It does not require profound building simulation knowledge to generate large volumes of data.<n>We demonstrate BuilDa by generating data and utilizing it for a transfer learning study involving the fine-tuning of 486 data-driven models.
arXiv Detail & Related papers (2025-11-29T13:31:02Z) - DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback [62.235925602004535]
DataEnvGym is a testbed of teacher environments for data generation agents.<n>It frames data generation as a sequential decision-making task, involving an agent and a data generation engine.<n>Students are iteratively trained and evaluated on generated data, and their feedback is reported to the agent after each iteration.
arXiv Detail & Related papers (2024-10-08T17:20:37Z) - A Benchmark Time Series Dataset for Semiconductor Fabrication Manufacturing Constructed using Component-based Discrete-Event Simulation Models [0.0]
This research is based on a benchmark model of an Intel semiconductor fabrication factory.
The time series dataset is constructed using discrete-event time trajectories.
The dataset can also be utilized in the machine learning community for behavioral analysis.
arXiv Detail & Related papers (2024-08-17T23:05:47Z) - Scaling Data-Driven Building Energy Modelling using Large Language Models [3.0309252269809264]
We propose a methodology to tackle the scalability challenges associated with the development of data-driven models for Building Management System.
We use Large Language Models (LLMs) to generate code that processes structured data from BMS and build data-driven models for BMS's specific requirements.
Our case study indicates that bi-sequential prompting under the prompt template can achieve a high success rate of code generation and code accuracy, and significantly reduce human labor costs.
arXiv Detail & Related papers (2024-07-03T19:34:24Z) - Heat Death of Generative Models in Closed-Loop Learning [63.83608300361159]
We study the learning dynamics of generative models that are fed back their own produced content in addition to their original training dataset.
We show that, unless a sufficient amount of external data is introduced at each iteration, any non-trivial temperature leads the model to degenerate.
arXiv Detail & Related papers (2024-04-02T21:51:39Z) - Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data [49.73114504515852]
We show that replacing the original real data by each generation's synthetic data does indeed tend towards model collapse.
We demonstrate that accumulating the successive generations of synthetic data alongside the original real data avoids model collapse.
arXiv Detail & Related papers (2024-04-01T18:31:24Z) - Spatio-Temporal Graph Convolutional Network Combined Large Language Model: A Deep Learning Framework for Bike Demand Forecasting [0.0]
This study presents a new deep learning framework, combining Spatio-Temporal Graph Convolutional Network (STGCN) with a Large Language Model (LLM)
The proposed STGCN-L model demonstrates competitive performance compared to existing models, showcasing its potential in predicting bike demand.
arXiv Detail & Related papers (2024-03-23T05:47:19Z) - Scalable Diffusion for Materials Generation [99.71001883652211]
We develop a unified crystal representation that can represent any crystal structure (UniMat)
UniMat can generate high fidelity crystal structures from larger and more complex chemical systems.
We propose additional metrics for evaluating generative models of materials.
arXiv Detail & Related papers (2023-10-18T15:49:39Z) - T1: Scaling Diffusion Probabilistic Fields to High-Resolution on Unified
Visual Modalities [69.16656086708291]
Diffusion Probabilistic Field (DPF) models the distribution of continuous functions defined over metric spaces.
We propose a new model comprising of a view-wise sampling algorithm to focus on local structure learning.
The model can be scaled to generate high-resolution data while unifying multiple modalities.
arXiv Detail & Related papers (2023-05-24T03:32:03Z) - Comparison of Transfer Learning based Additive Manufacturing Models via
A Case Study [3.759936323189418]
This paper defines a case study based on an open-source dataset about metal AM products.
Five TL methods are integrated with decision tree regression (DTR) and artificial neural network (ANN) to construct six TL-based models.
The comparisons are used to quantify the performance of applied TL methods and are discussed from the perspective of similarity, training data size, and data preprocessing.
arXiv Detail & Related papers (2023-05-17T00:29:25Z) - Mapping and Describing Geospatial Data to Generalize Complex Mapping and
Describing Geospatial Data to Generalize Complex Models: The Case of
LittoSIM-GEN Models [0.0]
We provide a mapping approach to structure, describe, and automatize the integration of geospatial data into agent-based models.
This paper was part of the LittoSIM-GEN project.
arXiv Detail & Related papers (2021-01-19T09:16:05Z) - Learning Discrete Energy-based Models via Auxiliary-variable Local
Exploration [130.89746032163106]
We propose ALOE, a new algorithm for learning conditional and unconditional EBMs for discrete structured data.
We show that the energy function and sampler can be trained efficiently via a new variational form of power iteration.
We present an energy model guided fuzzer for software testing that achieves comparable performance to well engineered fuzzing engines like libfuzzer.
arXiv Detail & Related papers (2020-11-10T19:31:29Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.