Related papers: PyTDC: A multimodal machine learning training, evaluation, and inference platform for biomedical foundation models

PyTDC: A multimodal machine learning training, evaluation, and inference platform for biomedical foundation models

URL: http://arxiv.org/abs/2505.05577v1
Date: Thu, 08 May 2025 18:15:38 GMT
Title: PyTDC: A multimodal machine learning training, evaluation, and inference platform for biomedical foundation models
Authors: Alejandro Velez-Arce, Marinka Zitnik,
Abstract summary: PyTDC is a machine-learning platform providing streamlined training, evaluation, and inference software for multimodal biological AI models.<n>This paper discusses the components of PyTDC's architecture and, to our knowledge, the first-of-its-kind case study on the introduced single-cell drug-target nomination ML task.
Score: 59.17570021208177
License: http://creativecommons.org/licenses/by/4.0/
Abstract: Existing biomedical benchmarks do not provide end-to-end infrastructure for training, evaluation, and inference of models that integrate multimodal biological data and a broad range of machine learning tasks in therapeutics. We present PyTDC, an open-source machine-learning platform providing streamlined training, evaluation, and inference software for multimodal biological AI models. PyTDC unifies distributed, heterogeneous, continuously updated data sources and model weights and standardizes benchmarking and inference endpoints. This paper discusses the components of PyTDC's architecture and, to our knowledge, the first-of-its-kind case study on the introduced single-cell drug-target nomination ML task. We find state-of-the-art methods in graph representation learning and domain-specific methods from graph theory perform poorly on this task. Though we find a context-aware geometric deep learning method that outperforms the evaluated SoTA and domain-specific baseline methods, the model is unable to generalize to unseen cell types or incorporate additional modalities, highlighting PyTDC's capacity to facilitate an exciting avenue of research developing multimodal, context-aware, foundation models for open problems in biomedical AI.

Related papers

Benchmarking Foundation Models with Multimodal Public Electronic Health Records [24.527782376051693]
We present a benchmark that evaluates the performance, fairness, and interpretability of foundation models.<n>We developed a standardized data processing pipeline that harmonizes heterogeneous clinical records into an analysis-ready format.<n>Our findings demonstrate that incorporating multiple data modalities leads to consistent improvements in predictive performance without introducing additional bias.
arXiv Detail & Related papers (2025-07-20T05:08:28Z)
Platform for Representation and Integration of multimodal Molecular Embeddings [43.54912893426355]
Existing machine learning methods for molecular embeddings are restricted to specific tasks or data modalities.<n>Existing embeddings capture largely non-overlapping molecular signals, highlighting the value of embedding integration.<n>We propose Platform for Representation and Integration of multimodal Molecular Embeddings (PRISME) to integrate heterogeneous embeddings into a unified multimodal representation.
arXiv Detail & Related papers (2025-07-10T01:18:50Z)
Biomedical Foundation Model: A Survey [84.26268124754792]
Foundation models are large-scale pre-trained models that learn from extensive unlabeled datasets.<n>These models can be adapted to various applications such as question answering and visual understanding.<n>This survey explores the potential of foundation models across diverse domains within biomedical fields.
arXiv Detail & Related papers (2025-03-03T22:42:00Z)
Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation [113.5002649181103]
Training open-source small multimodal models (SMMs) to bridge competency gaps for unmet clinical needs in radiology. For training, we assemble a large dataset of over 697 thousand radiology image-text pairs. For evaluation, we propose CheXprompt, a GPT-4-based metric for factuality evaluation, and demonstrate its parity with expert evaluation. The inference of LlaVA-Rad is fast and can be performed on a single V100 GPU in private settings, offering a promising state-of-the-art tool for real-world clinical applications.
arXiv Detail & Related papers (2024-03-12T18:12:02Z)
OpenMEDLab: An Open-source Platform for Multi-modality Foundation Models in Medicine [55.29668193415034]
We present OpenMEDLab, an open-source platform for multi-modality foundation models. It encapsulates solutions of pioneering attempts in prompting and fine-tuning large language and vision models for frontline clinical and bioinformatic applications. It opens access to a group of pre-trained foundation models for various medical image modalities, clinical text, protein engineering, etc.
arXiv Detail & Related papers (2024-02-28T03:51:02Z)
Advancing bioinformatics with large language models: components, applications and perspectives [12.728981464533918]
Large language models (LLMs) are a class of artificial intelligence models based on deep learning.<n>We will provide a comprehensive overview of the essential components of large language models (LLMs) in bioinformatics.<n>Key aspects covered include tokenization methods for diverse data types, the architecture of transformer models, and the core attention mechanism.
arXiv Detail & Related papers (2024-01-08T17:26:59Z)
HEALNet: Multimodal Fusion for Heterogeneous Biomedical Data [10.774128925670183]
This paper presents the Hybrid Early-fusion Attention Learning Network (HEALNet), a flexible multimodal fusion architecture. We conduct multimodal survival analysis on Whole Slide Images and Multi-omic data on four cancer datasets from The Cancer Genome Atlas (TCGA) HEALNet achieves state-of-the-art performance compared to other end-to-end trained fusion models.
arXiv Detail & Related papers (2023-11-15T17:06:26Z)
Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects [2.1070612998322438]
The paper explores the transformative potential of multimodal models for clinical predictions. Despite advancements, challenges such as data biases and the scarcity of "big data" in many biomedical domains persist.
arXiv Detail & Related papers (2023-11-04T05:42:51Z)
DIME: Fine-grained Interpretations of Multimodal Models via Disentangled Local Explanations [119.1953397679783]
We focus on advancing the state-of-the-art in interpreting multimodal models. Our proposed approach, DIME, enables accurate and fine-grained analysis of multimodal models.
arXiv Detail & Related papers (2022-03-03T20:52:47Z)
An Extensible Benchmark Suite for Learning to Simulate Physical Systems [60.249111272844374]
We introduce a set of benchmark problems to take a step towards unified benchmarks and evaluation protocols. We propose four representative physical systems, as well as a collection of both widely used classical time-based and representative data-driven methods.
arXiv Detail & Related papers (2021-08-09T17:39:09Z)

This list is automatically generated from the titles and abstracts of the papers in this site.