Skip to content
This repository was archived by the owner on Aug 5, 2025. It is now read-only.

Latest commit

 

History

142 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SPARCats

A o2S2PARC tool to accelerate data discovery by AI and bring breakthroughs in the laboratory into healthcare. SPARCats generates synthetic augmented physiological data available through the SPARC Data Portal to improve AI training.

Logo

Python 3 Contributors Stargazers License Contributor Covenant PyPI version fury.io Conventional Commits

Table of contents

About

This is the repository of team SPARCats (Team #F) of the 2025 SPARC Codeathon. Information about the 2025 SPARC Codeathon can be found here.

No work was done on this project prior to the Codeathon.

Introduction

The NIH Common fund program Stimulating Peripheral Activity to Relieve Conditions (SPARC) seeks to understand how electrical signals control internal organ function. In doing so it explores how therapeutic devices might modulate nerve activity to treat conditions like hypertension, heart failure, and gastrointestinal disorders. To this end, data have been compiled from 60+ research groups, involving 3900+ subjects across 8 species from >60 anatomies, orgnans and structures.

The SPARC Portal offers a user-friendly interface to access and share resources from the SPARC community. It features well-curated, high-impact data, SPARC projects, and computational simulations, all available under the “Find Data” section.

The problem

AI can produce powerful tools for predicting and analysing a wide range of physiological biomarkers. Researchers want to use physiological time series data train AI models to answer scientific questions and distill impactful medical insights. They need a tool that allows them to seamlessly discover and generate augmented data from publicly available physiological data.

Limited Accessibility:

  • SPARC boasts a trove of datasets and tools, but some of these can be difficult to access and deploy for users that are not proficient in coding.

Poor Findability

  • There is no tool to easily search and collect relevant datafiles from the data sets available within o2S2PARC, leaving data underutilised.

Difficulties in Reusability:

  • Without a standardized approach to data augmentation, researchers may struggle to replicate or pipelines for different datasets or research contexts.

Our solution - (SPARCats)

We have developed a robust SPARC augmented timeseries (SPARCats) toolkit that runs within o2S2PARC to generate synthetic datasets. This toolkit empowers researchers to easily create pipelines that incorporate novel or published data as augmented training data for AI models. SPARCats includes a number of cookie-cutter services:

  • Search - The search feature enables the user to search for related datasets by filtering on experiment type, species, sex, and organ. With a little further tweaking, this module can become a become a general purpose sparc portal search service embedded within o2S2PARC.
  • Augment - The augment feature enables the user to perform a configurable data augmentation through the time warping, adding noise and applying a drift. The resulting synthetic dataset can be saved or piped directly into training an AI model.
  • Split data - The split data feature allows for the user to perform a configurable test/train data split for training.
  • Download - The download service allows for the user to download their data.

An example workflow is available here: SPARCats Example Workflow

SPARCats includes support for more complex steps in the AI development process through modular example code:

  • Train - Showcases the training process, enables the user to pass data (usually augmented data) into a specific model for training. The resulting model can be saved, or used directly to make predictions on real data.
  • Predict - Showcases the prediction process, enables the user to pass real data into a trained AI model to generate predictions and outputs to evaluate.

SPARCats has been developed to adhere to and enhance the FAIR core to SPARC.

  • Findable - The SPARCats search tool directly enhances the findability of data contained within SPARC.
  • Accessible - The low-code nature of the SPARCats cookie-cutter modules and comprehensive tutorials make searching for data, augmentation of data, and AI training accessible to users with a wide range of backgrounds and skills. From wetlab scientists, to curious highschool students.
  • Interoperable - SPARCats incorporates many existing SPARC tools and can be incorporated into a wide range of workflows.
  • Reusable - The configurability of the SPARCats cookie-cutter modules enables repeatable, sharable pipelines while encoraging the use of published data.

Impact

SPARCats can be used to enhance understanding of the autonomic and peripheral nervous systems and accelerate the development of healthcare innovation.

Develops new capabilities of SPARC tools within o2S2PARC

The tools offered by SPARCats extends the capabilities of o2S2PARC to streamline the discovery of data and development of AI tools. This can be done in a no/low-code way, that is accessible to a wide audience, beyond just research scientists.

Increase visibility and value of SPARC's public data

The Search and Augment modules within SPARCats enables users to easily find relevant data to include in their AI project. The augmentation tools offered enable published data to be readily used as training data for AI models, increasing the value of existing datasets.

Using sparcats

Sparcats has been developed to be used within o2S2PARC with little to node coding experience required. To demonstrate this and showcase the steps required to implement we have a series of tutorials and demo usecases:

Tutorial Description
Tutorial 1: Getting started - In this tutorial we show the basic set up process to create your first workflow in o2S2PARC using SPARCats and demonstrate how data augmentation improves performance & generalization of the resulting AI.
Tutorial 2: Augmenting SPARC data - In this tutorial we use the SPARCats search service to collect Search existing data, augment it and, train an AI model.
Tutorial 3: Train existing model - In this tutorial we use an existing model
Tutorial 4: Transfer learning - In this tutorial we use transfer learning by searching for existing data, augment it, and use it to fine-tune an existing AI model.


Reporting issues

To report an issue or suggest a new feature, please use the issues page. Please check existing issues before submitting a new one.

Contributing

To contribute: fork this repository and submit a pull request. Before submitting a pull request, please read our Contributing Guidelines and Code of Conduct. If you found this tool helpful, please add a GitHub Star to support further developments!

Project structure

  • /src/ - Directory of sparcats python module.
  • /tutorials/ - Directory of tutorials and use cases showcasing sparcats python module in use.

License

Sparcats is open source and distributed under the Apache License 2.0. See LICENSE for more information.

Team

Acknowledgements

  • We would like to thank the 2025 SPARC Codeathon organizers for their guidance and support during this Codeathon.

About

No description, website, or topics provided.

Resources

Code of conduct

Contributing

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages