A o2S2PARC tool to accelerate data discovery by AI and bring breakthroughs in the laboratory into healthcare. SPARCats generates synthetic augmented physiological data available through the SPARC Data Portal to improve AI training.
- About
- Introduction
- The problem
- Our solution - (SPARCats)
- Impact
- Contributing
- Setting up sparcats
- Reporting issues
- Contributing
- FAIR practices
- License
- Team
- Acknowledgements
This is the repository of team SPARCats (Team #F) of the 2025 SPARC Codeathon. Information about the 2025 SPARC Codeathon can be found here.
No work was done on this project prior to the Codeathon.
The NIH Common fund program Stimulating Peripheral Activity to Relieve Conditions (SPARC) seeks to understand how electrical signals control internal organ function. In doing so it explores how therapeutic devices might modulate nerve activity to treat conditions like hypertension, heart failure, and gastrointestinal disorders. To this end, data have been compiled from 60+ research groups, involving 3900+ subjects across 8 species from >60 anatomies, orgnans and structures.
The SPARC Portal offers a user-friendly interface to access and share resources from the SPARC community. It features well-curated, high-impact data, SPARC projects, and computational simulations, all available under the “Find Data” section.
AI can produce powerful tools for predicting and analysing a wide range of physiological biomarkers. Researchers want to use physiological time series data train AI models to answer scientific questions and distill impactful medical insights. They need a tool that allows them to seamlessly discover and generate augmented data from publicly available physiological data.
- SPARC boasts a trove of datasets and tools, but some of these can be difficult to access and deploy for users that are not proficient in coding.
- There is no tool to easily search and collect relevant datafiles from the data sets available within o2S2PARC, leaving data underutilised.
- Without a standardized approach to data augmentation, researchers may struggle to replicate or pipelines for different datasets or research contexts.
We have developed a robust SPARC augmented timeseries (SPARCats) toolkit that runs within o2S2PARC to generate synthetic datasets. This toolkit empowers researchers to easily create pipelines that incorporate novel or published data as augmented training data for AI models. SPARCats includes a number of cookie-cutter services:
- Search - The search feature enables the user to search for related datasets by filtering on experiment type, species, sex, and organ. With a little further tweaking, this module can become a become a general purpose sparc portal search service embedded within o2S2PARC.
- Augment - The augment feature enables the user to perform a configurable data augmentation through the time warping, adding noise and applying a drift. The resulting synthetic dataset can be saved or piped directly into training an AI model.
- Split data - The split data feature allows for the user to perform a configurable test/train data split for training.
- Download - The download service allows for the user to download their data.
An example workflow is available here: SPARCats Example Workflow
SPARCats includes support for more complex steps in the AI development process through modular example code:
- Train - Showcases the training process, enables the user to pass data (usually augmented data) into a specific model for training. The resulting model can be saved, or used directly to make predictions on real data.
- Predict - Showcases the prediction process, enables the user to pass real data into a trained AI model to generate predictions and outputs to evaluate.
SPARCats has been developed to adhere to and enhance the FAIR core to SPARC.
- Findable - The SPARCats search tool directly enhances the findability of data contained within SPARC.
- Accessible - The low-code nature of the SPARCats cookie-cutter modules and comprehensive tutorials make searching for data, augmentation of data, and AI training accessible to users with a wide range of backgrounds and skills. From wetlab scientists, to curious highschool students.
- Interoperable - SPARCats incorporates many existing SPARC tools and can be incorporated into a wide range of workflows.
- Reusable - The configurability of the SPARCats cookie-cutter modules enables repeatable, sharable pipelines while encoraging the use of published data.
SPARCats can be used to enhance understanding of the autonomic and peripheral nervous systems and accelerate the development of healthcare innovation.
The tools offered by SPARCats extends the capabilities of o2S2PARC to streamline the discovery of data and development of AI tools. This can be done in a no/low-code way, that is accessible to a wide audience, beyond just research scientists.
The Search and Augment modules within SPARCats enables users to easily find relevant data to include in their AI project. The augmentation tools offered enable published data to be readily used as training data for AI models, increasing the value of existing datasets.
Sparcats has been developed to be used within o2S2PARC with little to node coding experience required. To demonstrate this and showcase the steps required to implement we have a series of tutorials and demo usecases:
| Tutorial | Description |
|---|---|
| Tutorial 1: | Getting started - In this tutorial we show the basic set up process to create your first workflow in o2S2PARC using SPARCats and demonstrate how data augmentation improves performance & generalization of the resulting AI. |
| Tutorial 2: | Augmenting SPARC data - In this tutorial we use the SPARCats search service to collect Search existing data, augment it and, train an AI model. |
| Tutorial 3: | Train existing model - In this tutorial we use an existing model |
| Tutorial 4: | Transfer learning - In this tutorial we use transfer learning by searching for existing data, augment it, and use it to fine-tune an existing AI model. |
To report an issue or suggest a new feature, please use the issues page. Please check existing issues before submitting a new one.
To contribute: fork this repository and submit a pull request. Before submitting a pull request, please read our Contributing Guidelines and Code of Conduct. If you found this tool helpful, please add a GitHub Star to support further developments!
/src/- Directory of sparcats python module./tutorials/- Directory of tutorials and use cases showcasing sparcats python module in use.
Sparcats is open source and distributed under the Apache License 2.0. See LICENSE for more information.
- Michael Hoffman (Writer)
- Mathias Roesler (Developer)
- Mishaim Malik (Developer)
- Omkar Athavale (Lead, SysAdmin)
- We would like to thank the 2025 SPARC Codeathon organizers for their guidance and support during this Codeathon.
