Skip to content

Latest commit

 

History

395 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

MIR Feature Extraction Framework

Comprehensive music feature extraction pipeline for conditioning Stable Audio Tools and similar audio generation models. Extracts 97+ numeric MIR features, 496 AI classification labels, and 5 natural language descriptions from audio files.

Status: Work-in-progress but functional. Core analysis scripts are tested; pipeline glue may lag behind. Scripts have built-in --help.

The pipeline — related repositories

This repository is one of three that together form a small pipeline for controllable AI music generation. It sits at the head of it: the features extracted here are the control targets everything downstream steers toward.

Repository Role
mir-feature-extraction (this repo) Extracts musical features from audio — rhythm/beat, loudness, spectral, harmonic, timbral, aesthetic — as whole-track timeseries and labels. The analysis backbone and the source of control targets. Also hosts a latent-space explorer/player for auditioning.
audio-tools-avp Trains the control layer: latent-control heads (LatCH) against those features, the FusionOpt optimizer, SAO-Small finetuning, and Stable Audio 3 control adapters. A fork of stable-audio-tools.
stable-audio-3 Runs Stable Audio 3 inference with those controls: LatCH-guided generation, LoRA finetuning, long-form rendering, and experimental samplers. A fork of stable-audio-3.

Flow: mir-feature-extraction (measure musical features) → audio-tools-avp (train heads/adapters that steer toward them) → stable-audio-3 (generate, steered). Because the control heads read the same latent space the generator carves, a feature this repo can measure becomes a knob you can steer — the throughline across all three.

What It Does

  1. Organizes audio files into structured folders
  2. Separates stems (drums, bass, other, vocals) via Demucs or BS-RoFormer
  3. Extracts rhythm, loudness, spectral, harmonic, timbral, and aesthetic features
  4. Classifies genre (400), mood (56), instruments (40) via Essentia
  5. Generates AI text descriptions via Music Flamingo (8B params), optionally condensed by Granite-tiny revision
  6. Benchmarks caption quality across Music Flamingo, LLM revision, and Qwen2.5-Omni
  7. Transcribes drums to MIDI via ADTOF-PyTorch
  8. Creates beat-aligned training crops with feature migration

All features are saved to .INFO JSON files with atomic writes (never overwrites).

Requirements

  • Python 3.12+
  • GPU: AMD ROCm 7.2+ (tested on RX 9070 XT / RDNA4) or NVIDIA CUDA
  • VRAM: 5-13 GB depending on workload (up to 10 GB for captioning benchmark)
  • OS: Linux (tested on Arch)

Key Dependencies

Package Purpose
PyTorch (ROCm/CUDA) GPU compute
Demucs / BS-RoFormer Stem separation
Essentia + ONNX Runtime Classification (genre/mood/instrument via MIGraphX EP; TF fallback)
llama.cpp (HIP build) Music Flamingo GGUF inference
llama-cpp-python LLM revision (captioning benchmark)
autoawq, qwen-omni-utils Qwen2.5-Omni-7B-AWQ (captioning benchmark)
librosa, soundfile Audio I/O and analysis
timbral_models Audio Commons perceptual features (patched, cloned via setup script)

See requirements.txt for the full list.

Quick Start

# Setup
python -m venv mir && source mir/bin/activate
pip install -r requirements.txt
bash scripts/setup_external_repos.sh

# Test all features on a single file
python src/test_all_features.py "/path/to/audio.flac"

# Full pipeline (config-driven)
python src/master_pipeline.py --config config/master_pipeline.yaml

# Audio captioning benchmark (compare Flamingo, LLM revision, Qwen-Omni)
python tests/poc_lmm_revise.py "/path/to/audio.flac" --genre "Goa Trance" -v

ROCm GPU Environment

All ROCm environment variables are centralized in src/core/rocm_env.py and documented in config/master_pipeline.yaml. Every GPU-using script calls setup_rocm_env() before importing torch.

Key variables (set automatically, shell exports override):

export FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE
export PYTORCH_TUNABLEOP_ENABLED=1
export PYTORCH_TUNABLEOP_TUNING=0
export PYTORCH_ALLOC_CONF=garbage_collection_threshold:0.8,max_split_size_mb:512
export HIP_FORCE_DEV_KERNARG=1
export TORCH_COMPILE=0   # buggy with FA on RDNA

Documentation

Project Layout

src/
  core/           # Utilities: JSON handler, file utils, rocm_env, text normalization
  preprocessing/  # File organization, stem separation (Demucs, BS-RoFormer), loudness
  rhythm/         # Beat detection, BPM, syncopation, onsets, per-stem rhythm
  spectral/       # Spectral features, multiband RMS
  harmonic/       # Chroma, per-stem harmonic movement
  timbral/        # Audio Commons features, AudioBox aesthetics
  classification/ # Essentia, Music Flamingo (GGUF + Transformers)
  transcription/  # MIDI drum transcription (ADTOF, Drumsep)
  tools/          # Metadata lookup, training crops, statistical analysis (VIF/PCA/MI)
  crops/          # Crop-specific pipeline and feature extraction
tests/            # Benchmarks (audio captioning comparison)
config/           # YAML pipeline configuration
models/           # GGUF model files (Qwen3, GPT-OSS, Granite, Music Flamingo)
repos/            # External repos (cloned by setup script, not tracked)

Acknowledgements

This project builds on the following open-source work:

Project Use
Essentia (MTG, Universitat Pompeu Fabra) Genre, mood, instrument, voice classification; danceability, atonality
AudioBox Aesthetics (Meta) Perceptual quality scores (enjoyment, usefulness, production quality/complexity)
Stable Audio Tools (Stability AI) Target model this pipeline conditions
Music Flamingo (Amazon) AI music descriptions (8B multimodal LLM)
Granite (IBM) Caption revision / condensation
Qwen2.5-Omni (Alibaba) Captioning benchmark reference model
llama.cpp (Georgi Gerganov et al.) GGUF inference for Music Flamingo and LLM revision
BS-RoFormer (Roman Solovyev et al.) High-quality stem separation
Hybrid Demucs (Meta) Fast stem separation
ADTOF (Mickael Zehren) Automatic drum transcription to MIDI
Drumsep (Fraunhofer HHI) Drum stem separation
madmom (CP-JKU Linz) Tempo estimation
timbral_models (AudioCommons) Perceptual timbral features (brightness, hardness, warmth, etc.)
Plotly Interactive feature explorer visualisations
librosa Beat tracking, onset detection, spectral analysis
Rubber Band Library (Breakfast Quay) Pitch shifting and time stretching (pitch shifter GUI)
Bungee / bungee-python Pitch shifting and time stretching (pitch shifter GUI)
Pedalboard (Spotify) Pitch shifting and time stretching (pitch shifter GUI)
SoX Time stretching and pitch shifting via WSOLA/OLA (pitch shifter GUI)

License

TBD

About

Music Information Retrieval feature extraction framework for Stable Audio Tools conditioning

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages