Skip to content

About

Pruning of a CNN-based classifier on CIFAR-10 via Genetic Algorithms.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Convolutional Neural Network Topology Reduction via Genetic Algorithms

Structured filter pruning of a CNN-based classifier using a genetic algorithm, evaluated on CIFAR-10. This project explores how much a trained convolutional network can be simplified, measuring parameter count, FLOPs, and inference time whilst preserving accuracy.

Overview

  • Structured pruning of convolutional filters where pruned filters are removed and all dependent layers (BatchNorm, subsequent convolutions, the final classification layer) are reshaped to match.
  • Genetic algorithm with binary filter masks as genetic encoding, microbial crossover, and elitism-preserving mutation.
  • Multi-objective fitness function trading off accuracy against parameter count, tunable via α/β weighting.
  • Evaluated across individual layers, combined multi-layer pruning, and fitness-hyperparameter configurations.

Architecture

The base model is a 5-layer CNN classifier trained on CIFAR-10 to a reference accuracy of 91.55%.

Base CNN architecture

Usage

# Optionally create and activate a virtual environment
python3 -m venv venv
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

# Change the working directory
cd src/

# Train the reference model
python train.py

# Run genetic-algorithm pruning (defaults: population 30, 20 generations)
python prune.py --layer_names conv4 conv5 --alpha 10 --beta 1e-6 --finetune

See config.py for the full set of tunable GA and pruning hyperparameters.

Methodology

Genetic encoding. Each individual is a binary mask per convolutional layer, indicating which filters are kept.

Selection. Top-K individuals by fitness score are retained each generation.

Crossover. Microbial crossover: the winning parent (by fitness) partially overwrites the losing parent's genes, per filter, with a fixed crossover probability.

Mutation. Random bit-flips per filter, with the current elite individual preserved unmutated across generations.

Fitness function:

fitness = α · accuracy − β · parameter_count

where α and β control the trade-off between preserving accuracy and minimizing model size. See Results for how different weightings affect outcomes.

Structured pruning implementation. Removing a filter from a convolutional layer requires updating every downstream dependency: the corresponding BatchNorm channel statistics, the input channels of the next convolutional layer, and the input features of the final linear classifier. This is handled explicitly rather than relying on masking alone, so pruned models are smaller (fewer real parameters, lower FLOPs), not just sparsified.

Acknowledgements:

  • Yang, C. et al., "Multi-Objective Pruning for CNNs Using Genetic Algorithm," 2019.
  • Harvey, I., "The Microbial Genetic Algorithm," 2009.

Experiments

1. Per-layer sensitivity

Hypothesis: earlier layers should tolerate less pruning (larger accuracy loss for smaller parameter savings), while deeper layers should tolerate more.

Each of the network's five convolutional layers was pruned independently, and accuracy vs. parameter reduction was measured across 10 runs per layer.

Accuracy and parameter reduction trade-off per layer

The results support the hypothesis: pruning conv1 costs almost no accuracy but also saves almost no parameters, while conv4 and conv5 tolerate substantial parameter reduction (34-45%) with a comparatively small accuracy drop.

2. Combined multi-layer pruning

Based on the per-layer results, the two deepest layers (conv4, conv5) were pruned simultaneously across several α/β fitness-weighting configurations, to see how the trade-off holds when compression is pushed further.

Accuracy across hyperparameter configurations Parameter reduction across hyperparameter configurations

Higher β (heavier penalty on parameter count) drives more aggressive compression, but the accuracy cost is not linear — the largest configuration (α=5, β=1e-5) trades a disproportionate amount of accuracy for a comparatively modest additional reduction. A Mann-Whitney U test between the two configurations (α=10, β=1e-7 vs. α=10, β=1e-6) found no statistically significant difference in accuracy (p > 0.05), so the weaker-compression configuration was dropped from further analysis in favor of the one with better parameter savings at equivalent accuracy.

3. Mutation rate sensitivity

As a follow-up question on the three remaining configurations: does the GA's mutation rate materially affect the final accuracy reached?

Impact of mutation rate across fitness configurations

The effect of mutation rate is not consistent across configurations, and the direction of the trend depends on how aggressively the network is being pruned. For the two less aggressive configurations (α=10, β=1e-6 and α=5, β=1e-6), lower mutation rates consistently outperform higher ones. This could be because with less aggressive pruning, the fitness landscape is comparatively smooth, and more exploitation (low mutation) is enough to converge well. For the most aggressive configuration (α=5, β=1e-5), this trend reverses: higher mutation rates outperform lower ones. One plausible explanation is that heavier pruning creates a bigger search space, where additional exploration is needed to escape weaker local optima.

Results

Variant # Parameters (×10⁶) Accuracy (%) ↑ GFLOPs ↓ Inference (ms) ↓
Reference 3.92 91.55 0.54 1.916
α=10, β=1×10⁻⁶ 1.66 (-58%) 85.86 0.41 1.643
α=10, β=1×10⁻⁶ (fine-tuned) — 90.01 — —
α=5, β=1×10⁻⁶ 1.55 (-60%) 84.61 0.40 1.583
α=5, β=1×10⁻⁵ 1.25 (-68%) 76.43 0.38 1.500

Inference time is averaged over 1000 runs, batch size 64, 32×32 inputs. Fine-tuning was performed for 5 epochs at a low learning rate. Bold indicates the best result in each column.

Fine-tuning the best-performing pruned configuration for just 5 epochs recovered accuracy to 90.01% while retaining a 58% parameter reduction compared to baseline, showing that most of the accuracy lost to pruning is recoverable with minimal additional training.

About

Pruning of a CNN-based classifier on CIFAR-10 via Genetic Algorithms.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages