Skip to content
@refresh-bio

REFRESH Bioinformatics Group

Pinned Loading

  1. KMC KMC Public

    Fast and frugal disk based k-mer counter: counts 55-mers in 36 gzipped FASTQ files totaling 614 GB (WGS paired-end reads of NA12878, Illumina HiSeq 2000) in under 1.5 h, using 12 threads, 32 GB of …

    C++ 341 85

  2. FAMSA FAMSA Public

    Algorithm for ultra-scale multiple protein sequence alignments: 3 million ABC transporters analyzed in 5 minutes and 18 GB of RAM.

    C++ 273 35

  3. colord colord Public

    A versatile compressor of third generation sequencing reads.

    C++ 53 15

  4. kmer-db kmer-db Public

    Fast and memory-efficient tool for large-scale k-mer analyses (indexing, querying, comparison): 16 million viral contigs analyzed in less than an hour.

    C++ 102 18

  5. agc agc Public

    Assembled Genomes Compressor (AGC) squeezing 290 GB of Human Pangenome Project assemblies down to <1.5 GB (>190× ratio), with sample and contig extraction in seconds.

    C++ 187 20

  6. vclust vclust Public

    Fast and accurate tool for calculating Average Nucleotide Identity (ANI) and clustering virus genomes and metagenomes

    Python 111 5

Repositories

Showing 10 of 31 repositories
  • agc Public

    Assembled Genomes Compressor (AGC) squeezing 290 GB of Human Pangenome Project assemblies down to <1.5 GB (>190× ratio), with sample and contig extraction in seconds.

    C++ 187 MIT 20 9 3 Updated Oct 5, 2026
  • GDC2 Public

    Genome Differential Compressor 2 (GDC2) for FASTA collections, achieving ~9,500× compression on 1000GP data at 200 MB/s, featuring 1 GB/s decompression and direct single-genome extraction.

    C++ 8 GPL-2.0 3 1 0 Updated Oct 5, 2026
  • MuGI Public

    Space-efficient MUlti-Genome Index (MuGI) for FASTA reference + VCF collections, squeezing 1,092 diploid human genomes (6.7 TB) into 6.5–32 GB RAM for microsecond exact and mismatch search.

    refresh-bio/MuGI's past year of commit activity
    C++ 1 GPL-2.0 1 1 0 Updated Oct 5, 2026
  • GTShark Public

    GTShark: Fast VCF GenoType compressor with rapid decompression and reference-based compression for external samples, squeezing a human genome down to ~65 KB.

    refresh-bio/GTShark's past year of commit activity
    C++ 3 GPL-3.0 4 2 0 Updated Oct 5, 2026
  • GTC Public

    Ultrafast VCF GenoType Compressor (GTC) with instant random access, shrinking the 4.3 TB HRC dataset (27k samples, 40M variants) down to under 4 GB.

    C++ 16 GPL-3.0 8 2 0 Updated Oct 5, 2026
  • TGC Public

    Thousands Genome Compressor (TGC) operating on stripped VCF data, squeezing an individual human genome down to ~400 KB.

    refresh-bio/TGC's past year of commit activity
    C++ 1 1 0 0 Updated Oct 5, 2026
  • VCFShark Public

    VCF file compressor reaching over 30-fold efficiency on massive genotype datasets, with processing speeds up to 100 MB/s and peak memory usage below 30 GB.

    refresh-bio/VCFShark's past year of commit activity
    C++ 19 GPL-3.0 1 1 1 Updated Oct 5, 2026
  • MKMC Public

    The tool for creating k-mers counts matrices of many samples and computing various k-mer based statistics without need of a reference genome. The matrix may be built in less than 2 hours using less than 20 GB of RAM for reads of raw size over 12 TB.

    refresh-bio/MKMC's past year of commit activity
    C++ 4 GPL-3.0 0 0 0 Updated Sep 17, 2026
  • RECKONER Public

    Illumina read substitution and indel error corrector. Processes human WGS 45x depth reads in about 4 hours, resulting in the reads quality similar to other top algorithms.

    refresh-bio/RECKONER's past year of commit activity
    C++ 10 1 2 0 Updated Sep 17, 2026
  • LZ-ANI Public

    Fast and accurate tool for calculating Average Nucleotide Identity (ANI) among virus and bacteria genomes

    refresh-bio/LZ-ANI's past year of commit activity
    C++ 13 GPL-3.0 2 2 0 Updated Sep 12, 2026

Top languages

Loading…

Most used topics

Loading…