Principal Bioinformatics Engineer

Nextflow pipelines and Azure Batch infrastructure for genomics

I build production genomic pipelines for WGS, pathogen surveillance, and multi-omics, and fix the Azure Batch and HPC bottlenecks that stall them.

Book a Technical Review

A bioinformatician who also runs the infrastructure

Bioinformatics engineer and consultant with a PhD and 15+ years in genomics. I write Nextflow and Snakemake pipelines and keep them running on Azure Batch and HPC clusters: the containerization, resource allocation, and cloud debugging that usually falls between the bioinformatics team and IT. Available for remote B2B work as an EU-based independent contractor (W-8BEN ready).

Core Services

01

Nextflow & Snakemake Pipeline Architecture

Reproducible pipelines for WGS, pathogen genomics, transcriptomics, AMR, and GMO tracking, built to run the same way on a laptop, an HPC cluster, or the cloud.

02

Azure Batch & HPC Troubleshooting

Deployment and debugging on Azure Batch and SLURM/SGE clusters. Docker containerization and resource tuning to cut cloud spend and clear the bottlenecks that stall large runs.

03

Scripting & Data Analytics

Automation and large-scale multi-omics data processing in Python, Bash, and R.

Case Study — European Food Safety Authority (EFSA)

WGS pipeline troubleshooting on Azure Batch

Challenge

As an external contractor on EFSA's high-throughput WGS pipeline (bacteria, fungi, and viruses), I traced two problems the internal IT team had been unable to resolve: storage limits on the container batch images, and inefficient resource allocation during large genomic runs.

Solution

Debugging the Docker containerization and tuning nextflow.config parameters, including custom Docker image mounts to work around Azure VM storage limits and distribute the HPC workload more evenly.

Impact

Storage bottlenecks gone, wasted cloud compute cut, and a stable pipeline that now supports ongoing pathogen and AMR surveillance.

Technical Stack

Pipeline Engines

  • Nextflow
  • Snakemake

Infrastructure & DevOps

  • Azure Batch
  • HPC Clusters (SLURM/SGE)
  • Docker
  • Git

Languages

  • Python
  • Bash
  • R

Domains

  • Genomics
  • Transcriptomics
  • Metagenomics
  • AMR

Let's scale your bioinformatics infrastructure.

If you have pipelines that need to move from a working prototype to something that runs reliably in production, get in touch.

Book a Technical Review ↗

or send a message directly

By submitting this form you agree to the Privacy Policy.