DevOps & SRE Engineer specializing in infrastructure reliability, incident management, and self-healing systems. Building AI-powered platforms for automated incident detection, diagnosis, and controlled remediation. Currently a master's student in International Software Systems Science at the University of Bamberg.
An AI-powered platform for automated incident detection, diagnosis, controlled remediation, and recovery verification. Implements a complete 9-state incident lifecycle with evidence collection, deterministic diagnosis, optional AI/RAG-enhanced root cause analysis, allowlisted remediation with risk-based approval workflow, and automatic recovery verification.
- Self-Healing Pipeline: Detection β Incident β Evidence β Diagnosis β Remediation β Verification β Resolution
- Key Features: HTTP health & Prometheus detectors, ChromaDB runbook retrieval, SLO/SLI monitoring (MTTD, MTTR), cost optimization, rate limiting, RBAC, audit logging
- Infrastructure: Docker Compose, Kubernetes/Helm, Terraform, Argo CD GitOps, GitHub Actions CI/CD with Trivy/Gitleaks security scanning
- Testing: 39 tests (unit, integration, end-to-end)
- Tools:
Pythonβ’FastAPIβ’PostgreSQLβ’Redisβ’Prometheusβ’Grafanaβ’ChromaDBβ’Dockerβ’Kubernetesβ’Helmβ’Terraformβ’Argo CDβ’GitHub Actions
An ML microservice for real-time and batch customer-churn prediction, with feature explanations and data-drift monitoring.
- Tools:
Pythonβ’Scikit-Learnβ’FastAPIβ’Streamlitβ’Evidently AIβ’Docker
A backend data pipeline for ETL processing, data validation, database integration, and analytics APIs.
- Tools:
Pythonβ’Pandasβ’SQLAlchemyβ’PostgreSQLβ’FastAPI
- DevOps & Infrastructure Automation
- Cloud Engineering & Platform Reliability
- Kubernetes & Container Orchestration
- CI/CD Pipelines & Infrastructure as Code
- System Observability & Monitoring
- AI-Assisted Operations & Self-Healing Systems
- SRE Practices: SLO/SLI, Incident Management, Root Cause Analysis
