Hello, I am

Djibril Faty

15 years operating critical platforms (SLAs, 24/7 on-call, P1 incidents) and a Cloud Native specialization (Kubernetes, Terraform, GitOps, SRE). Focused on DevOps transformations in demanding environments, especially Sovereign Cloud. I bring a reliability culture to teams industrializing their practices.

Cloud Engineer / Site Reliability Engineer (SRE)

My Approach

These three values guide my day-to-day work on critical production environments.

Reliability & Availability Icon

Reliability & Availability

15 years ensuring high availability and Disaster Recovery for critical production platforms: SLA compliance, 24/7 on-call, and P1 incident management, with a strong focus on service continuity.

Automation & SRE

Monitoring, logging, and automation (Ansible, scripting) to make operations more reliable, with an SRE approach to incident management.

Cloud Native & Sovereign

AWS Certified SysOps Administrator with a Sovereign Cloud / SRE specialization: Kubernetes, Terraform, and GitOps to industrialize reliable platforms, including on sovereign infrastructure (OVH, Proxmox).

Skills & Tools

The technologies and tools I master, drawn from 15 years of production operations and my Sovereign Cloud / SRE specialization.

Cloud & Containerization

Public and sovereign cloud infrastructure, virtualization, and container orchestration, down to Kubernetes networking.

AWS

OVH

Proxmox

Kubernetes

Helm

Docker

Ingress-NGINX

Traefik

cert-manager

Cilium

Infrastructure as Code & Automation

Infrastructure provisioned and configured as code: reproducible, version-controlled, and automated.

Terraform

Ansible

Python

Bash

CI/CD & GitOps

Continuous integration and delivery pipelines, GitOps deployments, and centralized management of secrets and images.

Git

GitLab CI/CD

GitHub Actions

ArgoCD

Vault

Harbor

Observability & SRE

Metrics, alerting, and centralized logging in service of high availability, with rigorous incident management through to root cause analysis.

Prometheus

Grafana

Alertmanager

Centralized logging

High availability & DR

RCA

Systems, Networks & Data

Linux administration, networking fundamentals, and distributed block and object storage for Kubernetes.

Linux (RHEL)

TCP/IP & DNS

Load balancing

SSH / SFTP

Longhorn

MinIO

SeaweedFS

My Journey

15 years of experience on critical production platforms, from engineering school to Cloud Engineering.

Sovereign Cloud Engineer / SRE Specialization

Intensive 450-hour production-oriented program: Kubernetes, Terraform, GitOps, observability, SRE practices, and Sovereign Cloud. One-month capstone project defended before a professional jury.

Cloud Certification — AWS SysOps Administrator

AWS SysOps Administrator certification, complementing hands-on Kubernetes, Terraform, and Ansible expertise.

Operation & Support Engineer

Led support for Cloud platforms (OVH, AWS), ensuring reliability, high availability, and Disaster Recovery across Big Data client accounts (geolocation, contextual marketing, SMS Gateway), with an SRE approach: monitoring, logging, automation, and incident management.

Technical Validation Lead

Integration and validation of IPTV/VOD service platforms, end-to-end interconnection testing, non-regression test automation, and coordination of internal and partner teams.

Contact

Feel free to reach out for any question or collaboration opportunity.

LinkedIn Icon

LinkedIn

linkedin.com/in/dfaty

Connect on LinkedIn Connect on LinkedIn Arrow

Let's Collaborate

A question, a project, or just want to talk Cloud and system reliability? Feel free to reach out.

Contact Me