Muhammad N.
Senior AWS DevOps & Kubernetes Engineer | Terraform, EKS, CI/CD
Most AWS environments I'm brought into work fine โ until traffic doubles, the bill jumps, or a release fails at 2am. I rebuild them so they stop breaking: production EKS, Terraform that holds, deploys nobody babysits. Top Rated Plus, 3,000+ hours. WHAT I'VE ACTUALLY DONE โ Cut a SaaS client's AWS bill from roughly $70K/month to $45K/month over three months, without touching the product roadmap โ Rebuilt a failing EKS cluster; it has held above 99.9% for nine months since, with no paging incidents โ Replaced an 18-step manual release with an 8-minute automated pipeline and a rollback that actually works โ Designed and delivered a HIPAA-ready AWS environment for a healthcare SaaS โ Built multi-account AWS infrastructure for a fintech payment platform โ Held one AWS DevOps engagement past 1,800 hours. Nobody keeps an infrastructure engineer that long unless the infrastructure works. WHERE I'M USUALLY BROUGHT IN โ AWS was built fast and now has to be production-grade โ An EKS or GKE cluster is unstable, expensive, or nobody on the team wants to touch it โ Terraform exists, but environments drift and every change feels risky โ Releases need a human watching them โ The cloud bill is growing faster than revenue โ There's no real alerting, dashboards, or incident visibility โ IAM, secrets, and network access need cleaning up before an audit โ You're moving off VMs or hand-built servers and can't afford downtime WHAT I DO AWS infrastructure โ VPC and network design, IAM, EKS, ECS, RDS, Lambda, S3, CloudFront, Route53, WAF, CloudWatch, SQS/SNS, API Gateway, autoscaling, and cost controls. AWS is my primary platform; I also work in GCP (GKE, Cloud SQL, Cloud Build). Production Kubernetes โ EKS and GKE with Helm, ArgoCD, ingress controllers, autoscaling, network policies, pod disruption budgets, resource right-sizing, and the cluster troubleshooting that tends to happen at bad hours. Terraform and Infrastructure as Code โ reusable, documented modules. Remote state, multi-environment layouts, drift reduction, and refactoring inherited Terraform that nobody on the team understands anymore. Also CloudFormation and CDK. CI/CD โ GitHub Actions, GitLab CI, Jenkins, Azure DevOps, CodePipeline. GitOps workflows, blue-green and staged rollouts, and rollbacks that hold under pressure. Observability and reliability โ Prometheus, Grafana, Loki, OpenTelemetry, Datadog, ELK, CloudWatch, PagerDuty. Alerts your team acts on instead of mutes. Cloud cost optimization โ right-sizing, orphaned resource cleanup, RDS and S3 lifecycle tuning, NAT gateway and data transfer review, Savings Plans and Reserved Instance analysis, Kubernetes resource efficiency, and a monthly cost report you can actually read. Security and production readiness โ least-privilege IAM, Vault and External Secrets, security groups, WAF, backup and DR planning, public exposure review, and SOC 2-ready cloud baselines. Architecture design โ when you need the plan before the build: target architecture, the tradeoffs behind it, a cost model, and a migration sequence. Clients have re-hired me for this specific piece more than any other. HOW I WORK I don't build demo infrastructure that only works during handoff. Your team has to be able to read it, run it, and change it after I'm gone: clean documentation, real automation, least-privilege access, monitoring that means something, and infrastructure that stays boring. I work overlapping your hours, respond within 4 hours, and tell you when something is a bad idea before you pay me to build it. CREDENTIALS AWS Certified Solutions Architect โ Professional AWS Certified Solutions Architect โ Associate HashiCorp Certified: Terraform Associate Microsoft Certified: Azure Fundamentals BS Engineering, Ghulam Ishaq Khan Institute (GIKI) TELL ME WHAT'S BROKEN Message me with what's failing right now โ the bill, the cluster, the pipeline, the 3am page. I'll tell you what I'd check first and whether I'm the right person for it. If I'm not, I'll say so.