Job Responsibilities:
· Monitoring & Operations
o Keep production healthy by watching the signals and acting on them. You’ll be the first set of eyes on alerts and the person flagging when costs drift.
o Monitor alerts across critical services using Azure Monitor, Log Analytics (KQL), Application Insights, Grafana, and Prometheus, and respond or escalate as appropriate.
o Track cloud costs, surface anomalies, and flag optimization opportunities to the team.
o Monitor and operate middleware services (PostgreSQL, Redis, Kafka); assist with root cause analysis on production issues.
· Infrastructure as Code & Automation
o Turn manual, repetitive operations into reliable, repeatable automation.
o Execute and maintain Terraform to provision and update Azure infrastructure.
o Write and maintain shell scripts that automate routine operational tasks.
o Use Helm and ArgoCD to support application deployment and configuration.
· Disaster Recovery
o Make sure the team can recover quickly and confidently when incidents happen.
o Build, document, and maintain disaster recovery runbooks for critical services.
o Validate recovery procedures regularly and keep runbooks accurate and current.
· Security & Governance
o Operate within established guardrails and help keep the estate clean and compliant.
o Apply least-privilege RBAC and work within existing Azure Policy and compliance guardrails.
o Handle day-to-day secrets operations using Key Vault and managed identities.
o Follow tagging and governance standards across subscriptions.
Required Qualifications:
· Solid hands-on experience operating production cloud infrastructure on Azure.
· Working knowledge of Azure core services: networking, compute, storage, and resource governance.
· Expert with Azure identity (Entra ID).
· Solid Linux and networking fundamentals.
· Practical Terraform experience provisioning and maintaining infrastructure.
· Strong scripting in Bash and Python.
· Experience with monitoring tooling: Azure Monitor, Log Analytics / KQL.
· Application Insights, Cost Management, Grafana, and Prometheus.
· Experience operating common middleware (PostgreSQL, Redis, or Kafka).
· Familiarity with RBAC, Key Vault, and operating within governance guardrails.
Preferred Qualifications:
· Azure certifications such as AZ-104 or AZ-400.
· Experience writing or maintaining disaster recovery runbooks.
· Experience with multi-subscription environments.
· FinOps / cost optimization exposure.