US Citizen or GC Holder Required - FedRAMP Requirement
REMOTE - prefer PST hours
Max vendor rate is $90/hr
5 days per week/ 8 hours per day
Senior Kubernetes Platform Engineer Lead
Top Skills Required :
1. Kubernetes (GKE / AKS / EKS) Production grade, regulated environments
2. CI/CD & Automation (GitHub Actions, Jenkins, ArgoCD, GitOps)
3. FedRAMP High / IL5 Compliance & Observability (Prometheus, Grafana, ELK, OpenTelemetry)
About the Role
We build technology that simply works reliable, secure, and easy to use. We re looking for a Senior Site Reliability Engineer (SRE) - Technical Leader to help us design, operate, and scale a Kubernetes-based platform supporting highly regulated environments, including FedRAMP High and DoD IL5.
This role sits at the intersection of software engineering and infrastructure. You ll work closely with engineers across the stack to ensure our platform is resilient, observable, compliant, and developer-friendly without slowing teams down.
What You ll Do
Design, build, and operate production-grade Kubernetes platforms in regulated environments
Improve system reliability through automation, thoughtful design, and continuous iteration
Define and drive SLOs, SLIs, and error budgets to guide reliability decisions
Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
Implement robust observability (metrics, logs, traces) to make systems understandable and actionable
Reduce operational toil by automating repetitive processes and improving workflows
Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity
Support ATO processes, including documentation, controls implementation, and audit readiness
Participate in on-call rotations supporting customer requests and paging alerts
Participate in incident response, blameless postmortems, and continuous improvement efforts
Help shape a platform that engineers enjoy using
What You Bring
10+ years of experience in SRE, DevOps, or infrastructure engineering
Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)
Hands-on experience working in FedRAMP High and/or DoD IL5 environments
Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals
Experience with Infrastructure as Code (Terraform preferred)
Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
Proficiency in scripting or programming (Python, Go)
Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)
Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)
Nice to Have
Experience with service mesh technologies (Istio, Linkerd)
Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
Experience with GitOps workflows
Exposure to multi-cluster or hybrid cloud architectures
Knowledge of FIPS-compliant systems or DoD Cloud SRG
Relevant certifications (CKA, CKS, cloud provider certs, Security+)
How We Work
We value simplicity, transparency, and collaboration
We believe in blameless culture and learning from incidents
We focus on building tools and platforms that empower other engineers
We balance reliability, security, and developer experience not one at the expense of the others
What Success Looks Like
Our platform is reliable, scalable, and easy to operate
Engineers can deploy confidently in high-compliance environments
Observability provides clear, actionable insights
Operational overhead is minimized through automation
Compliance requirements are met seamlessly as part of the platform