Talent.com
AgileEngine
DevOps / Site Reliability Engineer ID70127AgileEngine • Houston, TX, us
DevOps / Site Reliability Engineer ID70127

DevOps / Site Reliability Engineer ID70127

AgileEngine • Houston, TX, us
5 days ago
Job type
  • Full-time
Job description
Job Description
AgileEngine is an Inc. 5000 company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries. We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.

WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!

ABOUT THE ROLE
We are looking for a DevOps / Site Reliability Engineer to maintain operational resilience and 24/7 stability for a multi-cloud enterprise security program, serving as Incident Commander on major and critical incidents while also owning IaC, CI/CD pipelines, and CSPM telemetry. You will drive major-incident calls, own post-incident remediation follow-through, draft stakeholder communications, and develop divisional incident-management playbooks alongside multi-cloud security guardrails using Terraform and Wiz. The role requires 5+ years of SRE experience with hands-on incident command in a 24/7 financial services environment.

WHAT YOU WILL DO
- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams, making time-critical decisions, and coordinating cross-functional responders under pressure.
- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.

MUST HAVES
- You must be authorized to work for ANY employer in the US (e.g., Green card holders, TN visa holders, GC EAD, H4 EAD, U4U with EAD), as we are unable to sponsor or take over employment visa sponsorship at this time;
- 5+ years of experience.
- In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles.
- Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting.
- Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment.
- Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed.
- Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences.
- Direct experience authoring divisional/group incident-management playbooks and escalation procedures.
- Fully autonomous.
- Drives the architecture of complex automated runbooks and mentors Middle-level SREs.
- Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz.
- Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
- Upper-intermediate English level.

NICE TO HAVES
- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.

PERKS AND BENEFITS
- Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
- Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
- Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
- Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
- Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
- Well-being & support: access local well-being programs and people-focused support tailored to your location


Requirements
5+ years of experience. In-depth architectural expertise in multi-cloud defense, federated IAM, and zero-trust principles. Strong practical experience with Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting. Senior-level, hands-on incident-command experience driving major/critical incident calls to resolution in a 24x7 production environment. Proven track record of remediation follow-up — coordinating with teams and holding owners accountable until issues are fully closed. Demonstrated skill drafting and issuing incident notification communications to both technical and executive audiences. Direct experience authoring divisional/group incident-management playbooks and escalation procedures Fully autonomous. Drives the architecture of complex automated runbooks and mentors Middle-level SREs. Extensive experience deploying and tuning APIs from modern CNAPP/CSPM platforms, ideally Wiz. Prior experience building platforms subject to strict financial compliance standards (PCI-DSS, SOC2).
Create a job alert for this search

DevOps / Site Reliability Engineer ID70127 • Houston, TX, us

Similar jobs

Senior Manager of Site Reliability Engineering - Data Protection and Recovery

ChaseHouston, TX, United States
Full-time

Senior Manager Of Site Reliability EngineeringGuide and shape the future of technology at a globally recognized firm, driven by pride in ownership.As a Senior Manager of Site Reliability Engineerin... Show more

 • Promoted

Solutions Engineer - Global

SHIHouston, TX, United States
Full-time

The Global Solutions Engineer collaborates with account teams to assess customer data center environments and design tailored infrastructure solutions that align with business objectives.This role ... Show more

 • Promoted

NA IT Remote Site Associate

Edp Renovaveis Servicios Financieros SAHouston, TX, United States
Remote
Full-time +1

What you will doRole Overview :The Remote Site Associate will support the planning, coordinate, direct and design IT-related activities of new and current projects (including field development offi... Show more

 • Promoted

Systems Engineer

UpStream Global ServicesHouston, Houston, US
Full-time

We are seeking an experienced and detail-oriented Systems Engineer to join our Information Technology team.In this role, you will be responsible for designing, implementing, maintaining, and optimi... Show more

Solutions Engineer, Remote

Aperia TechnologiesHouston, TX, United States
Remote
Full-time

Aperia is unlocking a new era of efficiency and sustainability for commercial vehicle fleets, by developing innovative hardware and data analytics solutions.Inventors of the award- winning and disr... Show more

 • Promoted

Infrastructure Lead-Remote

SBS CorpHouston, TX, United States
Remote
Full-time

Client is looking for Infrastructure Lead Remote Contract AWS, EKS, ECS, S3, GPT2.CI-CD, DevOps, Lambda Please reach me at raghu(at)mysbscorp(dot)com. Show more

 • Promoted

Senior Substation Engineer - Level 4

QcellsHouston, TX, US
Remote
Full-time +1

Title: Senior Substation Engineer – Level 4.Department: Engineering Center.Supervisor: Manager of HV Engineering.Position Status: Permanent, Full-Time.Work Status (Remote/Hybrid/In-Office): Remote.... Show more

GNC Engineer - Remote

KBR Inc.Houston, TX, United States
Remote
Full-time

JOB DESCRIPTIONTitle:GNC Engineer - RemoteBelong.KBR! Around here, we define the future.We are a company of innovators, thinkers, creators, explorers, volunteers, and dreamers.But we all share one ... Show more

 • Promoted

Disaster Recovery (DR) Lead

Ace StackHouston, TX, United States
Full-time

Location: Houston, TX Duration: Contract.The Disaster Recovery (DR) Lead is responsible for designing, implementing, and managing the organizations disaster recovery and business continuity strateg... Show more

 • Promoted

Systems Engineer

Qureos IncHouston, Texas, United States
Full-time

We are seeking an experienced and detail-oriented Systems Engineer to join our Information Technology team.In this role, you will be responsible for designing, implementing, maintaining, and optimi... Show more

Remote Operations Support Engineer

Reasonn LLCHouston, TX, United States
Remote
Permanent

Position:Oil and Gas - Remote Operations Support Engineer Location:Houston, TX Duration:Longterm Contract Roles and Responsibilities:Deliver SAS support for Remote Operations Centers, including Wel... Show more

 • Promoted

Assistant Chief Engineer - Critical Environments

InfinitiveHouston, TX, United States
Full-time

Assistant Chief Engineer - Critical EnvironmentsAs an Assistant Chief Engineer for Data Center Facilities at JLL, you will play a crucial role in ensuring the smooth operation and maintenance of mi... Show more