Talent.com
Site Reliability Engineer

Site Reliability Engineer

DevOps projectsBerkeley, CA, United States
21 hours ago
Job type
  • Full-time
Job description

Site Reliability Engineer

About the Company

LMArena is an engineering-first startup redefining how the world evaluates large language models. Created in 2023 by UC Berkeley researchers, our neutral, community-driven benchmarking platform attracts over one million monthly users—pairwise comparing leading models from OpenAI, Google, Anthropic, and more—to deliver real-time insights into the rapidly evolving LLM landscape. LMArena is scaling fast to build the next generation of AI testing infrastructure and set the industry standard for model evaluation.

Position Overview

We are seeking an experienced, security‑minded Site Reliability Engineer to own and elevate our infrastructure, processes, and operational security. We are cognizant that much of the security domain sits within the SRE’s world these days and are building our team accordingly.

  • Take end‑to‑end ownership of infrastructure operations across Cloudflare, Vercel, and our CI / CD pipelines.
  • Embed security best practices into every layer of the stack, ensuring resilience against emerging threats.
  • Establish processes and procedures that promote efficient onboarding and ramping up new team members, and mentor incoming and more junior members of the team.

This role is ideal for a seasoned SRE who thrives at the intersection of reliability, performance, and security, and who brings the rigor needed to keep fast‑moving product teams focused on innovation.

Key Responsibilities

  • Infrastructure as Code – Manage Terraform modules and secrets pipelines; champion immutable, auditable infrastructure.
  • Cloudflare Operations – Configure, monitor, and harden WAF, DDoS protections, bot management, and CDN caching strategies.
  • Vercel & Edge Runtime – Own deployment architecture, performance tuning, and incident response for our Next.js‑based front end and Edge Functions.
  • CI / CD & Release Engineering – Design, implement, and maintain secure pipelines (GitHub Actions, Vercel integrations) with automated testing and vulnerability scanning.
  • Change Management & Documentation – Establish and enforce a lightweight but disciplined RFC / change‑control process; maintain comprehensive runbooks and architecture diagrams.
  • Observability & Incident Response – Expand monitoring, logging, and alerting; lead post‑incident reviews and drive continual improvement.
  • Mentorship – Provide day‑to‑day guidance to engineers and junior SREs, fostering a culture of ownership and learning.
  • Compliance Support – Partner with ProdSec and GRC teams on SOC2, ISO27001, and customer security questionnaires.
  • Manage and maintain internal and external facing infrastructure.
  • Maintain and configure log aggregation requirements, and the infrastructure used to store them across the business.
  • Required Qualifications

  • 7+years in SRE / DevOps roles for high‑traffic SaaS or consumer web products.
  • Proven expertise securing and scaling Cloudflare and Vercel (or comparable CDN / edge and serverless platforms).
  • Deep understanding of web application security, networking, TLS, and zero‑trust principles.
  • Strong proficiency with infrastructure as code (Terraform, Pulumi, or similar), and serverless build pipelines (GitHub Actions or similar).
  • Strong programming abilities (Golang, python, TypeScript) and scripting.
  • Demonstrated success designing and enforcing change‑management workflows.
  • Excellent written communication—able to produce clear runbooks and architecture docs.
  • Track record mentoring or leading junior engineers.
  • Experience with container orchestration (Kubernetes or Nomad).
  • Experience with serverless stacks.
  • Certifications such as AWS / GCP Professional, GIAC‑GCSA, CKS, or CISSP.
  • Why You’ll Love Working Here

  • Impact – You’ll set the foundation for reliability and security across a rapidly growing AI benchmarking platform.
  • Compensation – Competitive salary, meaningful equity, comprehensive benefits, and professional‑development budget.
  • #J-18808-Ljbffr

    Create a job alert for this search

    Site Reliability Engineer • Berkeley, CA, United States

    Related jobs
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    ConductorOneSan Francisco, CA, United States
    Full-time
    ConductorOne is the first AI-native identity security platform that protects every identity : human, non-human, and AI.With powerful automation, platform-level AI, and out-of-the-box connectors, it ...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer I

    Site Reliability Engineer I

    ProsperSan Francisco, CA, United States
    Full-time
    As a Site Reliability Engineer I at Prosper, you will play a crucial role in enhancing the reliability, scalability, and maintainability of our technology platform. This entry-level position is desi...Show moreLast updated: 23 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    Rivago Infotech IncSan Francisco, CA, United States
    Full-time
    Staff Site Reliability Engineer (SRE).As our Staff SRE, you'll be the primary expert responsible for our entire compute ecosystem. Your key responsibilities will include : .Design, implement, and lead...Show moreLast updated: 5 days ago
    • Promoted
    Senior Site Reliability Engineer – Platform

    Senior Site Reliability Engineer – Platform

    Icon VenturesSan Francisco, CA, United States
    Full-time
    At Quizlet, our mission is to help every learner achieve their outcomes in the most effective and delightful way.We blend cognitive science with machine learning to personalize and enhance the lear...Show moreLast updated: 1 day ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    CorelightSan Francisco, CA, United States
    Full-time
    Senior Site Reliability Engineer.We are looking for a Senior Site Reliability Engineer to design, automate, and scale cloud and hybrid platforms that power AI / ML workloads and SaaS services.You\'ll...Show moreLast updated: 4 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    Runloop AISan Francisco, CA, United States
    Full-time
    Runloop is building the foundational infrastructure for the next generation of AI development.We provide AI engineers and data scientists with lightning-fast, secure, and reproducible code sandboxe...Show moreLast updated: 27 days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    AlchemySan Francisco, CA, United States
    Full-time
    Our mission is to bring web3 to a billion people, by providing builders with the tools they need to build exceptional onchain products. Alchemy is the only complete developer platform that offers th...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    Together AISan Francisco, CA, United States
    Full-time
    As a Site Reliability Engineer (SRE) at Together, you are responsible for keeping all user-facing services and production systems running smoothly. You are a blend of a pragmatic operator and a soft...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    Alembic TechnologiesSan Francisco, CA, United States
    Full-time
    Senior Site Reliability Engineer.This range is provided by Alembic Technologies.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.We’re looking fo...Show moreLast updated: 1 day ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    AlembicSan Francisco, CA, United States
    Full-time
    We’re looking for an experienced.Site Reliability Engineer (SRE).You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure that powers our core platform—...Show moreLast updated: 2 days ago
    • Promoted
    Site Reliability Engineer (SRE)

    Site Reliability Engineer (SRE)

    BasetenSan Francisco, CA, United States
    Full-time
    Baseten powers inference for the world's most dynamic AI companies, like OpenEvidence, Clay, Mirage, Gamma, Sourcegraph, Writer, Abridge, Bland, and Zed. By uniting applied AI research, flexible inf...Show moreLast updated: 23 days ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    Loft OrbitalSan Francisco, CA, United States
    Full-time
    Senior Site Reliability Engineer.This range is provided by Loft Orbital.Your actual pay will be based on your skills and experience — talk with your recruiter to learn more.Loft Orbital is revoluti...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    PrimerSan Francisco, CA, United States
    Full-time
    Primer helps B2B products break out of the B2C-centric marketing box.Our platform turns consumer ad channels, data streams, and emerging AI workflows into measurable growth engines for go-to-market...Show moreLast updated: 30+ days ago
    • Promoted
    Senior Site Reliability Engineer

    Senior Site Reliability Engineer

    HiveSan Francisco, CA, United States
    Full-time
    Hive is the leading provider of cloud-based AI solutions to understand, search, and generate content, and is trusted by hundreds of the world's largest and most innovative organizations.The company...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    SpeakSan Francisco, CA, United States
    Full-time
    Our mission is to reinvent the way people learn, starting with language.Learning a language can change a life by opening doors to new cultures, careers, and communities. Two billion people around th...Show moreLast updated: 1 day ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    P2PSan Francisco, CA, United States
    Full-time
    Our mission is to bring web3 to a billion people, by providing builders with the tools they need to build exceptional onchain products. Alchemy is the only complete developer platform that offers th...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer II

    Site Reliability Engineer II

    Hinge HealthSan Francisco, CA, United States
    Full-time
    From scaling Kubernetes clusters to improving observability with Datadog, we build the tooling and automation that empower product teams to ship with confidence. Collaborate with engineering teams t...Show moreLast updated: 30+ days ago
    • Promoted
    Site Reliability Engineer

    Site Reliability Engineer

    Rockwoods IncPleasanton, CA, United States
    Full-time
    Note : Candidates must have relevant experience in Medical / Healthcare domains, this is mandatory.Senior SRE Engineer - Pleasanton, 5 days office. Primary work : 24x7 On-call support and setting up mo...Show moreLast updated: 30+ days ago