Talent.com

Cloud engineer Jobs in Renton, WA

Create a job alert for this search

Cloud engineer • renton wa

Last updated: 12 hours ago

Cloud & Customer Solutions Engineer - DC GPU

Advanced Micro Devices, IncBellevue, Washington, United States
Full-time +1

At AMD, we believe technology has the power to solve the world’s most important challenges.From advancing healthcare and scientific discovery to powering AI and the technologies people rely on ever... Show more

Senior Cloud Architect

Globenet Consulting CorpBELLEVUE, WA, US
$130,000.00 yearly
Full-time

Location: Fort Belvoir, VA 22060.Let’s Create Our Future Together at The AES Group!.We are seeking an experienced Senior Cloud Architect to lead the migration of applications and data from a closed... Show more

Director of Engineering - Application, Cloud and Offensive Security

SnowflakeBellevue, WA, United States
Full-time

At Snowflake, we are powering the era of the agentic enterprise.To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they ... Show more

Security Engineer

SnapchatBellevue, WA, United States
$199,000.00–$297,000.00 yearly
Full-time

Snap Inc is a technology company.We believe the camera presents the greatest opportunity to improve the way people live and communicate.Snap contributes to human progress by empowering people to ex... Show more

 • New!

Building Engineer

AA2ITMercer Island, WA, United States
Full-time

Bill Rate: 35-$38/Hr Hours: M-F 6am - 2:30pm Location: 3003 77th Ave SE, Mercer Island WA.Overview of Work Environment/Client Nuances: Class A office Building.There will be a lot of interaction wit... Show more

Amazon Dedicated Cloud Engineer, Kumo Enigma

Amazon Web Services, Inc.Bellevue, Washington, USA
Full-time

As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon’s unique experience and expertise, that are used by millions of companies ... Show more

Technical PM III (Government Cloud/M365)

ActalentBellevue, WA, United States
$70.00–$72.00 hourly
Full-time

This Technical Program Manager III will support the Hyper Control Environment product, a platform that supports government contract initiatives.The TPM will act as a key liaison between customers, ... Show more

 • New!

Sales Engineer

Olympus ControlsKent, WA, US
$78,000.00–$126,000.00 yearly
Full-time

Automation Company that has an opportunity for a Sales Engineer to join our growing company.We offer a dynamic environment with competitive salary and health benefits for the right type of highly m... Show more

Manufacturing Engineer

GpacRenton, Washington, United States
Full-time

A well-established manufacturer is searching for a Manufacturing Engineer to join their team!.A leader in their respective markets while also committed to both customers and employees.Great benefit... Show more

Weld Engineer

ProtingentBellevue, WA, US
$37.00–$74.00 hourly
Permanent

Protingent Staffing has an exciting contract Weld Engineer opportunity.Interface with various project teams to provide welding configuration, geometry and methods guidance for Company's reactor dev... Show more

Sr. Global Supply Manager, Amazon Cloud Logistics

Amazon Data Services, Inc.Bellevue, Washington, USA
Full-time

The AWS Cloud Logistics team is seeking a highly skilled and motivated Supply Chain Specialist in Bellevue, WA / Austin, TX / Florence, KY to manage day-to-day warehousing operations supporting Dat... Show more

DevOps Engineer

Glint Tech Solutions LLCBellevue, WA, USA
Full-time
Quick Apply

Glint Tech Solutions is a women-owned global staffing and IT recruiting firm connecting top technical talent with leading enterprise clients across the United States and Canada.Our client, a leadin... Show more

Engineer III

Puget Sound EnergyBellevue, WA, US
Full-time

Puget Sound Energy is looking to grow our community with top talented individuals like you!  With our rapidly growing, award winning energy efficiency programs, our pathway to an exciting and innov... Show more

Cloud & Customer Solutions Engineer - DC GPU

AMDBellevue, WA, US
Full-time +1

At AMD, we believe technology has the power to solve the world's most important challenges.From advancing healthcare and scientific discovery to powering AI and the technologies people rely on ever... Show more

Manufacturing Engineer

XANFAB INCSeatac, WA, US
$60,000.00–$95,000.00 yearly
Full-time +1

Xanfab Space is an Electronics Manufacturing Service (EMS) provider based in SeaTac, WA, specializing in space applications.We take a meticulous, detail-driven approach to manufacturing, and we are... Show more

Cloud Solutions Engineer – Azure

Bright Vision TechnologiesIssaquah, WA, US
$100,000.00–$150,000.00 yearly
Full-time
Quick Apply

Cloud Solutions Engineer – Azure - Remote    Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise s... Show more

Senior Cloud Marketing Manager

NetAppBellevue, WA, United States
Full-time

Senior Manager, Cloud Marketing.The Senior Manager, Cloud Marketing is the global marketing business owner for NetApp's premier partnerships with the hyperscalers.This role manages marketing's full... Show more

Network Engineer

PACCARRenton, WA, US
Full-time

This position may be eligible for an employee referral bonus.Please review the Employee Referral Bonus Instructions for information on steps or contact your HR representative.Fortune 500 company es... Show more

Cloud Technical Account Manager, ES - SI - Media and Entertainment

AmazonBellevue, WA, United States
Full-time

As part of the AWS Applied AI Solutions organization, we have a vision to provide business applications, leveraging Amazon's unique experience and expertise, that are used by millions of companies ... Show more

Project Engineer

PCL ConstructionBellevue, WA, US
Full-time

The future you want is within reach.At PCL Construction Services, Inc.PCL Family of Companies (PCL), we don’t just build projects—we build opportunities, careers and communities.We are 100% employe... Show more

People also ask
Cloud & Customer Solutions Engineer - DC GPU

Cloud & Customer Solutions Engineer - DC GPU

Advanced Micro Devices, IncBellevue, Washington, United States
9 days ago
Job type
  • Full-time
  • Permanent
Job description


ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.




THE TEAM:

AMD's Data Center GPU organization is transforming the industry with our AI based Graphic Processors. Our primary objective is to design exceptional products that drive the evolution of computing experiences, serving as the cornerstone for enterprise Data Centers, (AI) Artificial Intelligence, HPC and Embedded systems. If this resonates with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people.

THE ROLE:

As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers — frontier labs, NeoCloud providers, CSPs, and AI-native companies — to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets.

To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production outcomes, and be the engineer in the room when things break at scale. What you learn in the field, you convert into durable improvements — to ROCm, to the open-source serving ecosystem, and to the reference architectures every subsequent deployment inherits.

THE PERSON:

You are a strong production engineer who is energized rather than drained by ambiguity, customer pressure, and environments you do not control. You can debug a distributed training hang at 2am, explain the root cause to a customer VP at 9am, and land the fix upstream by the end of the week. You measure success by customer production outcomes, not code merged or tickets closed. When something is broken on a cluster you touch, it is your problem until it is fixed or explicitly handed off.

KEY RESPONSIBILITIES:

  • Own customer deployments end-to-end: cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets
  • Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer-specific workloads and SLOs across cloud, NeoCloud, and bare-metal environments
  • Lead root-cause analysis and resolution of production incidents on customer clusters, including Sev-1 response, and drive fixes to permanent closure
  • Deploy agentic AI solutions into customer environments in partnership with Agentic Data Engineers, and own their production behavior within the engagement
  • Build the observability, benchmarking, and validation tooling needed to certify clusters as production-ready and keep them there
  • Transfer operational capability to customer teams — documentation, runbooks, and hands-on enablement — moving customers up the operator-autonomy ladder from assisted operation to independent production ownership
  • Contribute field learnings to the Applied AI team's skills library and engagement memory databases, so deployment knowledge compounds across the practice
  • Convert field findings into upstream contributions — ROCm issues and patches, serving-framework improvements, reference-architecture updates — and provide structured field signal to AMD product, software, and silicon teams

PREFERRED EXPERIENCE:

  • 5+ years of production software or infrastructure engineering, including significant time operating or deploying systems in environments you did not build (level flexible for exceptional candidates)
  • Hands-on experience with GPU compute at scale: cluster deployment, distributed training or high-throughput inference, performance debugging, and workload optimization
  • Strong working knowledge of the modern AI infrastructure stack: Kubernetes and/or Slurm, containerized GPU workloads, collective communication libraries (RCCL/NCCL), high-performance networking (RoCE/InfiniBand), and observability tooling (Prometheus, Grafana)
  • Cloud platform depth (AWS, Azure, GCP, or NeoCloud environments), including hybrid and bare-metal deployment patterns
  • Proficiency in Python and at least one systems language; comfort navigating and modifying large codebases you did not write
  • Working familiarity with LLM application patterns — inference serving, RAG, and agentic workflows — sufficient to deploy and troubleshoot them in customer environments
  • Direct customer-facing experience: embedded deployments, technical escalations, on-site engagements, or equivalent
  • Experience with ROCm and AMD Instinct GPUs strongly preferred; deep CUDA-ecosystem experience with demonstrated ability to work cross-platform also valued
  • Open-source contribution history in AI/ML infrastructure projects is a plus

PREFERRED ACADEMIC CREDENTIALS:

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience

This role is not eligible for visa sponsorship.



#LI-RW1

#LI-HYBRID




Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

THE TEAM:

AMD's Data Center GPU organization is transforming the industry with our AI based Graphic Processors. Our primary objective is to design exceptional products that drive the evolution of computing experiences, serving as the cornerstone for enterprise Data Centers, (AI) Artificial Intelligence, HPC and Embedded systems. If this resonates with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people.

THE ROLE:

As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers — frontier labs, NeoCloud providers, CSPs, and AI-native companies — to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets.

To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production outcomes, and be the engineer in the room when things break at scale. What you learn in the field, you convert into durable improvements — to ROCm, to the open-source serving ecosystem, and to the reference architectures every subsequent deployment inherits.

THE PERSON:

You are a strong production engineer who is energized rather than drained by ambiguity, customer pressure, and environments you do not control. You can debug a distributed training hang at 2am, explain the root cause to a customer VP at 9am, and land the fix upstream by the end of the week. You measure success by customer production outcomes, not code merged or tickets closed. When something is broken on a cluster you touch, it is your problem until it is fixed or explicitly handed off.

KEY RESPONSIBILITIES:

  • Own customer deployments end-to-end: cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets
  • Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer-specific workloads and SLOs across cloud, NeoCloud, and bare-metal environments
  • Lead root-cause analysis and resolution of production incidents on customer clusters, including Sev-1 response, and drive fixes to permanent closure
  • Deploy agentic AI solutions into customer environments in partnership with Agentic Data Engineers, and own their production behavior within the engagement
  • Build the observability, benchmarking, and validation tooling needed to certify clusters as production-ready and keep them there
  • Transfer operational capability to customer teams — documentation, runbooks, and hands-on enablement — moving customers up the operator-autonomy ladder from assisted operation to independent production ownership
  • Contribute field learnings to the Applied AI team's skills library and engagement memory databases, so deployment knowledge compounds across the practice
  • Convert field findings into upstream contributions — ROCm issues and patches, serving-framework improvements, reference-architecture updates — and provide structured field signal to AMD product, software, and silicon teams

PREFERRED EXPERIENCE:

  • 5+ years of production software or infrastructure engineering, including significant time operating or deploying systems in environments you did not build (level flexible for exceptional candidates)
  • Hands-on experience with GPU compute at scale: cluster deployment, distributed training or high-throughput inference, performance debugging, and workload optimization
  • Strong working knowledge of the modern AI infrastructure stack: Kubernetes and/or Slurm, containerized GPU workloads, collective communication libraries (RCCL/NCCL), high-performance networking (RoCE/InfiniBand), and observability tooling (Prometheus, Grafana)
  • Cloud platform depth (AWS, Azure, GCP, or NeoCloud environments), including hybrid and bare-metal deployment patterns
  • Proficiency in Python and at least one systems language; comfort navigating and modifying large codebases you did not write
  • Working familiarity with LLM application patterns — inference serving, RAG, and agentic workflows — sufficient to deploy and troubleshoot them in customer environments
  • Direct customer-facing experience: embedded deployments, technical escalations, on-site engagements, or equivalent
  • Experience with ROCm and AMD Instinct GPUs strongly preferred; deep CUDA-ecosystem experience with demonstrated ability to work cross-platform also valued
  • Open-source contribution history in AI/ML infrastructure projects is a plus

PREFERRED ACADEMIC CREDENTIALS:

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience

This role is not eligible for visa sponsorship.



#LI-RW1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.