Talent.com

Java software developer Jobs in Santa Clara, CA

Create a job alert for this search

Java software developer • santa clara ca

Last updated: 7 hours ago

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, IncSan Jose, California, United States
Full-time

At AMD, we believe technology has the power to solve the world’s most important challenges.From advancing healthcare and scientific discovery to powering AI and the technologies people rely on ever... Show more

Software Engineer

TradeJobsWorkforce95153 San Jose, CA, US
Full-time

Software Engineer Job Duties: Develops information systems by designing, developing, and instal... Show more

 • Promoted

Backend Java Developer - Remote

YO AI LabsSan Jose, CA, United States
Remote
Full-time
Quick Apply

In this role, you will apply your backend development expertise to design, develop, review, and optimize scalable systems using.Java, Spring Boot, and microservices architecture.No prior AI experie... Show more

Sr Full stack Java Developer

Hudson ManpowerSan Jose, CA, US
Full-time

Designed and developed scalable.AI-powered full-stack applications.Java, Spring Boot, React/Angular, REST APIs, and cloud-native technologies.Generative AI and Large Language Models (LLMs).Built AI... Show more

Staff Software Engineer/Mern Stack developer

Tek Leaders IncSunnyvale, CA, California, USA
Full-time

Staff Software Engineer/Mern Stack developer</p> <p><span style="font-size:12pt"><span style="font-family:Aptos">5 days onsite in Sunnyvale, CA</span&... Show more

Software Engineer Manager, Switchstack Software

GoogleSunnyvale, CA, United States
Full-time

Like Google's own ambitions, the work of a Software Engineer goes beyond just Search.Software Engineering Managers have not only the technical expertise to take on and provide technical leadership ... Show more

 • New!

Software Manager

DeepSight TechnologySanta Clara, CA, USA
$180,000.00 yearly
Full-time
Quick Apply

Physical AI -scaling advanced multimodal sensing technologies so clinicians can see more, sense more, and guide procedures with greater precision and confidence.Our vision starts with developing a ... Show more

Software Engineer

ServiceNowSanta Clara, California, United States
$125,700.00–$194,800.00 yearly
Full-time

Build Agent is ServiceNow's AI coding assistant, purpose-built for the platform's metadata-driven substrate - operating natively across ServiceNow's scoped applications, tables, and metadata types.... Show more

Managing Staff Software Developer

Intuitive SurgicalSunnyvale, CA, United States
$190,300.00–$273,800.00 yearly
Full-time

It started with a simple idea: what if surgery could be less invasive and recovery less painful? Nearly 30 years later, that question still fuels everything we do at.We’re a team of engineers, clin... Show more

Java Full Stack Developer1016

Veracity Ventures Inc.San Jose, CA, US
Full-time

Java, Spring Boot, and Microservices.JavaScript/TypeScript, HTML5, and CSS3.Git, CI/CD, Docker, and Agile/Scrum.Strong communication and problem-solving skills.Develop and maintain full-stack appli... Show more

Java Developer – Windchill

Bright Vision TechnologiesMountain View, CA, US
$80,000.00–$110,000.00 yearly
Full-time
Quick Apply

Java Developer – Windchill – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across th... Show more

Java Architect

TradeJobsWorkForce95190 San Jose, CA, US
Full-time

Java Architect Job Duties: Achieves e-commerce information architecture operational obj... Show more

 • Promoted

Software Engineer

DataVisorMountain View, CA, US
Full-time
Quick Apply

DataVisor is the world’s leading AI-powered Fraud and Risk Platform that delivers the best overall detection coverage in the industry.With an open SaaS platform that supports easy consolidation and... Show more

Senior Java Software Engineer, 12-Month Contract

Next Step Systems – Recruiters for Information Technology Jobs Top IT Recruiting FirmSan Jose, CA, USA
Full-time +1

Senior Java Software Engineer, 12-Month Contract, San Jose, CA.We currently have a 12-month contract opportunity available for a Senior Java Software Engineer.You will work directly with a team of ... Show more

Java Developer

Eitacies IncSanta Clara, CA, US
$60.00 hourly
Full-time

Santa Clara, CA (Onsite; hybrid is possible, but no remote work).Identify business needs by establishing personal rapport with actual, potential, and internal clients.Design, develop, and implement... Show more

Java Developer, Hybrid - 69815

PRIMUS Global ServicesSunnyvale, CA, United States
Full-time
Quick Apply

Java Developer, Hybrid </b></p> <p>We have an immediate need for an experienced <b>Java Developer </b>with strong expertise in <b>Java, React JS, NoSQL database... Show more

Principal Software Developer – AI/ML Performance Validation & Systems Testing

AMDSan Jose, CA, US
Full-time

At AMD, we believe technology has the power to solve the world's most important challenges.From advancing healthcare and scientific discovery to powering AI and the technologies people rely on ever... Show more

Software Engineer

Sunrise SystemsSanta Clara, California, United States
Full-time
Quick Apply

Location: Onsite Candidates may be based in either the San Francisco Bay Area or the Des Moines Metro Area,.We are seeking a highly technical and self-directed Senior Software Engineer to contrib... Show more

Embedded Software Developer for RDK-B in Santa Clara, CA.(Onsite Position)

Aita Consulting Services, IncSanta Clara, CA, United States
Full-time
Quick Apply

Hi,</p> <p> </p> <p>One of our clients is looking for <b>Embedded Software Developer for RDK-B in Santa Clara, CA.Onsite Position)</b></p> <p> &lt... Show more

Sr. Software Developer: 25-05693 (No C2C)

Akraya IncSunnyvale, California, United States
$70.00–$75.00 hourly
Full-time
Quick Apply

Primary Skills: Java (Expert), MySQL (Expert), Distributed Systems Architecture (Expert),iOS (Proficient), Android (Proficient), Data Structure and Algorithms (Expert),.Duration: 6 Months with poss... Show more

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Principal Software Developer – AI/ML Performance Validation & Systems Testing

Advanced Micro Devices, IncSan Jose, California, United States
12 days ago
Job type
  • Full-time
Job description


ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career.




THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID




Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.

THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship — from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm — unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers — across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure — distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering — shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows — PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug — partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap — work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally — strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes — multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization — LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks — establishing reproducible methodology, baselines, and regression tracking.

PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:GPU software stacks (ROCm, CUDA, oneAPI, SYCL)AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)Linux kernel, GPU drivers, or accelerator firmwareDistributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.

ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).

LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

This posting is for an existing vacancy.