Job Title: Performance Engineer
Duration (Contract): 12 Months
Client Location: Dallas, TX
Location Preference: Hybrid
Job Description:
As a Performance Engineer, you will lead performance, scalability, resiliency, and capacity engineering initiatives for a high-volume Java-based platform. This individual contributor role focuses on identifying and resolving performance bottlenecks, analyzing production workloads, validating system resilience, and improving application reliability across distributed environments. You will work closely with development, architecture, infrastructure, and operations teams to evaluate application behavior, analyze JVM performance, design workload models, and recommend improvements across application code, databases, cloud infrastructure, networking, and platform configurations. The ideal candidate will possess deep Java expertise, strong diagnostic capabilities, and extensive experience in performance and resilience engineering.
Key Responsibilities:
- Independently plan and execute performance, scalability, resiliency, and capacity validation initiatives.
- Analyze application architecture, business workflows, dependencies, workloads, and production traffic patterns.
- Define performance requirements, workload models, acceptance criteria, service-level objectives, and reporting standards.
- Develop performance test scripts, workload generators, mock services, automation frameworks, and testing utilities.
- Execute load, stress, endurance, scalability, failover, recovery, degradation, and resiliency testing activities.
- Analyze source code, JVM behavior, thread activity, memory usage, garbage collection, latency, and system performance metrics.
- Investigate performance bottlenecks across applications, databases, infrastructure, networks, and distributed services.
- Recommend and validate improvements related to code optimization, system configuration, infrastructure design, capacity planning, and reliability engineering.
- Evaluate resiliency solutions including circuit breakers, traffic shaping, load balancing, and failover configurations.
- Utilize production metrics and workload patterns to support capacity planning, risk assessment, and performance validation.
- Develop and maintain dashboards, diagnostic tools, workload libraries, automation frameworks, reports, and operational runbooks.
- Assess and implement new performance testing, observability, profiling, and resilience engineering tools.
- Leverage approved automation and AI-assisted technologies to improve testing, diagnostics, analysis, and documentation processes.
- Communicate testing results, performance risks, recommendations, and remediation strategies to stakeholders.
Required Skills, Experiences, Education, and Competencies:
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
- 5+ years of experience in performance engineering, scalability engineering, software engineering, or production performance analysis.
- Strong expertise in performance testing, scalability testing, failover testing, disaster recovery validation, and resiliency engineering.
- Minimum 2 years of hands-on Java development experience.
- Deep understanding of Core Java, including JVM internals, concurrency, multithreading, memory management, garbage collection, thread management, and latency optimization.
- Ability to read and analyze source code to identify and resolve performance bottlenecks.
- Experience troubleshooting JVM issues using thread dumps, heap dumps, garbage collection logs, profilers, monitoring tools, and application metrics.
- Experience building workload generators, test frameworks, mock services, and performance automation using Java, Scala, Python, Shell Scripting, JMeter, Gatling, or similar technologies.
- Strong knowledge of Linux operating systems, databases, application servers, caching technologies, connection pools, networking, and distributed systems.
- Experience defining performance baselines, thresholds, acceptance criteria, capacity metrics, SLAs, SLOs, and KPIs.
- Proven ability to provide data-driven recommendations for performance optimization, system reliability, infrastructure scaling, and architectural improvements.
- Hands-on experience using monitoring and observability tools such as Splunk, Grafana, or similar platforms.
- Experience validating resiliency controls including Circuit Breakers, Traffic Shapers, Load Balancers, and Failover mechanisms.
- Strong analytical, troubleshooting, diagnostic, and problem-solving skills.
- Ability to work effectively with distributed teams and manage multiple priorities in fast-paced environments.
- Strong written and verbal communication skills with the ability to present technical findings to diverse audiences.
Preferred Skills, Experiences, Education, and Competencies:
- Experience supporting real-time, high-volume, transaction-intensive, or financial services platforms.
- Experience with Google Cloud Platform (GCP), Kubernetes, Docker, or other cloud-native technologies.
- Familiarity with AI-assisted analysis, intelligent automation, and Large Language Model (LLM)-based engineering tools.
- Experience using AI-powered development and analysis tools within Visual Studio, IntelliJ IDEA, or similar development environments.
- Experience with performance engineering for cloud-native, microservices, and distributed architectures.
- Knowledge of observability best practices, reliability engineering principles, and production support methodologies.
The hourly range for roles of this nature are $40.00 to $80.00/hr. Rates are heavily dependent on skills, experience, location, and industry.
cyberThink is an Equal Opportunity Employer.