Senior Kubernetes Engineer

Location: DMV area, NYC, NJ - Hybrid

The Opportunity

Strategio is supporting an enterprise client in building and operating mission-critical Amazon EKS infrastructure designed to power petabyte-scale data processing workloads. This is an opportunity to work on highly complex distributed systems where Kubernetes performance, scalability, resiliency, and infrastructure design are critical to supporting thousands of concurrent data processing jobs.

You will take ownership of large-scale Kubernetes environments operating within secure, air-gapped private VPCs and solve complex engineering challenges across Amazon EKS, Karpenter, and Apache Spark. This role is ideal for an experienced Kubernetes engineer who enjoys working deep within infrastructure, troubleshooting distributed systems under heavy load, and designing highly available platforms at significant scale.


About You


You are a highly experienced Kubernetes engineer with deep technical expertise in Kubernetes internals, Amazon EKS, and large-scale distributed systems. You have operated complex production environments where availability, performance, scalability, and fault tolerance are critical.

You are comfortable troubleshooting infrastructure issues beyond the surface level, including scheduling bottlenecks, node scaling delays, resource contention, executor failures, and cascading failures under heavy workloads. You combine strong Kubernetes engineering expertise with an understanding of large-scale data processing environments and Apache Spark workloads.



As a Senior Kubernetes Engineer, You will:

  • Design, deploy, and operate production-grade Amazon EKS clusters supporting petabyte-scale data processing workloads.
  • Architect highly available and fault-tolerant Kubernetes environments capable of supporting thousands of concurrent jobs.
  • Operate infrastructure within air-gapped private VPC environments without internet access, including the management of secure package repositories and container registries.
  • Investigate and resolve complex distributed systems issues across Amazon EKS, Karpenter, and Apache Spark.
  • Troubleshoot scheduling bottlenecks, node scaling delays, resource contention, and cascading failures under heavy workloads.
  • Implement and optimize Karpenter consolidation and disruption policies while balancing infrastructure costs with workload resiliency.
  • Develop and manage effective Spot and On-Demand instance strategies, including robust node interruption handling.
  • Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource allocation and prevent resource starvation.
  • Configure persistent volume claims and container-native storage using the Amazon EBS CSI and Amazon EFS CSI drivers.
  • Optimize storage architecture for resilient, high-throughput data processing workloads.
  • Implement comprehensive logging, monitoring, alerting, and anomaly detection to proactively identify provisioning failures, executor loss, and throughput degradation.
  • Design fault-tolerant architectures using retry strategies, checkpointing, and graceful degradation patterns to minimize re-computation following system failures.
  • Develop, manage, and maintain infrastructure configurations within highly secure private cloud environments.
  • Continuously improve platform scalability, reliability, performance, and operational efficiency.



Core Skills

  • Extensive hands-on experience designing, deploying, and operating production Kubernetes environments.
  • Deep expertise with Amazon EKS and AWS cloud infrastructure.
  • Strong understanding of Kubernetes internals, architecture, scheduling, resource management, and cluster operations.
  • Experience managing large-scale Kubernetes clusters supporting high-volume or data-intensive workloads.
  • Hands-on experience with Karpenter, including node provisioning, consolidation, disruption policies, and scaling strategies.
  • Experience supporting and optimizing Apache Spark workloads running on Kubernetes.
  • Strong understanding of distributed systems and the ability to troubleshoot complex performance and reliability issues.
  • Experience designing highly available and fault-tolerant infrastructure.
  • Strong knowledge of Kubernetes resource management, including ResourceQuotas, LimitRanges, and PriorityClasses.
  • Experience implementing Spot and On-Demand instance strategies and handling node interruptions.
  • Hands-on experience with Kubernetes storage architecture, persistent volumes, Amazon EBS CSI, and Amazon EFS CSI drivers.
  • Experience implementing infrastructure monitoring, logging, alerting, and anomaly detection.
  • Experience operating within secure private VPC or air-gapped environments.
  • Strong troubleshooting and root-cause analysis capabilities.



Nice-to-haves:

  • Certified Kubernetes Administrator (CKA), Certified Kubernetes Application Developer (CKAD), or Certified Kubernetes Security Specialist (CKS) certification.
  • AWS certifications.
  • Active contributions to open-source Kubernetes, Karpenter, or Apache Spark projects.
  • Experience implementing FinOps practices and optimizing cloud infrastructure costs at scale.
  • Previous experience working with large-scale data engineering or analytics platforms.


Your job search ends here.

If you’re eager to contribute to cutting-edge projects and want to work with a high-performance team, apply now.