Platform Engineer
Unisys is looking for Platform Engineer in Rockville, MD.
This local job opportunity with ID 3877201536 is live since 2026-10-08 20:40:15.
Need a Kubernetes SME / Kubernetes Architect / Kubernetes Engineer with
- Deep Kubernetes experience (Troubleshooting, configuring for scale), doing this at scale involves working with larger data sets (aka Big Data) and they use Spark on Kubernetes to run their processes.
- Spark: Spark would be 50% of the work.
- SQL querying will assess on this as well.
For a complete understanding of this opportunity, and what will be required to be a successful applicant, read on.
Senior Kubernetes Engineer Overview:
- We are seeking an expert level Kubernetes Engineer to configure, deploy, and operate mission-critical Amazon EKS infrastructure supporting petabyte-scale data processing workloads.
- This role requires deep technical expertise in Kubernetes internals, large-scale cluster management, and Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata.
Core Responsibilities:
- Design, deploy, and maintain production-grade Amazon EKS clusters architected for petabyte-scale data processing with high availability and fault tolerance.
- Operate in air-gapped private VPC environments without internet access, managing secure package repositories and container registries.
- Debug and resolve complex distributed systems issues across EKS, Karpenter, and Spark, including scheduling bottlenecks, node scaling delays, and cascading failures under heavy load.
- Implement Karpenter consolidation and disruption policies balancing cost optimization with job resiliency.
- Manage spot and on-demand instance strategies with robust node interruption handling.
- Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource distribution and prevent resource starvation.
- Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access.
- Optimize storage configurations for resiliency, high-throughput data processing workloads.
- Implement logging, alerting, and anomaly detection to identify provisioning failures, executor loss, and throughput degradation before broader system impact.
- Design fault-tolerant architectures with retry strategies, checkpointing, and graceful degradation patterns minimizing re-computation on failure.
- Develop and manage configs in a private VPC without internet access.
Preferred Qualifications:
- Kubernetes Certifications (CKA, CKAD, or CKS) and AWS Certifications.
- Active contributions to open-source Kubernetes, Karpenter or Spark projects.
- Familiarity with FinOps practices and cost optimization at scale. xevrcyc
- Background in data engineering or analytics platforms.
#LI-CGTS
#TS-3142