Job Description:
We are seeking a Senior DevOps Engineer to join their team on a remote-friendly contract basis. This is a hands-on role for an engineer with strong experience in AWS, Kubernetes, Terraform, and CI/CD who can independently drive reliable infrastructure decisions in ambiguous environments. You will be instrumental in automating and operating our cloud and Kubernetes platform, improving deployment flows, strengthening security and observability, and ensuring systems are scalable, reliable, and maintainable. This role also expects practical fluency with modern AI-assisted engineering workflows to accelerate analysis, troubleshooting, documentation, and delivery while maintaining high standards for correctness, security, and operational safety.
Responsibilities:
- Design, implement, and operate cloud infrastructure and automation using AWS and Terraform across services such as EC2, S3, EKS, ECR, Route 53, IAM, and core networking components.
- Own EKS and Kubernetes platform operations, including cluster lifecycle management, upgrades, node group strategy, autoscaling, workload scheduling, capacity planning, performance tuning, and cost optimization.
- Build and improve core Kubernetes platform capabilities such as networking, ingress, DNS integration, service discovery, RBAC, policy enforcement, secrets management, and multi-environment consistency.
- Develop and maintain deployment workflows using Helm, Kustomize, GitLab CI/CD, and GitOps approaches with tools such as Argo CD or Flux to improve release speed, consistency, traceability, and rollback safety.
- Deploy and evolve observability capabilities across infrastructure and Kubernetes environments, including metrics, logging, alerting, dashboards, and incident diagnostics.
- Support and optimize platform and application workloads on Kubernetes, improving deployment patterns, scaling behavior, runtime efficiency, resilience, and day-to-day operational support.
- Strengthen security across infrastructure, pipelines, and Kubernetes environments through IAM least-privilege access, secrets management, image and artifact scanning, admission controls, policy-as-code, workload isolation, and practical runtime hardening.
- Support the reliability, backup integrity, availability, and operational performance of PostgreSQL/RDS environments in partnership with application teams.
- Work closely with development and platform teams to improve deployment strategies, runtime reliability, developer experience, and operational standards for cloud-native systems.
- Monitor system health, investigate incidents, troubleshoot infrastructure and application issues, and drive timely resolution through strong root cause analysis and preventative improvements.
- Lead technical decision-making within the scope of the role by prioritizing work, evaluating tradeoffs, integrating inputs from multiple stakeholders, and driving issues through to completion with limited oversight.
- Use modern AI tools to accelerate infrastructure design exploration, Terraform authoring, Kubernetes troubleshooting, CI/CD workflow drafting, observability analysis, and operational documentation, while rigorously validating outputs for correctness, security, maintainability, and production readiness.
- Document and continually refine DevOps methodologies, infrastructure standards, deployment workflows, operational procedures, and support runbooks.