Best practices for upgrading EKS clusters with Karpenter - Zesty

Best practices for upgrading EKS clusters with Karpenter

By Omer Hamerman

Principal DevOps Engineer

Amazon Elastic Kubernetes Service (EKS) makes running Kubernetes a much smoother experience, but upgrading an EKS cluster can be anything but smooth. An upgrade touches both the control plane (the brain of your cluster) and the data plane (the worker nodes actually running workloads). Each component can break things if not handled properly.

Karpenter, the open-source cluster autoscaler, brings efficiency to node provisioning and scaling. But because it directly interacts with your workloads and AWS infrastructure, any EKS version change can affect how Karpenter behaves.

The reality: upgrading is unavoidable, but with the right plan you can do it without breaking production. In this guide, we’ll walk through the process end-to-end so you can upgrade smoothly and confidently.

Why you should upgrade your EKS cluster

Upgrading your cluster is not about chasing shiny features. It’s about avoiding risk and keeping your environment stable.

Delaying upgrades piles up technical debt. The further behind you fall, the riskier and more complex the eventual upgrade becomes.

Common issues during upgrades

Understanding the usual pain points will help you anticipate them.

Knowing these pitfalls upfront is the best way to avoid firefighting later. Next, let’s map out how to prepare properly.

Planning your upgrade

Preparation separates smooth upgrades from war stories.

Pre-upgrade checklist

Audit your cluster for deprecated APIs:

kubectl get --raw /metrics | grep deprecated

Staging environment

Backup strategy

Capacity planning

Once these basics are covered, you’re ready to start the upgrade with confidence.

Step by step: executing the upgrade with minimal disruption

Here’s the process that experienced practitioners follow to keep systems online.

Step 1: Upgrade the control plane

Verify the new control plane version:

aws eks describe-cluster --name my-cluster --query cluster.version

Step 2: Upgrade node groups

For self-managed nodes: do a rolling update. Drain each node gracefully:

kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data

Step 3: Handle Karpenter provisioners

Step 4: Upgrade networking and addons

Confirm pod networking works by deploying a simple busybox pod and testing DNS resolution:

kubectl run test-dns --image=busybox:1.28 --rm -it --restart=Never -- nslookup kubernetes.default

Tips for near-zero downtime

Once your nodes are upgraded and workloads stable, the next step is validation.

Post-upgrade validation and testing

Don’t assume the cluster is healthy just because the upgrade finished.

Cluster health checks

Confirm control plane components are ready:

kubectl get componentstatuses

Application smoke tests

Karpenter validation

Check logs for provisioning errors:

kubectl logs -n karpenter -l app.kubernetes.io/name=karpenter

Observability

This is the stage where small misconfigurations show up, so give it real attention.

Best practices and lessons learned to build a sustainable upgrade flow

After a few cycles, you’ll notice patterns. These practices will save you time and headaches:

Every upgrade is an opportunity to refine your process. Share lessons with your team so everyone benefits.

Upgrading without turning into a war story

Upgrading EKS clusters is a reality of operating Kubernetes in production. With careful planning, staged testing, and thorough validation, you can perform upgrades with minimal disruption.

Stay ahead of AWS deprecations, treat Karpenter as a first-class citizen in the upgrade, and keep your playbooks current. Done right, upgrades become routine rather than dreaded fire drills.