Skip to main content
Version: main 🚧

Run CAIPE on Amazon EKS

This guide walks you through creating an Amazon EKS (Elastic Kubernetes Service) cluster and deploying CAIPE (Community AI Platform Engineering) on it. No prior experience with CAIPE or EKS is required.

This guide opts in to EKS Auto Mode for a new cluster. Existing EKS clusters and Helm installations can keep their current cluster config and chart values. Do not apply the Auto Mode storage class or scheduling overlay to an existing cluster without planning volume migration. The original managed-node-group example remains at deploy/eks/dev-eks-cluster-config.yaml.example.

What is EKS? EKS is AWS’s managed Kubernetes service. You get a production-ready cluster without managing control-plane nodes yourself. eksctl is a simple CLI to create and manage EKS clusters with sensible defaults.

What you’ll do: Create an EKS cluster, install ArgoCD (optional, for GitOps-style deploys), then deploy CAIPE using the Helm chart. You’ll need an AWS account and the tools listed below.


Step 1: Clone the repository​

You need the repo to use the EKS cluster configuration example and to follow the same paths as this guide.

git clone https://github.com/caipe-io/ai-platform-engineering.git
cd ai-platform-engineering

The EKS config example lives under deploy/eks/. We’ll use it in a later step.


Step 2: Prerequisites​

Install and configure these before creating the cluster:

ToolPurpose
AWS CLIAuthenticate to AWS and run commands (install)
eksctlCreate and manage EKS clusters (install)
kubectlTalk to your Kubernetes cluster (install)
HelmInstall CAIPE and add-ons (install)

AWS account: Your user or role needs permissions for EC2, EKS, CloudFormation, and IAM (for cluster and node creation). See Required AWS permissions below.


Step 3: Configure AWS credentials​

Log in to AWS and confirm your identity:

# Configure AWS CLI (you’ll be prompted for Access Key ID and Secret)
aws configure

# Confirm credentials work
aws sts get-caller-identity

# Optional: set a default region
export AWS_DEFAULT_REGION=us-east-2

Use the same region in the next step when you create the cluster.


Step 4: Create the EKS cluster​

The repo includes a cluster config using EKS Auto Mode, which manages node provisioning via the built-in Karpenter controller. No additional autoscaler setup is required.

Prepare the cluster config​

Start from the example file so the KMS ARN placeholder is fresh. Run this each time you create the cluster, so it is safe to re-run after a previous cluster has been torn down:

# Always start from the example so the ARN placeholder is fresh
cp deploy/eks/dev-eks-auto-mode-cluster-config.yaml.example dev-eks-cluster-config.yaml

Required: update publicAccessCIDRs in dev-eks-cluster-config.yaml to your VPN or office egress CIDR before continuing. It ships with a non-routable placeholder (203.0.113.0/24) that blocks public API access until you replace it, so the example fails closed rather than exposing the control plane.

Create a KMS key for secrets encryption​

The cluster config encrypts Kubernetes secrets at rest using a customer-managed KMS key. Run this block in full each time you create the cluster. It creates a new key and writes its ARN into the config, so it is safe to re-run after a previous cluster has been torn down:

# Create a new key and capture its ARN
KEY_ARN=$(aws kms create-key \
--description "dev-eks-cluster secrets encryption" \
--query KeyMetadata.Arn --output text)

# Recreate the alias (delete first in case it exists from a previous run)
aws kms delete-alias --alias-name alias/dev-eks-cluster-secrets 2>/dev/null || true
aws kms create-alias \
--alias-name alias/dev-eks-cluster-secrets \
--target-key-id "$KEY_ARN"

# Write the new ARN into the config
sed -i.bak "s|arn:aws:kms:us-east-2:ACCOUNT_ID:key/KEY_ID|$KEY_ARN|" dev-eks-cluster-config.yaml
rm dev-eks-cluster-config.yaml.bak

Run eksctl​

Create the cluster. This usually takes 10–15 minutes:

eksctl create cluster -f dev-eks-cluster-config.yaml

If you see a "CloudFormation stack already exists" error, see Troubleshooting.

eksctl will:

  • Create a VPC and subnets
  • Set up the EKS control plane
  • Run aws eks update-kubeconfig --region us-east-2 --name dev-eks-cluster to configure your kubectl context to use the new cluster

Verify the cluster​

# List EKS clusters
eksctl get cluster

# Check that nodes are ready
kubectl get nodes

# Cluster and API server info
kubectl cluster-info

# Optional: list add-ons and system pods
eksctl get addons --cluster dev-eks-cluster
kubectl get pods -n kube-system

Once kubectl get nodes shows nodes in Ready state, continue to the next step.

Grant additional IAM access​

The cluster config uses authenticationMode: API, meaning cluster access is managed via IAM access entries, not the aws-auth ConfigMap. To grant another IAM user or role admin access, use eksctl create accessentry:

AWS_ACCOUNT_ID=$(aws sts get-caller-identity --query Account --output text)

# Grant an IAM user cluster-admin access
IAM_USER=USERNAME
eksctl create accessentry \
--cluster dev-eks-cluster \
--principal-arn "arn:aws:iam::${AWS_ACCOUNT_ID}:user/${IAM_USER}" \
--kubernetes-groups system:masters

# Grant an IAM role cluster-admin access (e.g. a CI/CD role)
IAM_ROLE=ROLE_NAME
eksctl create accessentry \
--cluster dev-eks-cluster \
--principal-arn "arn:aws:iam::${AWS_ACCOUNT_ID}:role/${IAM_ROLE}" \
--kubernetes-groups system:masters

Do not edit the aws-auth ConfigMap directly, as it has no effect in API auth mode.


Step 5: Create the Auto Mode StorageClass​

EKS Auto Mode ships the EBS CSI driver (ebs.csi.eks.amazonaws.com) but, by AWS design, creates no StorageClass. Apply the repo's gp3 class once, right after the new cluster is up. The manifest does not change a cluster's default StorageClass:

kubectl apply -f deploy/eks/storage/

On a new Auto Mode cluster with no existing default StorageClass, explicitly make auto-ebs-sc the default so chart PVCs without a storageClassName can bind. On an existing cluster, keep its current default and plan a separate volume migration before changing it:

kubectl annotate storageclass auto-ebs-sc \
storageclass.kubernetes.io/is-default-class=true --overwrite

Verify that (default) appears next to auto-ebs-sc:

kubectl get storageclass

By default, Auto Mode places all workloads on its built-in general-purpose pool. The memory-bound RAG stack benefits from a dedicated tier, so apply the rag NodePool to give it on-demand memory-optimised nodes that scale to zero when idle. Everything else stays on the Auto Mode general-purpose pool.

NodePoolWorkloadsInstance strategy
ragrag-server, agent-ontology, rag-ingestors, rag-redis, neo4j, milvus (+ its etcd/minio)On-demand, memory-optimised (r5/r6i)
general-purpose (built-in)Dynamic Agents, MCP servers (mcp-*), UI, Keycloak, OpenFGA, …Auto Mode managed
kubectl apply -f deploy/eks/karpenter/

Verify the NodePools are created and the built-in Auto Mode pools are present:

kubectl get nodepool

When deploying CAIPE in the next step, append -f charts/ai-platform-engineering/values-karpenter.yaml to the Helm install command to route workloads to the correct node tier.


Step 6: Deploy CAIPE on EKS​

You have two main options:

Option A: Install CAIPE with Helm​

Install the CAIPE Helm chart directly on the cluster. Configure secrets and LLM settings as described in the Helm guide.

helm upgrade --install ai-platform-engineering oci://ghcr.io/caipe-io/charts/ai-platform-engineering \
--version 1.1.0 \
--namespace ai-platform-engineering \
--create-namespace \
--set-string tags.caipe-ui=true \
--set-string tags.dynamic-agents=true \
--set-string tags.mcp-netutils=true

Then:

Full details: Deploy CAIPE with Helm.

Option B: Use ArgoCD, then deploy CAIPE​

ArgoCD keeps your cluster in sync with Git (or Helm) and is useful for ongoing updates. You can install ArgoCD first, then deploy the CAIPE chart through ArgoCD or Helm.

Install ArgoCD on the cluster:

kubectl create namespace argocd
kubectl apply -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml

Access the ArgoCD UI (optional):

kubectl port-forward svc/argocd-server -n argocd 8080:443

Open http://localhost:8080. Then deploy CAIPE via the Helm chart (as in Option A) or by defining an ArgoCD Application that points at the same chart. See the Helm deployment guide for the chart values and verification steps.


Step 7: Configure load balancing​

EKS Auto Mode includes load balancing for Service and Ingress resources. Follow AWS's Network Load Balancer guide for LoadBalancer services or its Application Load Balancer guide for HTTP ingress. The Auto Mode ALB path uses an IngressClass with the eks.amazonaws.com/alb controller.


Required AWS permissions​

Your AWS user or role needs permissions for:

  • EC2 — Instances, VPC, subnets, security groups
  • EKS — Cluster and node group management
  • CloudFormation — Stacks created by eksctl
  • IAM — Roles and policies for the cluster and node groups

If something fails with “access denied”, check IAM permissions for eksctl and your organisation’s policies.


Troubleshooting​

Insufficient permissions​

aws iam get-user
aws iam list-attached-user-policies --user-name YOUR_USERNAME

Fix by attaching the required policies or using a role that has them.

Region mismatch​

Ensure the region in dev-eks-cluster-config.yaml matches your AWS CLI default:

aws configure get region

Nodes remain pending​

  • Inspect the built-in pools: kubectl get nodepool
  • Check unschedulable pod events: kubectl describe pod <pod> -n <namespace>
  • See Troubleshoot EKS Auto Mode for provisioning failures.

CloudFormation stack already exists​

A previous cluster creation attempt failed and left a partial stack behind. Delete it before retrying:

aws cloudformation delete-stack --stack-name eksctl-dev-eks-cluster-cluster
aws cloudformation wait stack-delete-complete --stack-name eksctl-dev-eks-cluster-cluster

TerminationProtection is enabled​

eksctl enables CloudFormation termination protection on successfully created clusters. Disable it before deleting:

aws cloudformation update-termination-protection \
--no-enable-termination-protection \
--stack-name eksctl-dev-eks-cluster-cluster

Then retry the delete command.

kubectl can’t reach the cluster​

# Refresh kubeconfig for your cluster (use your region and cluster name)
aws eks update-kubeconfig --region us-east-2 --name dev-eks-cluster

# Confirm current context
kubectl config current-context

Cleanup​

When you’re done, delete the cluster to avoid ongoing AWS charges. eksctl enables CloudFormation termination protection on successfully created clusters, so disable it first:

aws cloudformation update-termination-protection \
--no-enable-termination-protection \
--stack-name eksctl-dev-eks-cluster-cluster

eksctl delete cluster -f dev-eks-cluster-config.yaml

Verify that CloudFormation stacks are gone:

aws cloudformation list-stacks --query 'StackSummaries[?contains(StackName, `eksctl-dev-cluster`)].{Name:StackName,Status:StackStatus}'

Important: Always tear down the cluster when you’re not using it to prevent unexpected charges.

Clean up CloudWatch logs​

EKS does not delete the control plane log group when the cluster is deleted. Remove it manually:

aws logs delete-log-group --log-group-name /aws/eks/dev-eks-cluster/cluster

Clean up the KMS key​

KMS keys cannot be deleted immediately; They must be scheduled for deletion with a minimum 7-day waiting period. Delete the alias first, then schedule the key:

# Look up the key ARN via the alias
KEY_ARN=$(aws kms describe-key \
--key-id alias/dev-eks-cluster-secrets \
--query KeyMetadata.Arn --output text)

aws kms delete-alias --alias-name alias/dev-eks-cluster-secrets

aws kms schedule-key-deletion \
--key-id "$KEY_ARN" \
--pending-window-in-days 7

The key will be permanently deleted after the pending window. You can cancel before then with aws kms cancel-key-deletion --key-id "$KEY_ARN".


Next steps​