Sample course · Some prior knowledge · 12 lessons

Kubernetes for data engineers

Run batch pipelines, Airflow and Spark on Kubernetes, from the first pod to the cost report

A practical course on Kubernetes for people who already build data pipelines and know Docker. You follow one team as it moves a nightly batch load onto a cluster, then adds Airflow and Spark, learning pods, Jobs and CronJobs, Services, configuration, storage, resources, permissions and autoscaling on the way. Afterwards you can run, size, secure, debug and cost data workloads on EKS, GKE, AKS or your own cluster.

What you'll learn

  • Explain how a Kubernetes cluster works and use kubectl to inspect and change it
  • Choose between Deployments, Jobs and CronJobs and write a reliable scheduled batch job
  • Expose services and inject configuration and credentials safely with ConfigMaps, Secrets and workload identity
  • Pick the right storage for pipeline data, from emptyDir and PersistentVolumeClaims to object storage
  • Set resource requests and limits that keep batch jobs from failing or wasting capacity
  • Share a cluster between teams with namespaces, RBAC and quotas
  • Run Airflow and Spark on Kubernetes and scale them with pod and node autoscalers
  • Debug failing pods systematically and find where a cluster's money goes

Who it's for

  • Data engineers who package pipelines in Docker and are moving them onto Kubernetes
  • Analytics engineers and data platform developers asked to run Airflow or Spark on a shared cluster
  • Engineers who use kubectl occasionally and want to understand why their jobs are Pending or OOMKilled

Syllabus

  1. 1.Kubernetes from a data engineer's seat

    What a cluster is, the pod as the unit of work, and how batch jobs differ from long-running services.

    1. Why a data team moves to Kubernetes
    2. Pods: the smallest thing Kubernetes runs· checkpoint
    3. Deployments, Jobs and CronJobs
  2. 2.Wiring the pipeline up

    Giving pods stable addresses, injecting configuration and credentials, and choosing where data is stored.

    1. Services: stable addresses for changing pods· checkpoint
    2. ConfigMaps and Secrets
    3. Storage: volumes, claims and object storage· checkpoint
  3. 3.Sharing the cluster safely and efficiently

    Sizing workloads with requests and limits, dividing the cluster between teams, and letting capacity follow demand.

    1. Requests, limits and the scheduler
    2. Namespaces, RBAC and quotas· checkpoint
    3. Autoscaling pods and nodes
  4. 4.Running data platforms on Kubernetes

    Airflow and Spark on the cluster, then the day-to-day work of debugging, monitoring and controlling cost.

    1. Running Airflow on Kubernetes· checkpoint
    2. Running Spark on Kubernetes
    3. Debugging, monitoring and cost· checkpoint

Lesson 1

Why a data team moves to Kubernetes

What you'll learn: What Kubernetes actually does for a data team, how a cluster is put together, and why "declaring the state you want" changes how you run pipelines.

Meet the pipeline we will move

Throughout this course we follow the data team at Harbour Lane Grocers, a regional chain of 40 shops. The team lead, Tomás, looks after a nightly batch pipeline called sales-load. Every night at 02:00 a cron entry on a single virtual machine starts a Python script. It pulls the day's till files from object storage, cleans them, and loads about 3 million rows into the warehouse before the morning sales report runs at 07:00.

It already runs in a Docker container, which is good news. The trouble is everything around the container:

  • When the VM was patched and rebooted at 02:05 one night, the load simply never happened, and nobody knew until the report was empty.
  • The VM is sized for the busiest night of the year, so it sits idle about 22 hours a day.
  • The team now wants Airflow for scheduling and Spark for a heavier pricing model, and nobody wants three more hand-managed servers.

Kubernetes is the tool Tomás is evaluating. By the end of this course the sales-load job, a small API, Airflow and Spark will all run on one cluster, and you will understand every piece.

What Kubernetes is, in one paragraph

Kubernetes (often written K8s) is an open-source system that runs containers across a group of machines. You tell it what you want running, such as "one copy of sales-load every night at 02:00, with 2 GB of memory", and it decides which machine runs it, starts it, restarts it if it crashes, and reports what happened. It does not build your images or replace your warehouse. It is the layer that keeps containers running where and when you asked.

The parts of a cluster

A Kubernetes cluster has two halves.

The control plane is the brain. It stores the desired state of everything in a database called etcd, exposes an API server that every tool talks to, runs a scheduler that picks a machine for each new workload, and runs controllers that notice when reality drifts from what you asked for.

The nodes are the muscle: the virtual or physical machines that actually run containers. Each node runs an agent called the kubelet, which takes instructions from the control plane and starts containers through a container runtime (usually containerd).

PartLives onJob
API serverControl planeThe front door; every command goes through it
etcdControl planeStores the desired and current state
SchedulerControl planeChooses a node for each new pod
ControllersControl planeKeep nudging reality towards the desired state
kubeletEvery nodeStarts and watches containers on that node

Desired state, not instructions

The idea that makes Kubernetes feel different from a cron VM is that you describe an outcome rather than a sequence of steps. Think of a thermostat. You do not tell the heating "run for 20 minutes, then stop". You set 20 degrees and the thermostat keeps checking the room, switching the heating on and off until the room matches. If someone opens a window, it reacts without you.

Kubernetes controllers work the same way. Tomás will write a short YAML file saying a job should exist with a given image and schedule, then apply it with kubectl apply -f sales-load.yaml. The controllers compare that description with what is running, and act on the difference. If a node dies halfway through the load, the controller sees that the job has not finished and starts it again on a healthy node. That alone fixes the "VM rebooted at 02:05" incident.

Every YAML file starts with two key lines: apiVersion: batch/v1 names the API version, and kind: CronJob says what sort of object it is. Later lessons build these files up a line at a time.

Talking to the cluster with kubectl

kubectl is the command-line tool for the API server. A few commands cover most of a data engineer's day:

  1. kubectl get pods lists what is running.
  2. kubectl describe pod sales-load-abc12 explains one object in detail, including recent events.
  3. kubectl logs sales-load-abc12 shows a container's output.
  4. kubectl apply -f file.yaml creates or updates objects from a file.
  5. kubectl delete -f file.yaml removes them.

Your kubectl talks to one cluster at a time, chosen by a context in your kubeconfig file. kubectl config current-context tells you which one, and checking it before a delete is a habit worth forming early.

Build your own or rent one

Running the control plane yourself is real work: upgrades roughly three times a year, etcd backups, certificates. Most data teams rent it. The three big managed offerings are Amazon EKS, Google Kubernetes Engine (GKE) and Azure Kubernetes Service (AKS). Each runs the control plane for you and charges either a small hourly fee per cluster or nothing at all on a basic tier, plus the normal price of the nodes. GKE Autopilot and EKS Auto Mode go further and manage the nodes too. Prices and tiers change, so check your provider's current page rather than trusting a number in a course.

For learning, a local cluster is plenty. kind, minikube, k3d and the Kubernetes option built into Docker Desktop all give you a one-node cluster on a laptop. Tomás starts with kind so the team can break things freely.

Is it worth it for a data team?

Kubernetes adds moving parts. For one nightly script, a serverless job service may be simpler. It pays off when, as at Harbour Lane, you have several workloads with different shapes (nightly batch, always-on APIs, bursty Spark) that you want on shared, self-healing infrastructure with one way of deploying. That is the case we will build.

Recap

  • Kubernetes runs containers across many machines and keeps them in the state you declared.
  • A cluster has a control plane (API server, etcd, scheduler, controllers) and nodes running the kubelet.
  • You describe the outcome in YAML and apply it; controllers close the gap, like a thermostat.
  • kubectl talks to the API server; always know which context you are pointed at.
  • EKS, GKE and AKS run the control plane for you; kind or minikube are fine for learning.

Lesson 2

Pods: the smallest thing Kubernetes runs

What you'll learn: What a pod is, why Kubernetes runs pods rather than bare containers, and how to run and inspect the sales-load container as a pod.

From container to pod

Tomás's sales-load image already runs with docker run harbourlane/sales-load:1.4. Kubernetes never runs a container on its own, though. The smallest thing it schedules is a pod: a wrapper around one or more containers that are always placed on the same node and started together.

Most pods hold exactly one container, and that is what Harbour Lane's will look like. The wrapper still matters, because the pod is where Kubernetes attaches the things a container needs from the outside world: an IP address, storage volumes, configuration and a restart policy.

Containers in a pod share a home

When a pod does hold more than one container, those containers live together closely. Think of flatmates sharing one flat. They have separate bedrooms (each container has its own filesystem and process), but they share the front door and the kitchen: one network address, so they can reach each other on localhost, and any volumes the pod mounts.

That makes multi-container pods useful for helpers that must sit right next to the main process:

  • An init container runs to completion before the main container starts, for example to wait until the warehouse accepts connections or to download a reference file.
  • A sidecar runs alongside the main container for its whole life, for example a log shipper or a proxy that handles credentials. Recent Kubernetes versions support sidecars natively, declared as init containers with restartPolicy: Always.

If two processes could live on different machines without problems, they belong in different pods. Putting the warehouse loader and a dashboard in one pod just because they are related is a common beginner mistake.

A first pod, one line at a time

Pods are described in YAML. Tomás writes sales-load-pod.yaml, and the lines that matter are these:

  • kind: Pod says what is being created.
  • name: sales-load-test (under metadata) is the pod's name, unique within its namespace.
  • image: harbourlane/sales-load:1.4 (under spec.containers) is the container image.
  • args: ["--date", "2026-09-30"] passes arguments to the script, here a date to reload.
  • restartPolicy: Never tells the kubelet not to restart the container when it exits.

He applies it with kubectl apply -f sales-load-pod.yaml. For a quick experiment there is also a one-line shortcut, kubectl run sales-load-test --image=harbourlane/sales-load:1.4 --restart=Never, but files are what the team will keep in Git.

The life of a pod

Every pod moves through a small set of phases, and reading them is the first debugging skill you need.

PhaseWhat it meansWhat Tomás would check
PendingAccepted, but not yet runningIs there a node with room? Is the image still downloading?
RunningAt least one container is runningLogs, to see progress
SucceededAll containers exited with code 0Nothing; the load worked
FailedA container exited with a non-zero codeLogs and the exit code
UnknownThe control plane cannot reach the nodeThe node's health

kubectl get pods shows a STATUS column that is a little more detailed than the phase. You will meet values such as ContainerCreating, Completed, CrashLoopBackOff (the container keeps failing and Kubernetes is waiting longer between restarts) and ImagePullBackOff (it cannot download the image). We come back to all of these in the final lesson.

Looking inside a running pod

Tomás's test pod ran for four minutes and ended in Completed. Three commands told him what happened:

  1. kubectl logs sales-load-test printed the script's output, including "loaded 2,981,442 rows".
  2. kubectl describe pod sales-load-test showed which node it ran on, the image it used, and an Events list (scheduled, pulled image, created container, started container).
  3. kubectl get pod sales-load-test -o yaml printed the full object, including fields Kubernetes filled in for him.

While a pod is running, kubectl exec -it sales-load-test -- sh opens a shell inside the container. It is handy for checking whether a file landed, but treat it as a diagnostic, never as a way to fix things by hand.

Pods are disposable

This is the idea that takes longest to accept. A pod is not a small server you look after. It is closer to a single run. When a pod dies, Kubernetes does not repair it; a controller creates a new pod, with a new name and a new IP address, from the same description. Anything written to the container's own filesystem is lost with it.

For the data team that has two consequences. First, the output of sales-load must go somewhere durable (the warehouse and object storage), never only to local disk. Second, the load must be safe to run twice. Tomás changes the script so that it deletes the target day's rows before inserting them, inside one transaction. Now a retry after a crash produces the same result as a clean run. This property, called idempotency, will matter in every lesson from here on.

Why not just create pods?

A bare pod has no controller watching it. If its node disappears, nothing brings it back, and nothing runs it again tomorrow night. So in practice you almost never create pods directly. You create a higher-level object (a Deployment, a Job or a CronJob) that creates pods for you and keeps them in line. Choosing between those is the next lesson.

Recap

  • A pod wraps one or more containers that share a node, a network address and volumes.
  • Most pods have one container; init containers and sidecars are for tightly coupled helpers.
  • Phases (Pending, Running, Succeeded, Failed) and statuses like CrashLoopBackOff tell you where a pod is stuck.
  • kubectl logs, kubectl describe and kubectl get -o yaml are the first three things to run.
  • Pods are disposable, so pipeline output must be durable and every load idempotent.

This lesson ends with a 3-question checkpoint, graded in the app.

10 more lessons in this course

Start it in Akadyo to read on, take the checkpoints and keep your place, with a tutor beside every lesson.