Lesson 1
Why a data team moves to Kubernetes
What you'll learn: What Kubernetes actually does for a data team, how a cluster is put together, and why "declaring the state you want" changes how you run pipelines.
Meet the pipeline we will move
Throughout this course we follow the data team at Harbour Lane Grocers, a regional chain of 40 shops. The team lead, Tomás, looks after a nightly batch pipeline called sales-load. Every night at 02:00 a cron entry on a single virtual machine starts a Python script. It pulls the day's till files from object storage, cleans them, and loads about 3 million rows into the warehouse before the morning sales report runs at 07:00.
It already runs in a Docker container, which is good news. The trouble is everything around the container:
- When the VM was patched and rebooted at 02:05 one night, the load simply never happened, and nobody knew until the report was empty.
- The VM is sized for the busiest night of the year, so it sits idle about 22 hours a day.
- The team now wants Airflow for scheduling and Spark for a heavier pricing model, and nobody wants three more hand-managed servers.
Kubernetes is the tool Tomás is evaluating. By the end of this course the sales-load job, a small API, Airflow and Spark will all run on one cluster, and you will understand every piece.
What Kubernetes is, in one paragraph
Kubernetes (often written K8s) is an open-source system that runs containers across a group of machines. You tell it what you want running, such as "one copy of sales-load every night at 02:00, with 2 GB of memory", and it decides which machine runs it, starts it, restarts it if it crashes, and reports what happened. It does not build your images or replace your warehouse. It is the layer that keeps containers running where and when you asked.
The parts of a cluster
A Kubernetes cluster has two halves.
The control plane is the brain. It stores the desired state of everything in a database called etcd, exposes an API server that every tool talks to, runs a scheduler that picks a machine for each new workload, and runs controllers that notice when reality drifts from what you asked for.
The nodes are the muscle: the virtual or physical machines that actually run containers. Each node runs an agent called the kubelet, which takes instructions from the control plane and starts containers through a container runtime (usually containerd).
| Part | Lives on | Job |
|---|---|---|
| API server | Control plane | The front door; every command goes through it |
| etcd | Control plane | Stores the desired and current state |
| Scheduler | Control plane | Chooses a node for each new pod |
| Controllers | Control plane | Keep nudging reality towards the desired state |
| kubelet | Every node | Starts and watches containers on that node |
Desired state, not instructions
The idea that makes Kubernetes feel different from a cron VM is that you describe an outcome rather than a sequence of steps. Think of a thermostat. You do not tell the heating "run for 20 minutes, then stop". You set 20 degrees and the thermostat keeps checking the room, switching the heating on and off until the room matches. If someone opens a window, it reacts without you.
Kubernetes controllers work the same way. Tomás will write a short YAML file saying a job should exist with a given image and schedule, then apply it with kubectl apply -f sales-load.yaml. The controllers compare that description with what is running, and act on the difference. If a node dies halfway through the load, the controller sees that the job has not finished and starts it again on a healthy node. That alone fixes the "VM rebooted at 02:05" incident.
Every YAML file starts with two key lines: apiVersion: batch/v1 names the API version, and kind: CronJob says what sort of object it is. Later lessons build these files up a line at a time.
Talking to the cluster with kubectl
kubectl is the command-line tool for the API server. A few commands cover most of a data engineer's day:
kubectl get podslists what is running.kubectl describe pod sales-load-abc12explains one object in detail, including recent events.kubectl logs sales-load-abc12shows a container's output.kubectl apply -f file.yamlcreates or updates objects from a file.kubectl delete -f file.yamlremoves them.
Your kubectl talks to one cluster at a time, chosen by a context in your kubeconfig file. kubectl config current-context tells you which one, and checking it before a delete is a habit worth forming early.
Build your own or rent one
Running the control plane yourself is real work: upgrades roughly three times a year, etcd backups, certificates. Most data teams rent it. The three big managed offerings are Amazon EKS, Google Kubernetes Engine (GKE) and Azure Kubernetes Service (AKS). Each runs the control plane for you and charges either a small hourly fee per cluster or nothing at all on a basic tier, plus the normal price of the nodes. GKE Autopilot and EKS Auto Mode go further and manage the nodes too. Prices and tiers change, so check your provider's current page rather than trusting a number in a course.
For learning, a local cluster is plenty. kind, minikube, k3d and the Kubernetes option built into Docker Desktop all give you a one-node cluster on a laptop. Tomás starts with kind so the team can break things freely.
Is it worth it for a data team?
Kubernetes adds moving parts. For one nightly script, a serverless job service may be simpler. It pays off when, as at Harbour Lane, you have several workloads with different shapes (nightly batch, always-on APIs, bursty Spark) that you want on shared, self-healing infrastructure with one way of deploying. That is the case we will build.
Recap
- Kubernetes runs containers across many machines and keeps them in the state you declared.
- A cluster has a control plane (API server, etcd, scheduler, controllers) and nodes running the kubelet.
- You describe the outcome in YAML and apply it; controllers close the gap, like a thermostat.
kubectltalks to the API server; always know which context you are pointed at.- EKS, GKE and AKS run the control plane for you; kind or minikube are fine for learning.