K8s STE Guide

Kubernetes for New Persons

Easy words. Short steps. Copy the commands. Each part gives the steps for a problem.

New? Start here
Pod Deployment Ingress Service ConfigMap Node PVC Job StatefulSet HPA RBAC NetworkPolicy Namespace TLS Helm Operators
1

Start here

New person? Read this first.

You do not need Kubernetes data to read this guide.

Read this part first. Then read the other parts.

Kubernetes controls a group of machines.

You give Kubernetes your application.

Kubernetes starts the application and keeps the application in operation.

If a part stops, Kubernetes starts the part again.

Think of it in easy words.

Think of the cluster as a group of machines.

Think of the node as one machine in the group.

Think of the pod as a unit that holds the application.

Think of the container as the application package in the pod.

Think of the namespace as one room for your objects.

Think of the manifest as the file that tells Kubernetes the number and the type.

Do these steps to use this guide.

  1. Read the words to know.
  2. Read the easy intro in each part.
  3. Run one command at a time.
  4. Compare your output with the GOOD and BAD outputs.

Do these steps to read a command.

  1. Replace your names for the example names.
  2. Use your namespace in each command.
  3. The text in < > is an example. You must change it.

GOOD means the part works. BAD means the part has a problem.

GOOD shows Running or Ready or an address.

BAD shows a problem state such as Pending or CrashLoopBackOff.

Find these words in your output.

2

The objects

Kubernetes holds your application in a cluster. A pod is the basic unit of the cluster. A deployment controls a group of pods. An ingress gives persons on the Internet access to a service in the cluster. Each part gives the steps for a problem.

Think of it in easy words.

Think of the pod as the smallest unit.

Think of the deployment as the part that keeps the number.

Think of the ingress as the part that gives access from the Internet.

Pod

A pod holds one container or more containers.

Each pod has an IP address.

The pod has a phase.

The phase shows the state of the pod.

The cluster gives the pod a node.

You can delete the pod.

Deployment

A deployment controls a group of pods.

You set the number of pods in the manifest.

The deployment keeps the number at this value.

The deployment gives new pods to the cluster.

This change is a rollout.

A rollout replaces the pods one group at a time.

Ingress

An ingress gives persons on the Internet access to a service in the cluster.

The ingress has one or more rules.

Each rule has a host and a path.

The ingress sends the traffic to the correct service.

An ingress controller does the work.

The controller reads the rules of the ingress.

3

Words to know

Read these words first. They help new persons.

pod
The pod is the unit of the cluster.
The pod holds the application.
deployment
The deployment controls the pods.
The deployment keeps the number of pods.
ingress
The ingress gives access to the service.
service
The service gives a stable address to the pods.
node
The node is one machine in the cluster.
cluster
The cluster is the group of machines.
container
The container holds the application in the pod.
manifest
The manifest is the file for the object.
The manifest tells Kubernetes the number and the type.
namespace
The namespace keeps your objects in one group.
event
The event shows a change in the cluster.
log
The log shows the output of the container.
endpoint
The endpoint is the address of a ready pod.
4

Before you do a check

  1. Run the commands in a terminal.
  2. Replace the example names with your names.
  3. Add the name of your namespace to each command.
  4. Do one command at a time.

The three command types. They help new persons.

Each part uses the same three types.

This command shows one object.

$
kubectl get <type> <name> -n <namespace>

This command shows one object.

This command shows the full data of one object.

$
kubectl describe <type> <name> -n <namespace>

This command shows the full data of one object.

This command shows the output of a container.

$
kubectl logs <pod-name> -n <namespace>

This command shows the output of a container.

5

Pod problems

New person? Read this first.

Think of the pod as the smallest unit. If the pod has a problem, the application stops.

Node
the machine
→
Pod
the unit
→
Container
the application
1

First, do a check of the pod state.

$
kubectl get pods -n <namespace>

This command shows each pod and the state.

GOOD
NAME      READY   STATUS    RESTARTS
my-app    1/1     Running   0

GOOD output shows Running.

BAD
NAME      READY   STATUS             RESTARTS
my-app    0/1     CrashLoopBackOff   5

BAD output shows a problem state.

2

Then, look at the events and the state of the pod.

$
kubectl describe pod <pod-name> -n <namespace>

This command shows the full data of the pod. Find Events at the end.

3

Then, look at the logs of the container.

$
kubectl logs <pod-name> -n <namespace>

This command shows the output of the application. Find the error at the end.

Read the items that follow.

Pending state

The cluster does not start the pod.

Do a check of the events.

The node does not have sufficient CPU or memory.

The image is not available.

CrashLoopBackOff state

The container stops again and again.

Look at the logs of the container.

Correct the error in the application.

Then start the pod again.

ImagePullBackOff state

The cluster cannot get the image.

Do a check of the image name.

Do a check of the secret of the registry.

OOMKilled state

The container uses more memory than the limit.

Change the memory limit in the manifest.

Or make the application use less memory.

Not ready

The pod is not ready.

The readiness probe does not work.

Do a check of the probe path and the probe port.

6

Deployment problems

New person? Read this first.

Think of the deployment as the part that keeps the number. If the number is not correct, look here.

Deployment
controls the group
→
ReplicaSet
the replica set
→
Pods
run the application
1

First, do a check of the deployment state.

$
kubectl get deployment <name> -n <namespace>

This command shows the number of ready pods and the number of necessary pods.

GOOD
NAME     READY   UP-TO-DATE   AVAILABLE
my-app   3/3     3            3

GOOD output shows the same number in READY and UP-TO-DATE.

BAD
NAME     READY   UP-TO-DATE   AVAILABLE
my-app   0/3     3            0

BAD output shows a small READY number.

2

Then, look at the rollout state.

$
kubectl rollout status deployment/<name> -n <namespace>

This command shows if the change to new pods is complete.

3

Then, look at the events of the deployment.

$
kubectl describe deployment <name> -n <namespace>

This command shows the full data of the deployment. Find Events at the end.

4

Then, look at the replica sets.

$
kubectl get replicaset -n <namespace>

This command shows each replica set and the number of pods.

Read the items that follow.

Pods not ready

The pods are not ready.

Look at the pod problems.

Rollout stops

The rollout does not move.

Do a check of the events.

The image is not available.

Or the image name is not correct.

No pods

The deployment does not create pods.

Do a check of the events.

The service account does not have sufficient access.

Memory use

The pods use too much memory.

Change the memory limit in the manifest.

7

Ingress problems

New person? Read this first.

Think of the ingress as the part that gives access from the Internet. If the browser cannot get the application, look here.

Ingress
the rules
→
Service
the port
→
Pod
the application
1

First, do a check of the ingress state.

$
kubectl get ingress -n <namespace>

This command shows each ingress and the address. Find your host in the output.

GOOD
NAME     CLASS   HOSTS           ADDRESS         PORTS
my-app   nginx   my-app.test     203.0.113.10   80

GOOD output shows an address.

BAD
NAME     CLASS   HOSTS           ADDRESS   PORTS
my-app   nginx   my-app.test               80

BAD output shows no address.

2

Then, look at the events of the ingress.

$
kubectl describe ingress <name> -n <namespace>

This command shows the full data of the ingress. Find Events at the end.

3

Then, do a check of the service and the endpoints.

$
kubectl get service,endpoints -n <namespace>

This command shows each service and each endpoint. An empty endpoint is BAD.

4

Then, do a check of the ingress controller.

$
kubectl get pods -n ingress-nginx

This command shows the pods of the controller. They must show Running.

Read the items that follow.

400 code

The ingress gives a 400 code.

The request is not correct.

Do a check of the request of the client.

401 code

The ingress gives a 401 code.

The request has no token.

Do a check of the token in the request.

403 code

The ingress gives a 403 code.

The role does not give access.

Do a check of the role of the token.

404 code

The ingress gives a 404 code.

The host or the path is not correct.

Compare the host in the ingress with the host in your request.

Compare the path in the ingress with the path in the service.

500 code

The ingress gives a 500 code.

The container has an error.

Look at the logs of the container.

502 code

The ingress gives a 502 code.

The port of the service is not correct.

Do a check of the target port.

503 code

The ingress gives a 503 code.

The endpoint is not available.

Do a check of the endpoints of the service.

504 code

The ingress gives a 504 code.

The pod is too slow.

Look at the logs of the pod.

No address

The ingress has no address.

The ingress controller is not available.

Or the controller does not have a load balancer.

Do a check of the controller pods.

Wrong host

The host is not correct.

Add the host to your DNS record.

Or use the IP address of the load balancer.

8

Service problems

New person? Read this first.

Think of the service as a stable address. Pods change, but the service address stays the same.

A service gives a stable address to a group of pods.

Service
stable address
→
Endpoints
the pod list
→
Pod
the application
1

First, do a check of the service.

$
kubectl get service <name> -n <namespace>

This command shows each service and the address.

2

Then, look at the service.

$
kubectl describe service <name> -n <namespace>

This command shows the full data of the service. Find Endpoints in the output.

3

Then, do a check of the endpoints.

$
kubectl get endpoints <name> -n <namespace>

This command shows the address of each ready pod. An empty result is BAD.

Read the items that follow.

Empty endpoints

The endpoints are empty.

The selector of the service does not match a pod label.

Do a check of the pod labels.

Wrong port

The port of the service is not correct.

Do a check of the port of the pod.

No IP address

The service has no IP address.

Look at the events of the service.

The service can have a bad type.

9

ConfigMap and Secret problems

New person? Read this first.

Think of the config map as the data for the application. Think of the secret as the keys.

A config map holds the data for the application. A secret holds the keys.

ConfigMap / Secret
data and keys
→
Pod
uses the data
1

First, do a check of the config maps.

$
kubectl get configmap -n <namespace>

This command shows each config map in the namespace.

2

Then, do a check of the secrets.

$
kubectl get secret -n <namespace>

This command shows each secret in the namespace.

3

Then, look at the config map.

$
kubectl describe configmap <name> -n <namespace>

This command shows the full data of the config map. Find your key in the output.

Read the items that follow.

Key not correct

The application has an error.

The key in the config map is not correct.

Compare the key in the manifest with the key in the application.

Secret not found

The application cannot get the key.

The secret is not in the namespace.

Do a check of the namespace of the secret.

Config map not found

The config map is not in the namespace.

Apply the manifest for the config map.

10

Node problems

New person? Read this first.

Think of the node as one machine. If the machine has a problem, all pods on the machine have a problem.

A node is a machine in the cluster.

Node
the machine
→
Pods
run on the node
1

First, do a check of the nodes.

$
kubectl get nodes

This command shows each machine and the state. Find Ready in the output.

GOOD
NAME      STATUS   ROLES
node-1    Ready    worker

GOOD output shows Ready.

BAD
NAME      STATUS     ROLES
node-1    NotReady   worker

BAD output shows NotReady.

2

Then, look at the node.

$
kubectl describe node <name>

This command shows the full data of the node. Find Conditions and Events at the end.

Read the items that follow.

NotReady

The node is NotReady.

Look at the events of the node.

The disk of the node can be full.

MemoryPressure

The node has MemoryPressure.

The node uses too much memory.

Remove a pod or add a node.

DiskPressure

The node has DiskPressure.

The disk of the node is full.

Remove data from the node.

No new pods

The node does not start new pods.

The node does not have sufficient CPU or memory.

Remove a pod or add a node.

11

PersistentVolumeClaim problems

New person? Read this first.

Think of the claim as a request for disk space. The pod asks, the cluster gives.

A pod uses a persistent volume claim for storage.

Pod
asks for storage
→
PVC
the claim
→
PV
the disk
1

First, do a check of the claims.

$
kubectl get pvc -n <namespace>

This command shows each claim and the state. Find Bound in the output.

GOOD
NAME      STATUS   VOLUME   CAPACITY
my-disk   Bound    pv-1     10Gi

GOOD output shows Bound.

BAD
NAME      STATUS    VOLUME   CAPACITY
my-disk   Pending             10Gi

BAD output shows Pending.

2

Then, look at the claim.

$
kubectl describe pvc <name> -n <namespace>

This command shows the full data of the claim. Find Events at the end.

Read the items that follow.

Pending state

The claim is in Pending state.

The storage class does not exist.

Do a check of the storage class.

Small storage

The pod does not get sufficient storage.

Increase the storage request in the manifest.

Volume not attached

The pod does not start.

The volume does not attach to the pod.

Look at the events of the pod.

12

Job problems

New person? Read this first.

Think of the job as work that runs one time. A deployment keeps pods in operation, a job stops after the work is complete.

A job runs a pod one time.

Job
runs one time
→
Pod
does the work
1

First, do a check of the jobs.

$
kubectl get jobs -n <namespace>

This command shows each job and if the work is complete.

2

Then, look at the job.

$
kubectl describe job <name> -n <namespace>

This command shows the full data of the job. Find Events at the end.

3

Then, look at the logs of the job pod.

$
kubectl logs job/<name> -n <namespace>

This command shows the output of the work. Find the error at the end.

Read the items that follow.

Job does not start

The job does not start.

Look at the events of the job.

The image is not available.

Job pod error

The pod of the job has an error.

Look at the logs of the pod.

Correct the error in the command.

BackoffLimitExceeded

The job shows BackoffLimitExceeded.

Look at the events of the job.

Correct the error in the command.

13

StatefulSet and DaemonSet problems

New person? Read this first.

Think of the stateful set as pods with stable names. Think of the daemon set as one pod on each machine.

A stateful set gives pods a stable name and a stable volume. A daemon set gives each node one pod.

StatefulSet
stable names + volumes
→
Pod-0
volume 0
→
Pod-1
volume 1
1

First, do a check of the stateful sets.

$
kubectl get statefulset -n <namespace>

This command shows each stateful set and the number of ready pods.

2

Then, do a check of the daemon sets.

$
kubectl get daemonset -n <namespace>

This command shows each daemon set and the number of ready pods on the nodes.

3

Then, look at the stateful set.

$
kubectl describe statefulset <name> -n <namespace>

This command shows the full data of the stateful set. Find Events at the end.

Read the items that follow.

No stable name

The pods have no stable name.

The stateful set does not have a service.

Do a check of the service of the stateful set.

DaemonSet gap

The daemon set does not start a pod on a node.

Look at the events of the daemon set.

The node can have a taint.

Pods do not start

The pods cannot start.

The persistent volume claim can be in Pending state.

Look at the claims.

14

HorizontalPodAutoscaler problems

New person? Read this first.

Think of the autoscaler as the part that adds pods.

The autoscaler adds pods when the load is large.

The autoscaler removes pods when the load is small.

The autoscaler changes the number of pods as a function of the load.

HPA
watches the load
→
Deployment
scales the pods
→
Pods
1 to 10
1

First, do a check of the autoscalers.

$
kubectl get hpa -n <namespace>

This command shows each autoscaler and the present number of pods.

2

Then, look at the autoscaler.

$
kubectl describe hpa <name> -n <namespace>

This command shows the full data of the autoscaler. Find Events at the end.

3

Then, look at the load of the pods.

$
kubectl top pods -n <namespace>

This command shows the CPU and memory of each pod.

Read the items that follow.

No scaling

The autoscaler does not change the pods.

The metrics are not available.

Do a check of the metrics server.

Wrong pod count

The autoscaler gives the wrong number of pods.

The limit in the manifest is not correct.

Do a check of the minimum and the maximum.

Slow pods

The pods are slow.

The limit for CPU is too small.

Change the CPU limit.

15

RBAC and ServiceAccount problems

New person? Read this first.

Think of the service account as the name of the application. Think of the role as the list of permitted actions.

A service account gives a workload an identity. A role gives the rules for that identity.

ServiceAccount
the identity
→
RoleBinding
the link
→
Role
the rules
1

First, do a check of the service accounts.

$
kubectl get serviceaccount -n <namespace>

This command shows each service account in the namespace.

2

Then, do a check of the roles.

$
kubectl get role,rolebinding -n <namespace>

This command shows each role and each binding in the namespace.

3

Then, do a check of the access.

$
kubectl auth can-i <verb> <resource> --as=system:serviceaccount:<namespace>:<name>

This command tells you if the service account can do one action. Change the verb and the resource for your test.

Read the items that follow.

Forbidden error

The application shows a forbidden error.

The role does not give the rule for the API.

Do a check of the rules of the role.

Unknown user

The application gives an unknown user error.

The service account is in the wrong namespace.

Do a check of the namespace of the service account.

Binding missing

The binding does not exist.

Apply the role binding manifest.

16

NetworkPolicy problems

New person? Read this first.

Think of the network policy as a firewall. It stops traffic that does not have a rule.

A network policy is a firewall between the pods.

Client Pod
the source
→
NetworkPolicy
the firewall
→
Server Pod
the target
1

First, do a check of the network policies.

$
kubectl get networkpolicy -n <namespace>

This command shows each network policy in the namespace.

2

Then, look at the network policy.

$
kubectl describe networkpolicy <name> -n <namespace>

This command shows the full data of the policy. Find your source in the rules.

Read the items that follow.

Traffic blocked

The traffic is blocked to a pod.

The network policy does not have a rule for the source.

Add a rule for the source.

All blocked

All the traffic is blocked.

The default policy stops all traffic.

Change the default policy.

Namespace blocked

The traffic is blocked from another namespace.

The rule does not have the namespace as a source.

Add the namespace to the rule.

17

Namespace problems

New person? Read this first.

Think of the namespace as one room for your objects. Each team can have a room.

A namespace is a part of the cluster for the objects.

Namespace
holds the objects
→
Deployment + Service + Pod
the group
1

First, do a check of the namespaces.

$
kubectl get namespaces

This command shows each room in the cluster.

2

Then, do a check of the objects in the namespace.

$
kubectl get all -n <namespace>

This command shows all objects in your room.

3

Then, look at the namespace.

$
kubectl describe namespace <name>

This command shows the full data of the namespace. Find the state at the start.

Read the items that follow.

Wrong namespace

The objects are not in the namespace.

You look in the wrong namespace.

Do a check of the namespace of the object.

Role misplaced

The service account cannot see the objects in another namespace.

The role is in the wrong namespace.

Do a check of the binding.

Terminating

The namespace is in the Terminating phase.

The objects do not delete.

Look at the events of the namespace.

18

Limits and quotas problems

New person? Read this first.

Think of the limit as the maximum for one pod. Think of the quota as the maximum for the room.

Limits give a maximum for a pod. Quotas give a maximum for the namespace.

ResourceQuota
namespace maximum
→
Namespace
the objects
→
Pod
limited
1

First, do a check of the limit ranges.

$
kubectl describe limitrange -n <namespace>

This command shows the default minimum and maximum for new pods.

2

Then, do a check of the quotas.

$
kubectl describe resourcequota -n <namespace>

This command shows how much of the quota the namespace uses at present.

Read the items that follow.

Limit over quota

The pod does not start.

The limit in the manifest is more than the quota.

Do a check of the quota of the namespace.

No CPU request

The pod does not have a request for CPU.

Set a default in the limit range.

Quota used

The quota uses all of the capacity.

Remove a pod or increase the quota.

19

TLS problems in ingress

New person? Read this first.

Think of TLS as a lock on the connection. The lock needs a certificate.

TLS gives a secure connection to the ingress.

Client
HTTPS
→
Ingress + cert
TLS section
→
Service
the port
→
Pod
the application
1

First, do a check of the ingress.

$
kubectl get ingress -n <namespace>

This command shows each ingress. Find the host and the TLS part.

2

Then, look at the secret.

$
kubectl describe secret <name> -n <namespace>

This command shows the full data of the secret. The secret holds the certificate.

3

Then, do a check of the certificate.

$
kubectl get certificate -n <namespace>

This command shows each certificate and if it is ready.

Read the items that follow.

Not secure

The connection is not secure.

The TLS section is not in the manifest.

Add the TLS section.

No certificate

The certificate is not available.

The issuer is not ready.

Look at the events of the certificate.

Wrong host

The host in the certificate is not correct.

Do a check of the host in the secret.

20

Helm problems

New person? Read this first.

Think of Helm as a tool that installs a group of manifests at the same time.

Helm is a package manager for the cluster.

Helm
the tool
→
Chart
the package
→
Manifests
the files
→
Cluster
the objects
1

First, do a check of the releases.

$
helm list -n <namespace>

This command shows each release and the state.

2

Then, look at the release.

$
helm status <name> -n <namespace>

This command shows the full data of the release and the present state.

3

Then, look at the history of the release.

$
helm history <name> -n <namespace>

This command shows each change of the release. Find the bad change.

Read the items that follow.

Install blocked

The release does not install.

The chart is not available.

Do a check of the name of the chart.

Upgrade error

The application has an error after the upgrade.

Do a rollback of the release.

Release pending

The release shows pending.

The previous command did not stop.

Delete the pending release and start again.

21

CustomResource and Operator problems

New person? Read this first.

Think of the custom resource as your own object. Think of the operator as the part that manages your object.

A custom resource is your own object in the cluster. An operator manages the object.

Operator
the manager
→
CustomResource
your object
→
Pod
created by the operator
1

First, do a check of the definitions.

$
kubectl get crd

This command shows each new object type in the cluster.

2

Then, do a check of the custom resources.

$
kubectl get <kind> -n <namespace>

This command shows each object of your type in the namespace.

3

Then, look at the custom resource.

$
kubectl describe <kind> <name> -n <namespace>

This command shows the full data of your object. Find Events at the end.

Read the items that follow.

No pods

The custom resource does not create pods.

The operator is not running.

Look at the logs of the operator.

API missing CR

The API does not have the custom resource.

The definition does not exist.

Do a check of the name of the definition.

Operator error

The operator shows an error.

The manifest is not correct.

Compare the manifest with the definition.

22

The request path

New person? Read this first.

Think of a request as a person who knocks on four doors. Each door checks the request.

A request moves from left to right.

Each part checks the request.

If all parts work, the ingress gives a 200 code.

Person
makes the request
→
Ingress
checks the host and the path
→
Service
has the port
→
Pod
has the container
200
all parts

The request works.

400
at the ingress

The request is not correct.

401
at the ingress

The request has no token.

403
at the ingress

The role does not give access.

404
at the ingress

The host or the path is not correct.

500
at the pod

The container has an error.

502
at the service

The port of the service is not correct.

503
at the pod

The endpoint is not available.

504
at the pod

The pod is too slow.

Find the code in the response. Then do the steps for the code.

23

A safe change

CAUTION: Do not change the production manifest without a copy.

An incorrect change can stop the service.

Make a copy of the manifest before you change it.

  1. Change one item at a time.
  2. Do a check of the result after each change.
  3. Look at the events before you look at the logs.
  4. Keep the manifest until the new pods work.