A pod holds one container or more containers.
Each pod has an IP address.
The pod has a phase.
The phase shows the state of the pod.
The cluster gives the pod a node.
You can delete the pod.
Easy words. Short steps. Copy the commands. Each part gives the steps for a problem.
New? Start hereNew person? Read this first.
You do not need Kubernetes data to read this guide.
Read this part first. Then read the other parts.
Kubernetes controls a group of machines.
You give Kubernetes your application.
Kubernetes starts the application and keeps the application in operation.
If a part stops, Kubernetes starts the part again.
Think of it in easy words.
Think of the cluster as a group of machines.
Think of the node as one machine in the group.
Think of the pod as a unit that holds the application.
Think of the container as the application package in the pod.
Think of the namespace as one room for your objects.
Think of the manifest as the file that tells Kubernetes the number and the type.
Do these steps to use this guide.
Do these steps to read a command.
GOOD means the part works. BAD means the part has a problem.
GOOD shows Running or Ready or an address.
BAD shows a problem state such as Pending or CrashLoopBackOff.
Find these words in your output.
Kubernetes holds your application in a cluster. A pod is the basic unit of the cluster. A deployment controls a group of pods. An ingress gives persons on the Internet access to a service in the cluster. Each part gives the steps for a problem.
Think of it in easy words.
Think of the pod as the smallest unit.
Think of the deployment as the part that keeps the number.
Think of the ingress as the part that gives access from the Internet.
A pod holds one container or more containers.
Each pod has an IP address.
The pod has a phase.
The phase shows the state of the pod.
The cluster gives the pod a node.
You can delete the pod.
A deployment controls a group of pods.
You set the number of pods in the manifest.
The deployment keeps the number at this value.
The deployment gives new pods to the cluster.
This change is a rollout.
A rollout replaces the pods one group at a time.
An ingress gives persons on the Internet access to a service in the cluster.
The ingress has one or more rules.
Each rule has a host and a path.
The ingress sends the traffic to the correct service.
An ingress controller does the work.
The controller reads the rules of the ingress.
Read these words first. They help new persons.
The three command types. They help new persons.
Each part uses the same three types.
This command shows one object.
kubectl get <type> <name> -n <namespace>
This command shows one object.
This command shows the full data of one object.
kubectl describe <type> <name> -n <namespace>
This command shows the full data of one object.
This command shows the output of a container.
kubectl logs <pod-name> -n <namespace>
This command shows the output of a container.
New person? Read this first.
Think of the pod as the smallest unit. If the pod has a problem, the application stops.
First, do a check of the pod state.
kubectl get pods -n <namespace>
This command shows each pod and the state.
NAME READY STATUS RESTARTS
my-app 1/1 Running 0
GOOD output shows Running.
NAME READY STATUS RESTARTS
my-app 0/1 CrashLoopBackOff 5
BAD output shows a problem state.
Then, look at the events and the state of the pod.
kubectl describe pod <pod-name> -n <namespace>
This command shows the full data of the pod. Find Events at the end.
Then, look at the logs of the container.
kubectl logs <pod-name> -n <namespace>
This command shows the output of the application. Find the error at the end.
Read the items that follow.
The cluster does not start the pod.
Do a check of the events.
The node does not have sufficient CPU or memory.
The image is not available.
The container stops again and again.
Look at the logs of the container.
Correct the error in the application.
Then start the pod again.
The cluster cannot get the image.
Do a check of the image name.
Do a check of the secret of the registry.
The container uses more memory than the limit.
Change the memory limit in the manifest.
Or make the application use less memory.
The pod is not ready.
The readiness probe does not work.
Do a check of the probe path and the probe port.
New person? Read this first.
Think of the deployment as the part that keeps the number. If the number is not correct, look here.
First, do a check of the deployment state.
kubectl get deployment <name> -n <namespace>
This command shows the number of ready pods and the number of necessary pods.
NAME READY UP-TO-DATE AVAILABLE
my-app 3/3 3 3
GOOD output shows the same number in READY and UP-TO-DATE.
NAME READY UP-TO-DATE AVAILABLE
my-app 0/3 3 0
BAD output shows a small READY number.
Then, look at the rollout state.
kubectl rollout status deployment/<name> -n <namespace>
This command shows if the change to new pods is complete.
Then, look at the events of the deployment.
kubectl describe deployment <name> -n <namespace>
This command shows the full data of the deployment. Find Events at the end.
Then, look at the replica sets.
kubectl get replicaset -n <namespace>
This command shows each replica set and the number of pods.
Read the items that follow.
The pods are not ready.
Look at the pod problems.
The rollout does not move.
Do a check of the events.
The image is not available.
Or the image name is not correct.
The deployment does not create pods.
Do a check of the events.
The service account does not have sufficient access.
The pods use too much memory.
Change the memory limit in the manifest.
New person? Read this first.
Think of the ingress as the part that gives access from the Internet. If the browser cannot get the application, look here.
First, do a check of the ingress state.
kubectl get ingress -n <namespace>
This command shows each ingress and the address. Find your host in the output.
NAME CLASS HOSTS ADDRESS PORTS
my-app nginx my-app.test 203.0.113.10 80
GOOD output shows an address.
NAME CLASS HOSTS ADDRESS PORTS
my-app nginx my-app.test 80
BAD output shows no address.
Then, look at the events of the ingress.
kubectl describe ingress <name> -n <namespace>
This command shows the full data of the ingress. Find Events at the end.
Then, do a check of the service and the endpoints.
kubectl get service,endpoints -n <namespace>
This command shows each service and each endpoint. An empty endpoint is BAD.
Then, do a check of the ingress controller.
kubectl get pods -n ingress-nginx
This command shows the pods of the controller. They must show Running.
Read the items that follow.
The ingress gives a 400 code.
The request is not correct.
Do a check of the request of the client.
The ingress gives a 401 code.
The request has no token.
Do a check of the token in the request.
The ingress gives a 403 code.
The role does not give access.
Do a check of the role of the token.
The ingress gives a 404 code.
The host or the path is not correct.
Compare the host in the ingress with the host in your request.
Compare the path in the ingress with the path in the service.
The ingress gives a 500 code.
The container has an error.
Look at the logs of the container.
The ingress gives a 502 code.
The port of the service is not correct.
Do a check of the target port.
The ingress gives a 503 code.
The endpoint is not available.
Do a check of the endpoints of the service.
The ingress gives a 504 code.
The pod is too slow.
Look at the logs of the pod.
The ingress has no address.
The ingress controller is not available.
Or the controller does not have a load balancer.
Do a check of the controller pods.
The host is not correct.
Add the host to your DNS record.
Or use the IP address of the load balancer.
New person? Read this first.
Think of the service as a stable address. Pods change, but the service address stays the same.
A service gives a stable address to a group of pods.
First, do a check of the service.
kubectl get service <name> -n <namespace>
This command shows each service and the address.
Then, look at the service.
kubectl describe service <name> -n <namespace>
This command shows the full data of the service. Find Endpoints in the output.
Then, do a check of the endpoints.
kubectl get endpoints <name> -n <namespace>
This command shows the address of each ready pod. An empty result is BAD.
Read the items that follow.
The endpoints are empty.
The selector of the service does not match a pod label.
Do a check of the pod labels.
The port of the service is not correct.
Do a check of the port of the pod.
The service has no IP address.
Look at the events of the service.
The service can have a bad type.
New person? Read this first.
Think of the config map as the data for the application. Think of the secret as the keys.
A config map holds the data for the application. A secret holds the keys.
First, do a check of the config maps.
kubectl get configmap -n <namespace>
This command shows each config map in the namespace.
Then, do a check of the secrets.
kubectl get secret -n <namespace>
This command shows each secret in the namespace.
Then, look at the config map.
kubectl describe configmap <name> -n <namespace>
This command shows the full data of the config map. Find your key in the output.
Read the items that follow.
The application has an error.
The key in the config map is not correct.
Compare the key in the manifest with the key in the application.
The application cannot get the key.
The secret is not in the namespace.
Do a check of the namespace of the secret.
The config map is not in the namespace.
Apply the manifest for the config map.
New person? Read this first.
Think of the node as one machine. If the machine has a problem, all pods on the machine have a problem.
A node is a machine in the cluster.
First, do a check of the nodes.
kubectl get nodes
This command shows each machine and the state. Find Ready in the output.
NAME STATUS ROLES
node-1 Ready worker
GOOD output shows Ready.
NAME STATUS ROLES
node-1 NotReady worker
BAD output shows NotReady.
Then, look at the node.
kubectl describe node <name>
This command shows the full data of the node. Find Conditions and Events at the end.
Read the items that follow.
The node is NotReady.
Look at the events of the node.
The disk of the node can be full.
The node has MemoryPressure.
The node uses too much memory.
Remove a pod or add a node.
The node has DiskPressure.
The disk of the node is full.
Remove data from the node.
The node does not start new pods.
The node does not have sufficient CPU or memory.
Remove a pod or add a node.
New person? Read this first.
Think of the claim as a request for disk space. The pod asks, the cluster gives.
A pod uses a persistent volume claim for storage.
First, do a check of the claims.
kubectl get pvc -n <namespace>
This command shows each claim and the state. Find Bound in the output.
NAME STATUS VOLUME CAPACITY
my-disk Bound pv-1 10Gi
GOOD output shows Bound.
NAME STATUS VOLUME CAPACITY
my-disk Pending 10Gi
BAD output shows Pending.
Then, look at the claim.
kubectl describe pvc <name> -n <namespace>
This command shows the full data of the claim. Find Events at the end.
Read the items that follow.
The claim is in Pending state.
The storage class does not exist.
Do a check of the storage class.
The pod does not get sufficient storage.
Increase the storage request in the manifest.
The pod does not start.
The volume does not attach to the pod.
Look at the events of the pod.
New person? Read this first.
Think of the job as work that runs one time. A deployment keeps pods in operation, a job stops after the work is complete.
A job runs a pod one time.
First, do a check of the jobs.
kubectl get jobs -n <namespace>
This command shows each job and if the work is complete.
Then, look at the job.
kubectl describe job <name> -n <namespace>
This command shows the full data of the job. Find Events at the end.
Then, look at the logs of the job pod.
kubectl logs job/<name> -n <namespace>
This command shows the output of the work. Find the error at the end.
Read the items that follow.
The job does not start.
Look at the events of the job.
The image is not available.
The pod of the job has an error.
Look at the logs of the pod.
Correct the error in the command.
The job shows BackoffLimitExceeded.
Look at the events of the job.
Correct the error in the command.
New person? Read this first.
Think of the stateful set as pods with stable names. Think of the daemon set as one pod on each machine.
A stateful set gives pods a stable name and a stable volume. A daemon set gives each node one pod.
First, do a check of the stateful sets.
kubectl get statefulset -n <namespace>
This command shows each stateful set and the number of ready pods.
Then, do a check of the daemon sets.
kubectl get daemonset -n <namespace>
This command shows each daemon set and the number of ready pods on the nodes.
Then, look at the stateful set.
kubectl describe statefulset <name> -n <namespace>
This command shows the full data of the stateful set. Find Events at the end.
Read the items that follow.
The pods have no stable name.
The stateful set does not have a service.
Do a check of the service of the stateful set.
The daemon set does not start a pod on a node.
Look at the events of the daemon set.
The node can have a taint.
The pods cannot start.
The persistent volume claim can be in Pending state.
Look at the claims.
New person? Read this first.
Think of the autoscaler as the part that adds pods.
The autoscaler adds pods when the load is large.
The autoscaler removes pods when the load is small.
The autoscaler changes the number of pods as a function of the load.
First, do a check of the autoscalers.
kubectl get hpa -n <namespace>
This command shows each autoscaler and the present number of pods.
Then, look at the autoscaler.
kubectl describe hpa <name> -n <namespace>
This command shows the full data of the autoscaler. Find Events at the end.
Then, look at the load of the pods.
kubectl top pods -n <namespace>
This command shows the CPU and memory of each pod.
Read the items that follow.
The autoscaler does not change the pods.
The metrics are not available.
Do a check of the metrics server.
The autoscaler gives the wrong number of pods.
The limit in the manifest is not correct.
Do a check of the minimum and the maximum.
The pods are slow.
The limit for CPU is too small.
Change the CPU limit.
New person? Read this first.
Think of the service account as the name of the application. Think of the role as the list of permitted actions.
A service account gives a workload an identity. A role gives the rules for that identity.
First, do a check of the service accounts.
kubectl get serviceaccount -n <namespace>
This command shows each service account in the namespace.
Then, do a check of the roles.
kubectl get role,rolebinding -n <namespace>
This command shows each role and each binding in the namespace.
Then, do a check of the access.
kubectl auth can-i <verb> <resource> --as=system:serviceaccount:<namespace>:<name>
This command tells you if the service account can do one action. Change the verb and the resource for your test.
Read the items that follow.
The application shows a forbidden error.
The role does not give the rule for the API.
Do a check of the rules of the role.
The application gives an unknown user error.
The service account is in the wrong namespace.
Do a check of the namespace of the service account.
The binding does not exist.
Apply the role binding manifest.
New person? Read this first.
Think of the network policy as a firewall. It stops traffic that does not have a rule.
A network policy is a firewall between the pods.
First, do a check of the network policies.
kubectl get networkpolicy -n <namespace>
This command shows each network policy in the namespace.
Then, look at the network policy.
kubectl describe networkpolicy <name> -n <namespace>
This command shows the full data of the policy. Find your source in the rules.
Read the items that follow.
The traffic is blocked to a pod.
The network policy does not have a rule for the source.
Add a rule for the source.
All the traffic is blocked.
The default policy stops all traffic.
Change the default policy.
The traffic is blocked from another namespace.
The rule does not have the namespace as a source.
Add the namespace to the rule.
New person? Read this first.
Think of the namespace as one room for your objects. Each team can have a room.
A namespace is a part of the cluster for the objects.
First, do a check of the namespaces.
kubectl get namespaces
This command shows each room in the cluster.
Then, do a check of the objects in the namespace.
kubectl get all -n <namespace>
This command shows all objects in your room.
Then, look at the namespace.
kubectl describe namespace <name>
This command shows the full data of the namespace. Find the state at the start.
Read the items that follow.
The objects are not in the namespace.
You look in the wrong namespace.
Do a check of the namespace of the object.
The service account cannot see the objects in another namespace.
The role is in the wrong namespace.
Do a check of the binding.
The namespace is in the Terminating phase.
The objects do not delete.
Look at the events of the namespace.
New person? Read this first.
Think of the limit as the maximum for one pod. Think of the quota as the maximum for the room.
Limits give a maximum for a pod. Quotas give a maximum for the namespace.
First, do a check of the limit ranges.
kubectl describe limitrange -n <namespace>
This command shows the default minimum and maximum for new pods.
Then, do a check of the quotas.
kubectl describe resourcequota -n <namespace>
This command shows how much of the quota the namespace uses at present.
Read the items that follow.
The pod does not start.
The limit in the manifest is more than the quota.
Do a check of the quota of the namespace.
The pod does not have a request for CPU.
Set a default in the limit range.
The quota uses all of the capacity.
Remove a pod or increase the quota.
New person? Read this first.
Think of TLS as a lock on the connection. The lock needs a certificate.
TLS gives a secure connection to the ingress.
First, do a check of the ingress.
kubectl get ingress -n <namespace>
This command shows each ingress. Find the host and the TLS part.
Then, look at the secret.
kubectl describe secret <name> -n <namespace>
This command shows the full data of the secret. The secret holds the certificate.
Then, do a check of the certificate.
kubectl get certificate -n <namespace>
This command shows each certificate and if it is ready.
Read the items that follow.
The connection is not secure.
The TLS section is not in the manifest.
Add the TLS section.
The certificate is not available.
The issuer is not ready.
Look at the events of the certificate.
The host in the certificate is not correct.
Do a check of the host in the secret.
New person? Read this first.
Think of Helm as a tool that installs a group of manifests at the same time.
Helm is a package manager for the cluster.
First, do a check of the releases.
helm list -n <namespace>
This command shows each release and the state.
Then, look at the release.
helm status <name> -n <namespace>
This command shows the full data of the release and the present state.
Then, look at the history of the release.
helm history <name> -n <namespace>
This command shows each change of the release. Find the bad change.
Read the items that follow.
The release does not install.
The chart is not available.
Do a check of the name of the chart.
The application has an error after the upgrade.
Do a rollback of the release.
The release shows pending.
The previous command did not stop.
Delete the pending release and start again.
New person? Read this first.
Think of the custom resource as your own object. Think of the operator as the part that manages your object.
A custom resource is your own object in the cluster. An operator manages the object.
First, do a check of the definitions.
kubectl get crd
This command shows each new object type in the cluster.
Then, do a check of the custom resources.
kubectl get <kind> -n <namespace>
This command shows each object of your type in the namespace.
Then, look at the custom resource.
kubectl describe <kind> <name> -n <namespace>
This command shows the full data of your object. Find Events at the end.
Read the items that follow.
The custom resource does not create pods.
The operator is not running.
Look at the logs of the operator.
The API does not have the custom resource.
The definition does not exist.
Do a check of the name of the definition.
The operator shows an error.
The manifest is not correct.
Compare the manifest with the definition.
New person? Read this first.
Think of a request as a person who knocks on four doors. Each door checks the request.
A request moves from left to right.
Each part checks the request.
If all parts work, the ingress gives a 200 code.
The request works.
The request is not correct.
The request has no token.
The role does not give access.
The host or the path is not correct.
The container has an error.
The port of the service is not correct.
The endpoint is not available.
The pod is too slow.
Find the code in the response. Then do the steps for the code.
CAUTION: Do not change the production manifest without a copy.
An incorrect change can stop the service.
Make a copy of the manifest before you change it.