“Autoscaling” names three different mechanisms in Kubernetes, and confusing them is the most common reason a cluster that should scale does not. Pods scale with a HorizontalPodAutoscaler, nodes scale with worker-pool bounds, and the control plane scales with the load on the API. This article explains each one, then walks through a real test we ran on a Kube-DC Managed Cluster, with the manifests so you can repeat it.

Three things can scale

Layer What scales On a Kube-DC Managed Cluster
Pods The number of replicas of a Deployment or StatefulSet, driven by a metric such as CPU Standard HorizontalPodAutoscaler, with metrics-server already installed. Tested in this article
Nodes The number of worker machines available to run those pods Worker pools with minimum and maximum bounds per pool, as documented for the platform. Not part of our test
Control plane The resources of the API server and etcd as cluster load grows Vertical autoscaling of both, on by default and within configured bounds, as documented. Not part of our test

We label each layer with what we tested and what we only know from the documentation, because the difference matters when you plan capacity.

What the HorizontalPodAutoscaler actually does

The HorizontalPodAutoscaler (HPA) is a control loop that runs inside the Kubernetes control plane. By default it wakes up every 15 seconds, reads the metrics for the pods it targets, and adjusts the replica count of the Deployment. Understanding four details explains almost every surprise:

  • Utilization is relative to the request. A target of 50% CPU means 50% of the CPU requested by the pod’s containers, not 50% of the node. If a container has no CPU request, utilization is undefined and the autoscaler takes no action on that metric.
  • It needs the Metrics API. The metrics.k8s.io API is normally provided by the Metrics Server add-on, which has to be installed separately on a self-managed cluster.
  • The arithmetic is a ratio. Desired replicas equal the current replicas multiplied by the ratio of the current metric to the target, rounded up. If the ratio is within a tolerance of 0.1 of 1.0, nothing happens.
  • Scaling down is deliberately slow. The controller remembers recent recommendations and acts on the highest one in a window that defaults to five minutes, which smooths out fluctuating metrics and prevents thrashing.

For example, one replica running at 57% CPU against a 50% target gives a ratio of 1.14, and rounding up 1 × 1.14 gives 2 replicas. That is exactly the first step we observed in the test below.

The test: scaling a Managed Cluster under real load

In our August 2026 end-to-end test we created a Managed Cluster named test2 with two worker nodes, each sized at 1 vCPU and 2 GB of memory to fit the headroom left in the package’s quota. The cluster’s API endpoint is private by default, so we reached it from the project’s built-in Cloud Shell through the cluster’s internal service address rather than from the internet. Both workers were Ready, and metrics-server was already installed and working, which the HPA requires.

We deployed a small web application with an HPA targeting 50% CPU and a range of 1 to 6 replicas. These are example manifests in the same shape as the test; the CPU request values are illustrative:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: hello
spec:
  replicas: 1
  selector:
    matchLabels:
      app: hello
  template:
    metadata:
      labels:
        app: hello
    spec:
      containers:
      - name: hello
        image: nginxdemos/hello
        ports:
        - containerPort: 80
        resources:
          requests:
            cpu: 100m
            memory: 64Mi
---
apiVersion: v1
kind: Service
metadata:
  name: hello
spec:
  selector:
    app: hello
  ports:
  - port: 80
    targetPort: 80
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: hello
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: hello
  minReplicas: 1
  maxReplicas: 6
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 50

To generate sustained load we started several throw-away pods inside the cluster, each looping HTTP requests at the service. Run the first command in two or three shells with different names, and watch the autoscaler in another:

kubectl run load-1 --image=busybox:1.36 --restart=Never -- /bin/sh -c "while true; do wget -q -O- http://hello > /dev/null; done"

kubectl get hpa hello --watch
Result: CPU rose to 57% under load, the HPA scaled the deployment from 1 to 2 and then 3 replicas, and average CPU settled at roughly 35%. When the load generators were removed, the replica count returned to 1 without any manual action.

Two honest notes about that result. We did not time the scale-down, and the default five-minute stabilization window is why it lags behind the scale-up. And the test proves the pod layer only: it shows that the HPA, the Metrics API and the cluster’s scheduling work together on a Kube-DC Managed Cluster, not how the worker pools behave under node-level scaling.

When you finish, remove the load with kubectl delete pod load-1 (and any other generators you started), then the test workload with kubectl delete hpa,svc,deploy hello.

What pods cannot do alone: capacity

The HPA adds pods, and pods need somewhere to run. If the workers are full, the extra replicas sit in Pending until capacity appears, so the HPA’s maxReplicas is only meaningful if the pool underneath can hold that many pods. Multiply the maximum replica count by each pod’s requests and compare the result with the free capacity of your worker pool.

On Kube-DC, worker pools are defined per pool (CPU, memory, disk and image) with autoscaling bounds per pool, and the pool draws from your package’s quota. That quota is pooled across everything you run: containers, virtual machines and Managed Clusters share the same vCPU and memory allowance. We saw the consequence in our test, where an existing Managed Cluster and a virtual machine used about 7.5 of the 8 vCPU in a Pro package and left no room for a GPU pod. It is why we sized the test workers at 1 vCPU and 2 GB. Treat your package as the ceiling for every scaling layer, and check remaining headroom before you raise a maximum.

The control plane scales too

A cluster with many pods and frequent scaling puts load on the API server and etcd. Kube-DC’s documentation states that control-plane and etcd resources are right-sized automatically within configured bounds as cluster load grows, with vertical autoscaling on by default. On a self-managed cluster, sizing those components correctly is your job, and etcd is sensitive to starvation, as we explain in our article on managed versus self-hosted Kubernetes.

Common mistakes

  • No resource requests. Without a CPU request, a utilization target cannot be calculated and the HPA does nothing for that metric.
  • Ignoring startup behaviour. A pod that burns CPU while it initializes can trigger extra scale-ups. The Kubernetes documentation recommends a startup probe, or a readiness probe that only passes after the spike, and the controller ignores CPU samples from a pod during its initialization period.
  • Scaling on CPU when CPU is not the bottleneck. The autoscaling/v2 API also supports memory and custom metrics, which suit queue-driven or I/O-bound workloads better.
  • Setting the maximum without checking capacity. As above, replicas beyond what the pool can hold just wait in Pending.
  • Running a single replica in production. A minimum of one means one pod while load is low. Use a higher minimum for anything that must survive a node loss.

Sources

Related reading

Run this test on your own cluster

Every Kube-DC package includes nested Kubernetes clusters, created from your project in a self-service portal, with metrics-server ready and your quota visible so you can plan capacity. Behind them sit dedicated vCPU, Ceph-backed NVMe storage, a dedicated public IPv4 and S3-compatible object storage, with 24/7 email and ticket support and a 4-hour resolution target. Repeat the walkthrough above and see the replicas move for yourself.

Explore Kube-DC Kubernetes