PHP and Kubernetes: Horizontal Scaling of Queue Workers Based on Redis Queue Length
Introduction: Why CPU-Based HPA Is Not Enough for Queue Workers
Horizontal scaling of queue workers is one of the classic challenges in high-load architectures. At first glance, it might seem that the standard Horizontal Pod Autoscaler (HPA) in Kubernetes handles this out of the box: load increases, CPU consumption rises, and replicas are added. In practice, however, this approach works poorly for queue workers.
The problem is that a PHP queue worker consumes CPU not when tasks accumulate, but when they are being processed. If 10,000 tasks have piled up in the Redis queue and only a few workers are running, each worker's CPU will be at 100% — but HPA always lags behind: it first needs to detect high CPU usage, then average the metric over an observation window (typically 1–3 minutes), and only then decide to scale. Meanwhile, the queue keeps growing.
The second scenario is even worse: tasks are lightweight (e.g., sending emails via an external API), and the worker spends 90% of its time waiting for I/O. CPU stays low, HPA doesn't react, and the queue grows for hours. In both cases, CPU is an indirect and lagging indicator.
The correct signal for scaling queue workers is the queue length itself. The more tasks waiting to be processed, the more workers need to be started. This is exactly what KEDA (Kubernetes Event-driven Autoscaling) is designed for — a tool that scales Deployments directly based on metrics from external sources, including Redis. In this article, we'll walk through a production-ready setup of such autoscaling for PHP applications.
Architecture Overview
The solution architecture consists of three key components:
- PHP worker — a process that reads tasks from the Redis queue and executes them. This can be a Laravel Queue Worker (
php artisan queue:work) or a native PHP script with a processing loop. - Redis as the message broker — stores tasks as lists (LIST). Laravel by default uses the
BLPOPcommand for blocking reads from a key likequeues:default. - KEDA — a Kubernetes operator installed in the cluster that adds a new resource type,
ScaledObject. KEDA polls Redis (or another source), retrieves the list length, and uses it to manage the number of Deployment replicas.
The data flow looks like this: a producer (web request, cron job, or another service) pushes tasks into Redis. KEDA periodically (every 30 seconds by default) queries the list length using the LLEN command and calculates the desired replica count: ceil(queue_length / targetValue). Kubernetes then adjusts the number of Deployment Pods, respecting the minReplicaCount and maxReplicaCount constraints.
Preparing the PHP Application: Dockerfile for the Worker
For a production environment, the worker should run in a separate Docker image optimized for a long-running process. Below is an example Dockerfile for a Laravel worker:
FROM php:8.3-cli-alpine
RUN apk add --no-cache \
linux-headers \
$PHPIZE_DEPS \
&& pecl install redis \
&& docker-php-ext-enable redis \
&& pecl install pcntl \
&& docker-php-ext-install pcntl \
&& apk del $PHPIZE_DEPS
WORKDIR /app
COPY --from=composer:2 /usr/bin/composer /usr/bin/composer
COPY composer.json composer.lock ./
RUN composer install --no-dev --no-scripts --prefer-dist --optimize-autoloader
COPY . .
RUN php artisan config:cache \
&& php artisan route:cache \
&& php artisan event:cache
USER nobody
CMD ["php", "artisan", "queue:work", "redis", \
"--queue=default", \
"--sleep=3", \
"--tries=3", \
"--max-time=3600", \
"--stop-when-empty"]
Several important decisions in this Dockerfile:
php:8.3-cli-alpineis used — a minimal image without unnecessary web server components.- The
--stop-when-emptyflag causes the worker to exit when the queue is empty. This is critical for correct scale-down: KEDA reduces replicas, the Pod receivesSIGTERM, the worker finishes the current task and exits. - The
--max-time=3600flag limits the worker's lifetime to one hour, preventing memory leaks in long-running PHP processes. - Running as the
nobodyuser follows the principle of least privilege.
The Redis connection configuration is set via environment variables in config/queue.php:
'redis' => [
'driver' => 'redis',
'connection' => 'default',
'queue' => env('REDIS_QUEUE', 'default'),
'retry_after' => 90,
'block_for' => 5,
],
Installing KEDA in a Kubernetes Cluster
KEDA is installed via Helm — the most convenient method for production:
helm repo add kedacore https://kedacore.github.io/charts
helm repo update
helm install keda kedacore/keda \
--namespace keda \
--create-namespace \
--set prometheus.metricServer.enabled=true \
--set prometheus.operator.enabled=true \
--version 2.14.0
After installation, verify that the KEDA pods are running:
kubectl get pods -n keda
# NAME READY STATUS RESTARTS AGE
# keda-operator-7d8f9b6c4-xk2pq 1/1 Running 0 2m
# keda-operator-metrics-apiserver-xxx 1/1 Running 0 2m
Now let's create the Deployment for PHP workers. Note: you don't need to specify replicas here — KEDA manages that:
apiVersion: apps/v1
kind: Deployment
metadata:
name: queue-worker
namespace: app
spec:
selector:
matchLabels:
app: queue-worker
template:
metadata:
labels:
app: queue-worker
spec:
terminationGracePeriodSeconds: 120
containers:
- name: worker
image: your-registry/php-worker:latest
env:
- name: REDIS_HOST
valueFrom:
secretKeyRef:
name: redis-secret
key: host
- name: REDIS_PASSWORD
valueFrom:
secretKeyRef:
name: redis-secret
key: password
- name: REDIS_QUEUE
value: "default"
resources:
requests:
cpu: "250m"
memory: "256Mi"
limits:
cpu: "1000m"
memory: "512Mi"
lifecycle:
preStop:
exec:
command: ["sh", "-c", "sleep 5"]
The terminationGracePeriodSeconds: 120 parameter gives the worker up to 2 minutes to finish its current task upon receiving a termination signal — this is a key parameter for graceful shutdown.
Configuring the ScaledObject for Redis
Now we create the ScaledObject — the main KEDA resource that links the Deployment to the metrics source:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: queue-worker-scaler
namespace: app
spec:
scaleTargetRef:
name: queue-worker
minReplicaCount: 0
maxReplicaCount: 50
pollingInterval: 15
cooldownPeriod: 60
triggers:
- type: redis
metadata:
address: redis-service.app.svc.cluster.local:6379
listName: queues:default
listLength: "10"
enableTLS: "false"
databaseIndex: "0"
authenticationRef:
name: redis-trigger-auth
---
apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
name: redis-trigger-auth
namespace: app
spec:
secretTargetRef:
- parameter: password
name: redis-secret
key: password
Let's break down the key parameters:
minReplicaCount: 0— KEDA can scale the Deployment down to zero replicas when the queue is empty, saving cluster resources.maxReplicaCount: 50— the upper limit of workers. Set this based on the throughput capacity of Redis and your database.pollingInterval: 15— KEDA checks the queue length every 15 seconds.cooldownPeriod: 60— after the queue reaches zero, KEDA waits 60 seconds before scaling down to zero to avoid flapping.listLength: "10"— the target number of tasks per worker. With 100 tasks in the queue, 10 workers will be started; with 500 tasks, 50 (up to the maximum).
Working with Multiple Queues of Different Priority
Real-world applications often use multiple queues with different priorities: critical, default, and low. It's best to create a separate Deployment and ScaledObject for each priority group with different parameters:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: critical-worker-scaler
namespace: app
spec:
scaleTargetRef:
name: critical-queue-worker
minReplicaCount: 2
maxReplicaCount: 20
pollingInterval: 10
cooldownPeriod: 30
triggers:
- type: redis
metadata:
address: redis-service.app.svc.cluster.local:6379
listName: queues:critical
listLength: "5"
enableTLS: "false"
authenticationRef:
name: redis-trigger-auth
Key differences for the critical queue:
minReplicaCount: 2— always keep at least 2 workers running so critical tasks are processed without cold-start delays.listLength: "5"— more aggressive scaling: one worker per every 5 tasks.pollingInterval: 10— check the queue more frequently.cooldownPeriod: 30— scale down faster after a peak.
For the low-priority queue (low), you can conversely set minReplicaCount: 0, listLength: "50", and cooldownPeriod: 300.
Monitoring: KEDA Metrics, Prometheus, and Grafana
KEDA exports metrics in Prometheus format out of the box. After installation with the prometheus.metricServer.enabled=true flag, metrics are available on port 9022 of the keda-operator-metrics-apiserver service.
Create a ServiceMonitor for the Prometheus Operator:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: keda-metrics
namespace: monitoring
spec:
namespaceSelector:
matchNames:
- keda
selector:
matchLabels:
app: keda-operator-metrics-apiserver
endpoints:
- port: metrics
interval: 30s
path: /metrics
Key metrics for the Grafana dashboard:
keda_scaler_metrics_value— the current value of the metric (Redis queue length).keda_scaled_object_paused— whether the ScaledObject is paused.keda_scaler_active— whether the scaler is active (i.e., tasks are present in the queue).kube_deployment_spec_replicas— the desired number of replicas.kube_deployment_status_replicas_available— the actual number of running replicas.
Example PromQL query for displaying scaling lag:
keda_scaler_metrics_value{scaledObject="queue-worker-scaler"}
/ on() group_left()
kube_deployment_status_replicas_available{deployment="queue-worker"}
This query shows the average number of tasks per active worker — a key SLO indicator.
Testing Autoscaling
Before going to production, load testing is essential. The simplest way is to push tasks directly into Redis:
kubectl run redis-load-test --image=redis:7-alpine --rm -it --restart=Never -- \
sh -c 'for i in $(seq 1 500); do \
redis-cli -h redis-service.app.svc.cluster.local \
RPUSH queues:default "{\"uuid\":\"test-$i\",\"displayName\":\"TestJob\",\"job\":\"Illuminate\\\\Queue\\\\CallQueuedHandler@call\",\"data\":{\"command\":\"test\"}}"; \
done'
Monitor scaling in real time:
watch -n 5 kubectl get pods -n app -l app=queue-worker
Typical behavior with a correct configuration: after pushing 500 tasks with listLength: 10, KEDA should bring up 50 workers within 15–30 seconds (if maxReplicaCount allows). As tasks are processed, the replica count will decrease incrementally, and after cooldownPeriod following queue exhaustion, it will return to zero.
Check the ScaledObject events for diagnostics:
kubectl describe scaledobject queue-worker-scaler -n app
Common Pitfalls: Graceful Shutdown and Task Loss
The most critical aspect of scaling down is graceful shutdown of workers. When Kubernetes deletes a Pod, it sends SIGTERM. If the worker is processing a task at that moment, the task must either complete correctly or be returned to the queue.
Laravel Queue Worker handles SIGTERM correctly starting from version 8.x: upon receiving the signal, the worker finishes the current task and exits. For this to work, you need to ensure:
- The
pcntlextension is installed in PHP (included in the Dockerfile above). terminationGracePeriodSecondsin the Pod spec is greater than the maximum execution time of a single task.- A
sleep 5is added to thepreStophook — this gives kube-proxy time to remove the Pod from the load balancer beforeSIGTERMarrives.
The second issue is flapping (constant scale-up and scale-down with a steady stream of tasks). If tasks arrive at a rate that causes the queue to alternately empty and refill, KEDA will continuously create and destroy Pods. Solutions include:
- Increase
cooldownPeriodto 120–300 seconds for non-critical queues. - Set
minReplicaCount: 1to avoid scaling to zero. - Use the
advanced.horizontalPodAutoscalerConfig.behaviorparameter in the ScaledObject to fine-tune the scale-down speed.
The third issue is tasks in the reserved/processing state. Laravel places a task in a temporary Redis key (queues:default:reserved) during processing. If the worker is killed without a graceful shutdown, the task remains reserved and only returns to the queue after the retry_after timeout expires (90 seconds by default). Make sure your CI/CD pipeline uses the RollingUpdate strategy with correct maxUnavailable: 0 settings during rolling updates.
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 5
maxUnavailable: 0
Conclusion
Horizontal autoscaling of PHP queue workers based on Redis queue length using KEDA is a mature, production-ready solution that has become the standard for high-load PHP applications on Kubernetes in 2026. The PHP + Redis + Docker + KEDA stack delivers precise responses to real load rather than relying on indirect CPU metrics.
Key takeaways from this article:
- CPU-based HPA is not suitable for queue workers — use KEDA with a Redis trigger.
minReplicaCount: 0saves resources but requires careful graceful shutdown configuration.- Separate queues of different priorities into dedicated Deployments with their own ScaledObjects.
- Monitoring via Prometheus and Grafana is mandatory in production: without metrics, you won't see bottlenecks.
- The
--max-timeflag andterminationGracePeriodSecondsare essential parameters for stable operation.
A properly configured event-driven autoscaling setup allows you to handle load spikes dozens of times more efficiently than a statically configured worker pool, while minimizing costs during periods of low activity.
Technologies
Tags
Ruslan Ismailov
Senior Web / Backend Developer. Senior web/backend developer with 9 years of experience. Stack: PHP, Laravel, PostgreSQL, Redis, Docker, Kubernetes, REST, microservices, CI/CD. More about me →