Building a Self-Hosted Observability Stack for PHP Microservices: OpenTelemetry, Prometheus, and Loki Without Cloud Dependencies
Introduction: The Three Pillars of Observability and Why Self-Hosted Matters in 2026
Modern distributed systems cannot be debugged or scaled blindly. Observability is the ability to understand the internal state of a system from its external signals. The three fundamental pillars of observability are metrics, tracing, and logs. Metrics provide an aggregated view of performance, traces let you follow a request's path through a chain of services, and structured logs reveal the details of every event.
In 2026, the self-hosted approach is more relevant than ever: cloud solutions like Datadog or New Relic cost tens of thousands of dollars per year at high data volumes, and data sovereignty requirements along with GDPR/regulatory compliance force companies to keep telemetry within their own perimeter. A custom stack built on OpenTelemetry, Prometheus, and Loki gives you full control, predictable costs, and zero vendor lock-in.
This article is a practical guide for DevOps engineers and PHP backend developers who want to deploy a fully featured observability stack on their own infrastructure.
Stack Architecture Overview
The stack consists of the following components, each solving a specific problem:
- OpenTelemetry PHP SDK — application instrumentation, generating traces, metrics, and logs.
- OpenTelemetry Collector — receiving, processing, and routing telemetry. Accepts data via OTLP/gRPC and OTLP/HTTP protocols.
- Prometheus — storing metrics in time series format, scraping endpoints.
- Loki — log storage with label-based indexing, compatible with Grafana.
- Tempo — distributed trace storage, integrated with Grafana for TraceQL queries.
- Grafana — a unified UI for metrics (Prometheus), logs (Loki), and traces (Tempo) with cross-signal correlation.
- Alertmanager — routing alerts from Prometheus to Slack, PagerDuty, and email.
The data flow works as follows: the PHP service sends traces and logs to the Collector via OTLP; the Collector exports traces to Tempo, logs to Loki via the Loki exporter, while metrics are scraped by Prometheus directly from the application's /metrics endpoint. Grafana reads from all three sources and lets you navigate from a metric to a trace, and from a trace to logs with a single click.
Instrumenting a PHP Application with the OpenTelemetry PHP SDK
OpenTelemetry provides an official PHP SDK. Install the required packages via Composer:
composer require open-telemetry/sdk \
open-telemetry/opentelemetry-auto-laravel \
open-telemetry/exporter-otlp \
open-telemetry/transport-grpc \
php-http/guzzle7-adapter
For automatic Laravel instrumentation, install the extension via PECL:
pecl install opentelemetry
# Add to php.ini:
extension=opentelemetry.so
Once installed, the opentelemetry-auto-laravel package automatically creates spans for incoming HTTP requests, Eloquent queries, and queues without any changes to application code.
For manual instrumentation of critical code sections:
<?php
use OpenTelemetry\API\Globals;
use OpenTelemetry\API\Trace\SpanKind;
use OpenTelemetry\API\Trace\StatusCode;
class OrderService
{
public function processOrder(int $orderId): void
{
$tracer = Globals::tracerProvider()->getTracer('order-service');
$span = $tracer->spanBuilder('processOrder')
->setSpanKind(SpanKind::KIND_INTERNAL)
->startSpan();
$scope = $span->activate();
try {
$span->setAttribute('order.id', $orderId);
$span->setAttribute('order.source', 'api');
// business logic
$this->chargePayment($orderId);
$this->sendNotification($orderId);
$span->setStatus(StatusCode::STATUS_OK);
} catch (\Throwable $e) {
$span->recordException($e);
$span->setStatus(StatusCode::STATUS_ERROR, $e->getMessage());
throw $e;
} finally {
$scope->detach();
$span->end();
}
}
}
Collecting Traces: Laravel Integration and Context Propagation Between Services
Configuring the SDK via environment variables is the recommended approach for Laravel applications, since it requires no code changes when switching environments:
# .env
OTEL_SERVICE_NAME=order-service
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_PROPAGATORS=tracecontext,baggage
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.1
To propagate context between services on outgoing HTTP requests via Guzzle, add a middleware:
<?php
use GuzzleHttp\Client;
use OpenTelemetry\API\Globals;
use OpenTelemetry\Context\Propagation\ArrayAccessGetterSetter;
use OpenTelemetry\SDK\Common\Export\Http\PsrTransportFactory;
class TracingHttpClient
{
private Client $client;
public function __construct()
{
$this->client = new Client();
}
public function get(string $url): string
{
$headers = [];
$propagator = Globals::propagator();
$propagator->inject($headers, ArrayAccessGetterSetter::getInstance());
$response = $this->client->get($url, ['headers' => $headers]);
return $response->getBody()->getContents();
}
}
The traceparent and tracestate headers (W3C TraceContext standard) are automatically forwarded to downstream services, enabling a unified trace tree across multiple PHP microservices.
PHP Metrics: Custom Counters and Histograms
We export Prometheus metrics using the promphp/prometheus_client_php library or the OpenTelemetry Metrics API. Here we demonstrate the native Prometheus client approach for maximum flexibility:
composer require promphp/prometheus_client_php
<?php
use Prometheus\CollectorRegistry;
use Prometheus\Storage\APCu;
use Prometheus\RenderTextFormat;
class MetricsService
{
private CollectorRegistry $registry;
public function __construct()
{
$this->registry = new CollectorRegistry(new APCu());
}
public function recordHttpRequest(
string $method,
string $route,
int $statusCode,
float $durationSeconds
): void {
// Request counter
$counter = $this->registry->getOrRegisterCounter(
'app',
'http_requests_total',
'Total HTTP requests',
['method', 'route', 'status_code']
);
$counter->inc([$method, $route, (string)$statusCode]);
// Latency histogram (RED metrics)
$histogram = $this->registry->getOrRegisterHistogram(
'app',
'http_request_duration_seconds',
'HTTP request duration in seconds',
['method', 'route'],
[0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5]
);
$histogram->observe($durationSeconds, [$method, $route]);
}
public function renderMetrics(): string
{
$renderer = new RenderTextFormat();
return $renderer->render($this->registry->getMetricFamilySamples());
}
}
Register a route for Prometheus scraping in Laravel:
<?php
// routes/web.php
Route::get('/metrics', function (MetricsService $metrics) {
return response($metrics->renderMetrics(), 200)
->header('Content-Type', \Prometheus\RenderTextFormat::MIME_TYPE);
})->middleware('throttle:60,1');
Structured Logging in PHP: JSON Logs and Correlation with Trace ID
Structured logs allow Loki to efficiently index and filter data. We configure Monolog with a JSON formatter and inject trace_id and span_id from the active OpenTelemetry context:
<?php
// config/logging.php
use Monolog\Formatter\JsonFormatter;
use Monolog\Handler\StreamHandler;
use OpenTelemetry\API\Globals;
'channels' => [
'json' => [
'driver' => 'monolog',
'handler' => StreamHandler::class,
'handler_with' => [
'stream' => 'php://stdout',
'level' => env('LOG_LEVEL', 'debug'),
],
'formatter' => JsonFormatter::class,
'tap' => [App\Logging\AddTraceContext::class],
],
],
<?php
// app/Logging/AddTraceContext.php
namespace App\Logging;
use Monolog\LogRecord;
use Monolog\Processor\ProcessorInterface;
use OpenTelemetry\API\Globals;
class AddTraceContext implements ProcessorInterface
{
public function __invoke(LogRecord $record): LogRecord
{
$span = Globals::tracerProvider()->getTracer('app')
->spanBuilder('noop')->startSpan();
$spanContext = $span->getContext();
$span->end();
// Retrieve the current active span from context
$currentSpan = \OpenTelemetry\API\Trace\Span::getCurrent();
$ctx = $currentSpan->getContext();
return $record->with(extra: array_merge($record->extra, [
'trace_id' => $ctx->getTraceId(),
'span_id' => $ctx->getSpanId(),
'service' => env('OTEL_SERVICE_NAME', 'unknown'),
]));
}
}
JSON-formatted logs from container stdout are collected by Promtail or Vector and sent to Loki with labels such as service, level, and env.
Configuring the OpenTelemetry Collector
The Collector is the central element of the pipeline. We configure OTLP reception, processing, and export:
# otel-collector-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
batch:
timeout: 1s
send_batch_size: 1024
resource:
attributes:
- key: deployment.environment
value: production
action: upsert
memory_limiter:
limit_mib: 512
spike_limit_mib: 128
check_interval: 5s
exporters:
otlp/tempo:
endpoint: tempo:4317
tls:
insecure: true
loki:
endpoint: http://loki:3100/loki/api/v1/push
default_labels_enabled:
exporter: false
job: true
prometheus:
endpoint: 0.0.0.0:8889
namespace: otelcol
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch, resource]
exporters: [otlp/tempo]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [prometheus]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loki]
Docker Compose Configuration for the Full Stack in Local Development
# docker-compose.yml
version: "3.9"
services:
app:
build: ./app
environment:
OTEL_SERVICE_NAME: order-service
OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318
OTEL_EXPORTER_OTLP_PROTOCOL: http/protobuf
OTEL_PROPAGATORS: tracecontext,baggage
depends_on: [otel-collector]
otel-collector:
image: otel/opentelemetry-collector-contrib:0.96.0
volumes:
- ./otel-collector-config.yaml:/etc/otelcol-contrib/config.yaml
ports:
- "4317:4317"
- "4318:4318"
- "8889:8889"
prometheus:
image: prom/prometheus:v2.51.0
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- prometheus_data:/prometheus
ports:
- "9090:9090"
loki:
image: grafana/loki:2.9.5
volumes:
- ./loki-config.yaml:/etc/loki/local-config.yaml
- loki_data:/loki
ports:
- "3100:3100"
tempo:
image: grafana/tempo:2.4.1
volumes:
- ./tempo-config.yaml:/etc/tempo.yaml
- tempo_data:/var/tempo
ports:
- "3200:3200"
grafana:
image: grafana/grafana:10.3.3
environment:
GF_SECURITY_ADMIN_PASSWORD: secret
volumes:
- ./grafana/provisioning:/etc/grafana/provisioning
- grafana_data:/var/lib/grafana
ports:
- "3000:3000"
depends_on: [prometheus, loki, tempo]
promtail:
image: grafana/promtail:2.9.5
volumes:
- /var/lib/docker/containers:/var/lib/docker/containers:ro
- ./promtail-config.yaml:/etc/promtail/config.yml
volumes:
prometheus_data:
loki_data:
tempo_data:
grafana_data:
Deploying the Stack to Kubernetes with Helm Charts
For production deployments in Kubernetes, we use official Helm charts. Add the repositories and install the components:
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo add grafana https://grafana.github.io/helm-charts
helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm repo update
# Prometheus + Alertmanager
helm upgrade --install prometheus prometheus-community/kube-prometheus-stack \
--namespace monitoring --create-namespace \
--values prometheus-values.yaml
# Loki (single binary for getting started)
helm upgrade --install loki grafana/loki \
--namespace monitoring \
--set loki.commonConfig.replication_factor=1 \
--set loki.storage.type=filesystem
# Tempo
helm upgrade --install tempo grafana/tempo \
--namespace monitoring
# OpenTelemetry Collector as a DaemonSet
helm upgrade --install otel-collector open-telemetry/opentelemetry-collector \
--namespace monitoring \
--values otel-collector-values.yaml
For Prometheus and Loki data persistence, always configure a PersistentVolumeClaim:
# prometheus-values.yaml
prometheus:
prometheusSpec:
retention: 30d
storageSpec:
volumeClaimTemplate:
spec:
storageClassName: fast-ssd
accessModes: [ReadWriteOnce]
resources:
requests:
storage: 100Gi
Grafana Dashboards: RED Metrics and Service Dependency Map
RED metrics (Rate, Errors, Duration) are the standard for service monitoring. In Grafana, we create $service and $route variables based on Prometheus label values and build the following panels:
- Rate:
sum(rate(app_http_requests_total{service="$service"}[5m])) by (route) - Error rate:
sum(rate(app_http_requests_total{status_code=~"5.."}[5m])) / sum(rate(app_http_requests_total[5m])) - Duration P99:
histogram_quantile(0.99, sum(rate(app_http_request_duration_seconds_bucket[5m])) by (le, route))
To find slow requests, use a Table panel sorted by P99 latency. The service dependency map is built using the Grafana Service Graph plugin based on span metrics from Tempo — it automatically renders a call graph between PHP microservices.
The key feature: clicking on an anomalous spike in a metrics graph, Grafana offers to navigate to Tempo traces for the same time range. From a trace, you can jump to related logs in Loki using the trace_id. This is the three-pillar observability correlation in action.
Alerting: Prometheus Alertmanager for PHP Services
A typical set of alerting rules for PHP microservices:
# alerts/php-services.yaml
groups:
- name: php-microservices
interval: 30s
rules:
- alert: HighErrorRate
expr: |
sum(rate(app_http_requests_total{status_code=~"5.."}[5m]))
/
sum(rate(app_http_requests_total[5m])) > 0.05
for: 2m
labels:
severity: critical
annotations:
summary: "High error rate on {{ $labels.service }}"
description: "Error rate is {{ humanizePercentage $value }} for last 5m"
- alert: SlowResponseTime
expr: |
histogram_quantile(0.95,
sum(rate(app_http_request_duration_seconds_bucket[5m])) by (le, service)
) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "P95 latency > 2s on {{ $labels.service }}"
- alert: PHPFPMQueueFull
expr: phpfpm_listen_queue > 10
for: 1m
labels:
severity: critical
annotations:
summary: "PHP-FPM queue is filling up on {{ $labels.instance }}"
- alert: HighMemoryUsage
expr: |
process_resident_memory_bytes{job="php-app"} > 512 * 1024 * 1024
for: 5m
labels:
severity: warning
annotations:
summary: "PHP process memory > 512MB"
Alertmanager routes alerts to the team's Slack channel and to PagerDuty for critical-severity alerts. Configuring silence windows during planned deployments prevents alert fatigue.
Conclusion and Tips for Scaling the Stack
The assembled stack — OpenTelemetry Collector, Prometheus, Loki, Tempo, and Grafana — covers all three pillars of observability for PHP microservices without a single cloud dependency. Here are the key recommendations for further development:
- Trace sampling: under high load, use tail-based sampling in the OpenTelemetry Collector — keep 100% of error traces and 1–10% of successful ones.
- Scaling Loki: when log volume exceeds 10 GB/day, switch to Loki in distributed mode with S3-compatible storage (MinIO for self-hosted).
- Scaling Prometheus: use Thanos or VictoriaMetrics for long-term storage and horizontal scaling.
- Security: restrict Collector, Prometheus, and Loki endpoints via Kubernetes network policies; use mTLS for OTLP gRPC between services.
- Auto-discovery: configure Prometheus Operator ServiceMonitors to automatically add new PHP services to scraping without changing configuration.
- SLO dashboards: use the Grafana SLO plugin or pyrra to automatically calculate error budgets based on collected metrics.
A self-hosted observability stack is an investment in infrastructure maturity. A pipeline set up correctly once dramatically reduces MTTR (Mean Time To Recovery) and gives the team confidence when deploying to production.
Technologies
Tags
Ruslan Ismailov
Senior Web / Backend Developer. Senior web/backend developer with 9 years of experience. Stack: PHP, Laravel, PostgreSQL, Redis, Docker, Kubernetes, REST, microservices, CI/CD. More about me →