DevOps

Building a Self-Hosted Observability Stack for PHP Microservices: OpenTelemetry, Prometheus, and Loki Without Cloud Dependencies

Ruslan Ismailov Published 14 min read
B

Introduction: The Three Pillars of Observability and Why Self-Hosted Matters in 2026

Modern distributed systems cannot be debugged or scaled blindly. Observability is the ability to understand the internal state of a system from its external signals. The three fundamental pillars of observability are metrics, tracing, and logs. Metrics provide an aggregated view of performance, traces let you follow a request's path through a chain of services, and structured logs reveal the details of every event.

In 2026, the self-hosted approach is more relevant than ever: cloud solutions like Datadog or New Relic cost tens of thousands of dollars per year at high data volumes, and data sovereignty requirements along with GDPR/regulatory compliance force companies to keep telemetry within their own perimeter. A custom stack built on OpenTelemetry, Prometheus, and Loki gives you full control, predictable costs, and zero vendor lock-in.

This article is a practical guide for DevOps engineers and PHP backend developers who want to deploy a fully featured observability stack on their own infrastructure.

Stack Architecture Overview

The stack consists of the following components, each solving a specific problem:

  • OpenTelemetry PHP SDK — application instrumentation, generating traces, metrics, and logs.
  • OpenTelemetry Collector — receiving, processing, and routing telemetry. Accepts data via OTLP/gRPC and OTLP/HTTP protocols.
  • Prometheus — storing metrics in time series format, scraping endpoints.
  • Loki — log storage with label-based indexing, compatible with Grafana.
  • Tempo — distributed trace storage, integrated with Grafana for TraceQL queries.
  • Grafana — a unified UI for metrics (Prometheus), logs (Loki), and traces (Tempo) with cross-signal correlation.
  • Alertmanager — routing alerts from Prometheus to Slack, PagerDuty, and email.

The data flow works as follows: the PHP service sends traces and logs to the Collector via OTLP; the Collector exports traces to Tempo, logs to Loki via the Loki exporter, while metrics are scraped by Prometheus directly from the application's /metrics endpoint. Grafana reads from all three sources and lets you navigate from a metric to a trace, and from a trace to logs with a single click.

Instrumenting a PHP Application with the OpenTelemetry PHP SDK

OpenTelemetry provides an official PHP SDK. Install the required packages via Composer:

composer require open-telemetry/sdk \
  open-telemetry/opentelemetry-auto-laravel \
  open-telemetry/exporter-otlp \
  open-telemetry/transport-grpc \
  php-http/guzzle7-adapter

For automatic Laravel instrumentation, install the extension via PECL:

pecl install opentelemetry
# Add to php.ini:
extension=opentelemetry.so

Once installed, the opentelemetry-auto-laravel package automatically creates spans for incoming HTTP requests, Eloquent queries, and queues without any changes to application code.

For manual instrumentation of critical code sections:

<?php

use OpenTelemetry\API\Globals;
use OpenTelemetry\API\Trace\SpanKind;
use OpenTelemetry\API\Trace\StatusCode;

class OrderService
{
    public function processOrder(int $orderId): void
    {
        $tracer = Globals::tracerProvider()->getTracer('order-service');

        $span = $tracer->spanBuilder('processOrder')
            ->setSpanKind(SpanKind::KIND_INTERNAL)
            ->startSpan();

        $scope = $span->activate();

        try {
            $span->setAttribute('order.id', $orderId);
            $span->setAttribute('order.source', 'api');

            // business logic
            $this->chargePayment($orderId);
            $this->sendNotification($orderId);

            $span->setStatus(StatusCode::STATUS_OK);
        } catch (\Throwable $e) {
            $span->recordException($e);
            $span->setStatus(StatusCode::STATUS_ERROR, $e->getMessage());
            throw $e;
        } finally {
            $scope->detach();
            $span->end();
        }
    }
}

Collecting Traces: Laravel Integration and Context Propagation Between Services

Configuring the SDK via environment variables is the recommended approach for Laravel applications, since it requires no code changes when switching environments:

# .env
OTEL_SERVICE_NAME=order-service
OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4318
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_PROPAGATORS=tracecontext,baggage
OTEL_TRACES_SAMPLER=parentbased_traceidratio
OTEL_TRACES_SAMPLER_ARG=0.1

To propagate context between services on outgoing HTTP requests via Guzzle, add a middleware:

<?php

use GuzzleHttp\Client;
use OpenTelemetry\API\Globals;
use OpenTelemetry\Context\Propagation\ArrayAccessGetterSetter;
use OpenTelemetry\SDK\Common\Export\Http\PsrTransportFactory;

class TracingHttpClient
{
    private Client $client;

    public function __construct()
    {
        $this->client = new Client();
    }

    public function get(string $url): string
    {
        $headers = [];
        $propagator = Globals::propagator();
        $propagator->inject($headers, ArrayAccessGetterSetter::getInstance());

        $response = $this->client->get($url, ['headers' => $headers]);
        return $response->getBody()->getContents();
    }
}

The traceparent and tracestate headers (W3C TraceContext standard) are automatically forwarded to downstream services, enabling a unified trace tree across multiple PHP microservices.

PHP Metrics: Custom Counters and Histograms

We export Prometheus metrics using the promphp/prometheus_client_php library or the OpenTelemetry Metrics API. Here we demonstrate the native Prometheus client approach for maximum flexibility:

composer require promphp/prometheus_client_php
<?php

use Prometheus\CollectorRegistry;
use Prometheus\Storage\APCu;
use Prometheus\RenderTextFormat;

class MetricsService
{
    private CollectorRegistry $registry;

    public function __construct()
    {
        $this->registry = new CollectorRegistry(new APCu());
    }

    public function recordHttpRequest(
        string $method,
        string $route,
        int $statusCode,
        float $durationSeconds
    ): void {
        // Request counter
        $counter = $this->registry->getOrRegisterCounter(
            'app',
            'http_requests_total',
            'Total HTTP requests',
            ['method', 'route', 'status_code']
        );
        $counter->inc([$method, $route, (string)$statusCode]);

        // Latency histogram (RED metrics)
        $histogram = $this->registry->getOrRegisterHistogram(
            'app',
            'http_request_duration_seconds',
            'HTTP request duration in seconds',
            ['method', 'route'],
            [0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5]
        );
        $histogram->observe($durationSeconds, [$method, $route]);
    }

    public function renderMetrics(): string
    {
        $renderer = new RenderTextFormat();
        return $renderer->render($this->registry->getMetricFamilySamples());
    }
}

Register a route for Prometheus scraping in Laravel:

<?php
// routes/web.php
Route::get('/metrics', function (MetricsService $metrics) {
    return response($metrics->renderMetrics(), 200)
        ->header('Content-Type', \Prometheus\RenderTextFormat::MIME_TYPE);
})->middleware('throttle:60,1');

Structured Logging in PHP: JSON Logs and Correlation with Trace ID

Structured logs allow Loki to efficiently index and filter data. We configure Monolog with a JSON formatter and inject trace_id and span_id from the active OpenTelemetry context:

<?php
// config/logging.php

use Monolog\Formatter\JsonFormatter;
use Monolog\Handler\StreamHandler;
use OpenTelemetry\API\Globals;

'channels' => [
    'json' => [
        'driver' => 'monolog',
        'handler' => StreamHandler::class,
        'handler_with' => [
            'stream' => 'php://stdout',
            'level' => env('LOG_LEVEL', 'debug'),
        ],
        'formatter' => JsonFormatter::class,
        'tap' => [App\Logging\AddTraceContext::class],
    ],
],
<?php
// app/Logging/AddTraceContext.php

namespace App\Logging;

use Monolog\LogRecord;
use Monolog\Processor\ProcessorInterface;
use OpenTelemetry\API\Globals;

class AddTraceContext implements ProcessorInterface
{
    public function __invoke(LogRecord $record): LogRecord
    {
        $span = Globals::tracerProvider()->getTracer('app')
            ->spanBuilder('noop')->startSpan();
        $spanContext = $span->getContext();
        $span->end();

        // Retrieve the current active span from context
        $currentSpan = \OpenTelemetry\API\Trace\Span::getCurrent();
        $ctx = $currentSpan->getContext();

        return $record->with(extra: array_merge($record->extra, [
            'trace_id' => $ctx->getTraceId(),
            'span_id'  => $ctx->getSpanId(),
            'service'  => env('OTEL_SERVICE_NAME', 'unknown'),
        ]));
    }
}

JSON-formatted logs from container stdout are collected by Promtail or Vector and sent to Loki with labels such as service, level, and env.

Configuring the OpenTelemetry Collector

The Collector is the central element of the pipeline. We configure OTLP reception, processing, and export:

# otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 1s
    send_batch_size: 1024
  resource:
    attributes:
      - key: deployment.environment
        value: production
        action: upsert
  memory_limiter:
    limit_mib: 512
    spike_limit_mib: 128
    check_interval: 5s

exporters:
  otlp/tempo:
    endpoint: tempo:4317
    tls:
      insecure: true
  loki:
    endpoint: http://loki:3100/loki/api/v1/push
    default_labels_enabled:
      exporter: false
      job: true
  prometheus:
    endpoint: 0.0.0.0:8889
    namespace: otelcol

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch, resource]
      exporters: [otlp/tempo]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [prometheus]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [loki]

Docker Compose Configuration for the Full Stack in Local Development

# docker-compose.yml
version: "3.9"

services:
  app:
    build: ./app
    environment:
      OTEL_SERVICE_NAME: order-service
      OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4318
      OTEL_EXPORTER_OTLP_PROTOCOL: http/protobuf
      OTEL_PROPAGATORS: tracecontext,baggage
    depends_on: [otel-collector]

  otel-collector:
    image: otel/opentelemetry-collector-contrib:0.96.0
    volumes:
      - ./otel-collector-config.yaml:/etc/otelcol-contrib/config.yaml
    ports:
      - "4317:4317"
      - "4318:4318"
      - "8889:8889"

  prometheus:
    image: prom/prometheus:v2.51.0
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml
      - prometheus_data:/prometheus
    ports:
      - "9090:9090"

  loki:
    image: grafana/loki:2.9.5
    volumes:
      - ./loki-config.yaml:/etc/loki/local-config.yaml
      - loki_data:/loki
    ports:
      - "3100:3100"

  tempo:
    image: grafana/tempo:2.4.1
    volumes:
      - ./tempo-config.yaml:/etc/tempo.yaml
      - tempo_data:/var/tempo
    ports:
      - "3200:3200"

  grafana:
    image: grafana/grafana:10.3.3
    environment:
      GF_SECURITY_ADMIN_PASSWORD: secret
    volumes:
      - ./grafana/provisioning:/etc/grafana/provisioning
      - grafana_data:/var/lib/grafana
    ports:
      - "3000:3000"
    depends_on: [prometheus, loki, tempo]

  promtail:
    image: grafana/promtail:2.9.5
    volumes:
      - /var/lib/docker/containers:/var/lib/docker/containers:ro
      - ./promtail-config.yaml:/etc/promtail/config.yml

volumes:
  prometheus_data:
  loki_data:
  tempo_data:
  grafana_data:

Deploying the Stack to Kubernetes with Helm Charts

For production deployments in Kubernetes, we use official Helm charts. Add the repositories and install the components:

helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo add grafana https://grafana.github.io/helm-charts
helm repo add open-telemetry https://open-telemetry.github.io/opentelemetry-helm-charts
helm repo update

# Prometheus + Alertmanager
helm upgrade --install prometheus prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace \
  --values prometheus-values.yaml

# Loki (single binary for getting started)
helm upgrade --install loki grafana/loki \
  --namespace monitoring \
  --set loki.commonConfig.replication_factor=1 \
  --set loki.storage.type=filesystem

# Tempo
helm upgrade --install tempo grafana/tempo \
  --namespace monitoring

# OpenTelemetry Collector as a DaemonSet
helm upgrade --install otel-collector open-telemetry/opentelemetry-collector \
  --namespace monitoring \
  --values otel-collector-values.yaml

For Prometheus and Loki data persistence, always configure a PersistentVolumeClaim:

# prometheus-values.yaml
prometheus:
  prometheusSpec:
    retention: 30d
    storageSpec:
      volumeClaimTemplate:
        spec:
          storageClassName: fast-ssd
          accessModes: [ReadWriteOnce]
          resources:
            requests:
              storage: 100Gi

Grafana Dashboards: RED Metrics and Service Dependency Map

RED metrics (Rate, Errors, Duration) are the standard for service monitoring. In Grafana, we create $service and $route variables based on Prometheus label values and build the following panels:

  • Rate: sum(rate(app_http_requests_total{service="$service"}[5m])) by (route)
  • Error rate: sum(rate(app_http_requests_total{status_code=~"5.."}[5m])) / sum(rate(app_http_requests_total[5m]))
  • Duration P99: histogram_quantile(0.99, sum(rate(app_http_request_duration_seconds_bucket[5m])) by (le, route))

To find slow requests, use a Table panel sorted by P99 latency. The service dependency map is built using the Grafana Service Graph plugin based on span metrics from Tempo — it automatically renders a call graph between PHP microservices.

The key feature: clicking on an anomalous spike in a metrics graph, Grafana offers to navigate to Tempo traces for the same time range. From a trace, you can jump to related logs in Loki using the trace_id. This is the three-pillar observability correlation in action.

Alerting: Prometheus Alertmanager for PHP Services

A typical set of alerting rules for PHP microservices:

# alerts/php-services.yaml
groups:
  - name: php-microservices
    interval: 30s
    rules:
      - alert: HighErrorRate
        expr: |
          sum(rate(app_http_requests_total{status_code=~"5.."}[5m]))
          /
          sum(rate(app_http_requests_total[5m])) > 0.05
        for: 2m
        labels:
          severity: critical
        annotations:
          summary: "High error rate on {{ $labels.service }}"
          description: "Error rate is {{ humanizePercentage $value }} for last 5m"

      - alert: SlowResponseTime
        expr: |
          histogram_quantile(0.95,
            sum(rate(app_http_request_duration_seconds_bucket[5m])) by (le, service)
          ) > 2
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "P95 latency > 2s on {{ $labels.service }}"

      - alert: PHPFPMQueueFull
        expr: phpfpm_listen_queue > 10
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "PHP-FPM queue is filling up on {{ $labels.instance }}"

      - alert: HighMemoryUsage
        expr: |
          process_resident_memory_bytes{job="php-app"} > 512 * 1024 * 1024
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "PHP process memory > 512MB"

Alertmanager routes alerts to the team's Slack channel and to PagerDuty for critical-severity alerts. Configuring silence windows during planned deployments prevents alert fatigue.

Conclusion and Tips for Scaling the Stack

The assembled stack — OpenTelemetry Collector, Prometheus, Loki, Tempo, and Grafana — covers all three pillars of observability for PHP microservices without a single cloud dependency. Here are the key recommendations for further development:

  • Trace sampling: under high load, use tail-based sampling in the OpenTelemetry Collector — keep 100% of error traces and 1–10% of successful ones.
  • Scaling Loki: when log volume exceeds 10 GB/day, switch to Loki in distributed mode with S3-compatible storage (MinIO for self-hosted).
  • Scaling Prometheus: use Thanos or VictoriaMetrics for long-term storage and horizontal scaling.
  • Security: restrict Collector, Prometheus, and Loki endpoints via Kubernetes network policies; use mTLS for OTLP gRPC between services.
  • Auto-discovery: configure Prometheus Operator ServiceMonitors to automatically add new PHP services to scraping without changing configuration.
  • SLO dashboards: use the Grafana SLO plugin or pyrra to automatically calculate error budgets based on collected metrics.

A self-hosted observability stack is an investment in infrastructure maturity. A pipeline set up correctly once dramatically reduces MTTR (Mean Time To Recovery) and gives the team confidence when deploying to production.

Technologies

Tags

Ruslan Ismailov

Senior Web / Backend Developer. Senior web/backend developer with 9 years of experience. Stack: PHP, Laravel, PostgreSQL, Redis, Docker, Kubernetes, REST, microservices, CI/CD. More about me →