DevOps

Elasticsearch Index Lifecycle Management (ILM) in 2026: Automated Index Management and Storage Optimization

Ruslan Ismailov Published 12 min read
E

Introduction: The Problem of Index Growth and Manual Management

In a production Elasticsearch environment, indices grow continuously. Application logs, metrics, audit events — all of these generate new documents every second. Without automation, engineers are forced to manually create new indices, migrate old data to cold nodes, delete outdated shards, and monitor disk space balance. At volumes of hundreds of gigabytes per day, this becomes a full-time operational burden.

Index Lifecycle Management (ILM) is Elasticsearch's built-in mechanism that solves this problem declaratively. You define a lifecycle policy once, attach it to an index template or data stream, and the cluster automatically handles data movement, optimization, and deletion. By 2026, ILM has become the de facto standard for any high-load Elasticsearch project.

ILM Concepts: Phases and Their Purpose

An ILM policy describes a sequence of phases that an index passes through during its lifetime. Each phase corresponds to a specific data state — from active writes to archival.

Hot

The active write and frequent read phase. The index resides on fast nodes (typically SSD). This is where rollover is configured — the condition that triggers the creation of a new index. It is the key mechanism for preventing a single index from growing to unmanageable sizes.

Warm

Data is no longer written but is still read frequently enough. In this phase, force merge (merging segments to reduce file count) and shrink (reducing the number of primary shards) are performed. Warm nodes typically use slower, cheaper disks.

Cold

Data is rarely read. Elasticsearch can move the index to cold nodes or apply searchable snapshots — a mechanism where the index is mounted directly from object storage (S3, GCS, Azure Blob) without fully copying data to the node's disk.

Frozen

Maximum resource savings: the index is fully unloaded from memory and loaded only on demand. Response times are significantly higher than in previous phases, but storage costs are minimal. Suitable for data that is needed occasionally — for example, for compliance audits.

Delete

The index is deleted after a specified age or other conditions are met. This is the final phase of a retention policy.

Configuring an ILM Policy via REST API: A Step-by-Step Guide

An ILM policy can be created through Kibana (Stack Management → Index Lifecycle Policies) or directly via the REST API. Let's focus on the API approach, as it is the most reproducible in CI/CD pipelines.

Example policy for application logs with a 90-day retention:

PUT _ilm/policy/app-logs-policy\n{\n  \"policy\": {\n    \"phases\": {\n      \"hot\": {\n        \"min_age\": \"0ms\",\n        \"actions\": {\n          \"rollover\": {\n            \"max_primary_shard_size\": \"50gb\",\n            \"max_age\": \"1d\",\n            \"max_docs\": 50000000\n          },\n          \"set_priority\": {\n            \"priority\": 100\n          }\n        }\n      },\n      \"warm\": {\n        \"min_age\": \"3d\",\n        \"actions\": {\n          \"forcemerge\": {\n            \"max_num_segments\": 1\n          },\n          \"shrink\": {\n            \"number_of_shards\": 1\n          },\n          \"set_priority\": {\n            \"priority\": 50\n          }\n        }\n      },\n      \"cold\": {\n        \"min_age\": \"14d\",\n        \"actions\": {\n          \"searchable_snapshot\": {\n            \"snapshot_repository\": \"my-s3-repository\"\n          },\n          \"set_priority\": {\n            \"priority\": 0\n          }\n        }\n      },\n      \"frozen\": {\n        \"min_age\": \"45d\",\n        \"actions\": {\n          \"searchable_snapshot\": {\n            \"snapshot_repository\": \"my-s3-repository\"\n          }\n        }\n      },\n      \"delete\": {\n        \"min_age\": \"90d\",\n        \"actions\": {\n          \"delete\": {}\n        }\n      }\n    }\n  }\n}

Note: min_age in the warm, cold, and subsequent phases is measured from the moment of rollover (when the new index is created), not from the index creation date itself. This is important when diagnosing delays in phase transitions.

Attaching a Policy to an Index Template and Data Stream

An ILM policy is not applied directly to each index upon creation — it is attached to an index template. All new indices matching the template automatically inherit the policy.

Example of creating a component template with ILM settings:

PUT _component_template/app-logs-settings\n{\n  \"template\": {\n    \"settings\": {\n      \"index.lifecycle.name\": \"app-logs-policy\",\n      \"index.lifecycle.rollover_alias\": \"app-logs\",\n      \"number_of_shards\": 3,\n      \"number_of_replicas\": 1\n    }\n  }\n}

Next, create an index template that references the component template:

PUT _index_template/app-logs-template\n{\n  \"index_patterns\": [\"app-logs-*\"],\n  \"data_stream\": {},\n  \"composed_of\": [\"app-logs-settings\"],\n  \"priority\": 200\n}

When using a data stream (the recommended approach since Elasticsearch 7.9+), the "data_stream": {} field automatically activates the rollover mechanism without the need to manage an alias manually. A data stream always writes to the current backing index, and a new one is created on rollover.

Rollovers: Conditions and Configuration Details

Rollover is the heart of ILM in the hot phase. It creates a new successor index and switches writes to it. Rollover conditions work on an OR basis — any single condition being met is sufficient to trigger it.

  • max_primary_shard_size — maximum primary shard size (e.g., 50gb). The recommended value for most use cases.
  • max_age — maximum index age (e.g., 1d for daily log rotation).
  • max_docs — maximum number of documents. Useful when data volume is uneven.
  • max_size — total index size (all shards). Deprecated in favor of max_primary_shard_size, but still supported.

For classic indices (not data streams), you must manually create the initial index with an alias:

PUT app-logs-000001\n{\n  \"aliases\": {\n    \"app-logs\": {\n      \"is_write_index\": true\n    }\n  }\n}

After that, ILM will automatically create app-logs-000002, app-logs-000003, and so on whenever rollover conditions are met.

Storage Optimization: Force Merge, Shrink, Freeze, and Searchable Snapshots

Force Merge

Elasticsearch stores data in Lucene segments. Each indexed document creates new segments, which are periodically merged by a background process. Force merge compulsorily merges all segments of an index into one (or a specified number), reducing heap memory usage, speeding up search, and decreasing the number of file descriptors. It is only performed on read-only indices.

Shrink

The shrink operation reduces the number of primary shards in an index. If an index was created with 3 shards and data is no longer growing actively, shrink can collapse it to 1 shard, reducing cluster overhead. The new shard count must be a divisor of the original count.

Searchable Snapshots

This is one of the most powerful storage cost optimization tools in 2026. Instead of storing the index on node disks, Elasticsearch mounts it directly from a snapshot repository (S3, GCS). Data is cached locally only when accessed. In the frozen phase, the cache is minimal, making storage costs comparable to raw object storage.

To use searchable snapshots, you need to register a repository:

PUT _snapshot/my-s3-repository\n{\n  \"type\": \"s3\",\n  \"settings\": {\n    \"bucket\": \"my-elasticsearch-snapshots\",\n    \"region\": \"eu-west-1\",\n    \"base_path\": \"ilm-snapshots\"\n  }\n}

Monitoring Policy Status: Explain API and Diagnostics

To check the current lifecycle state of an index, use the Explain API:

GET app-logs-000001/_ilm/explain

The response contains the current phase, action, step, and the time of the last transition. Key fields to analyze: phase, action, step, step_info (contains the error message on failure).

The most common issues and their causes:

  • Index stuck in the warm/cold phase — typically, nodes with the required attribute (data_warm, data_cold) are unavailable or not present in the cluster. Check node allocation settings.
  • Shrink is hanging — the index was not set to read-only before the shrink. ILM does this automatically, but if there are active write requests on the index, the operation will be blocked.
  • Rollover is not triggering — the alias is not configured as is_write_index, or for a data stream, the first backing index was not created by initializing the data stream.
  • Policy is not applied to new indices — the index template has a lower priority than a competing template. Check the priority field in the template.

To view all indices with ILM errors, you can use:

GET */_ilm/explain?only_errors=true

Integrating ILM with Microservice Architecture

In a microservices environment, a common question arises: who is responsible for creating indices and applying policies? The 2026 best practice suggests the following approach.

The index template and ILM policy are created by the platform team (DevOps/SRE) as part of infrastructure-as-code — via Terraform, Ansible, or a GitOps pipeline with Elasticsearch REST API calls. The microservices themselves write data to a data stream (or alias) without any knowledge of sharding and index rotation details.

This separation of responsibilities allows you to:

  • Change the retention policy without redeploying the application.
  • Centrally manage storage costs.
  • Apply different policies for different environments (dev, staging, production) using different templates with different name patterns.

When automatically creating a data stream via the REST API, a single request is sufficient — Elasticsearch will create the first backing index and apply the policy from the template:

PUT _data_stream/app-logs

Practical Use Case: Configuring ILM for Logs with 90-Day Retention

Let's walk through a complete scenario for a typical production application: a microservice generates structured JSON logs that need to be retained for 90 days. The first 3 days — hot access, days 3–14 — warm, days 14–45 — cold (searchable snapshot), days 45–90 — frozen, after that — deletion.

Step 1: Create the policy (see the example above in the REST API section).

Step 2: Create a mapping via a component template:

PUT _component_template/app-logs-mappings\n{\n  \"template\": {\n    \"mappings\": {\n      \"properties\": {\n        \"@timestamp\": { \"type\": \"date\" },\n        \"service\": { \"type\": \"keyword\" },\n        \"level\": { \"type\": \"keyword\" },\n        \"message\": { \"type\": \"text\" },\n        \"trace_id\": { \"type\": \"keyword\" }\n      }\n    }\n  }\n}

Step 3: Create the final index template that combines everything:

PUT _index_template/app-logs-template\n{\n  \"index_patterns\": [\"app-logs\"],\n  \"data_stream\": {},\n  \"composed_of\": [\"app-logs-settings\", \"app-logs-mappings\"],\n  \"priority\": 200,\n  \"_meta\": {\n    \"description\": \"Template for application logs with 90-day retention\"\n  }\n}

Step 4: Activate the data stream:

PUT _data_stream/app-logs

After this, the microservice can write logs to the app-logs index via the standard bulk API. ILM will automatically execute the entire chain of transitions over 90 days.

To verify the setup a few hours after launch:

GET _data_stream/app-logs\nGET app-logs/_ilm/explain

Conclusion

Elasticsearch Index Lifecycle Management in 2026 is not just a convenient feature — it is an essential component of any production architecture built on Elasticsearch. A properly configured ILM policy automatically handles index rotation, storage optimization through force merge and shrink, cost savings through searchable snapshots, and retention compliance through automatic deletion.

Key recommendations for practical implementation: use data streams instead of classic indices with aliases, separate responsibilities between the platform team and microservice developers, regularly monitor policy status via the Explain API, and test phase transitions in a staging environment before applying them to production. The time invested in a proper ILM setup pays off many times over through reduced operational burden and optimized infrastructure costs.

Technologies

Tags

Ruslan Ismailov

Senior Web / Backend Developer. Senior web/backend developer with 9 years of experience. Stack: PHP, Laravel, PostgreSQL, Redis, Docker, Kubernetes, REST, microservices, CI/CD. More about me →