[{"Value":"","Discard":false,"Expires":9999999999}] Cloud archivos - CloudArch https://cloudarch.es/category/cloud/ Blog sobre arquitectura en la nube Sun, 08 Feb 2026 19:39:52 +0000 en-US hourly 1 https://wordpress.org/?v=7.1.3 https://cloudarch.es/wp-content/uploads/2024/02/cropped-CloudArch-1-32x32.png Cloud archivos - CloudArch https://cloudarch.es/category/cloud/ 32 32 228797714 OpenSearch for SREs: The Open-Source Observability Powerhouse on Kubernetes https://cloudarch.es/opensearch-for-sres/ https://cloudarch.es/opensearch-for-sres/#respond Sun, 08 Feb 2026 18:54:11 +0000 https://cloudarch.es/?p=767 The Rise of the Open-Source Alternative: OpenSearch for the Modern SRE In the dynamic world of cloud-native infrastructure, robust observability […]

La entrada OpenSearch for SREs: The Open-Source Observability Powerhouse on Kubernetes se publicó primero en CloudArch.

]]>
The Rise of the Open-Source Alternative: OpenSearch for the Modern SRE

In the dynamic world of cloud-native infrastructure, robust observability is not just a nice-to-have; it’s a foundational requirement for any Site Reliability Engineer (SRE). For years, Elasticsearch dominated the landscape for log aggregation and search. However, a significant licensing change by Elastic in early 2021 created a void, prompting AWS to fork the last Apache 2.0 licensed version and launch OpenSearch.

What is OpenSearch? At its core, OpenSearch is a distributed, RESTful search and analytics engine built on Apache Lucene. It’s designed for high-volume data ingestion and rapid querying across massive datasets. Think of it as a specialized database for semi-structured data like logs, metrics, and traces.

Why is its open-source nature important for SREs? The “open” in OpenSearch isn’t just a marketing buzzword; it’s critical for SREs:

  • No Vendor Lock-in: You’re free from proprietary licenses and sudden feature restrictions. This gives you long-term stability and predictability in your tech stack.
  • Community-Driven Development: Bugs are squashed faster, features are added based on broad community needs, and you can even contribute fixes yourself.
  • Auditability & Transparency: You can inspect the source code, understand exactly how your data is being handled, and ensure there are no hidden surprises—crucial for security-conscious environments.
  • Cost Predictability: While managed services exist, you always have the option to self-host without escalating license costs as your data grows.

For an SRE, OpenSearch provides the backbone for what’s often called the “OSD Stack” (OpenSearch, OpenSearch Dashboards, and Data Prepper/Fluent Bit), empowering proactive monitoring, rapid incident response, and deep operational insights.

Core Concepts for the SRE Toolkit

To wield OpenSearch effectively, you need to understand its fundamental building blocks:

  1. Cluster & Nodes:
    • An OpenSearch Cluster is a group of interconnected servers (nodes) that work together to store and search your data.
    • Nodes specialize: some are cluster_manager (formerly “master”) nodes (brains for cluster state), others are data nodes (muscles for storing and searching data), and some can be ingest nodes (for pre-processing data). SREs design clusters with dedicated nodes for stability and scale.
  2. Indices & Documents:
    • An Index is like a logical database table, a collection of related JSON documents. For logs, you typically create time-based indices (e.g., application-logs-2026.02.08).
    • A Document is a single JSON record within an index—e.g., one log line, one metric data point, one trace span.
  3. Shards & Replicas:
    • To handle massive data, OpenSearch horizontally scales using Shards. An index is split into primary shards, distributed across data nodes.
    • Replicas are copies of primary shards. They provide high availability (if a node fails, a replica becomes primary) and scale read operations. SREs fine-tune shard/replica counts for performance and resilience.
  4. OpenSearch Dashboards:
    • The browser-based UI for visualizing, analyzing, and managing your OpenSearch data. It’s where you build dashboards, run queries, and configure alerts.
  5. Index State Management (ISM):
    • An automated policy engine. SREs use ISM to define lifecycle rules for indices, like moving old data to cheaper storage, taking snapshots, or deleting it after a set period. This prevents disk exhaustion and controls costs.

Getting Hands-On: Your Production-Lite OpenSearch Lab on Kubernetes

The best way to understand OpenSearch is to run it. We’ll set up a mini-stack on Minikube (or any Kubernetes cluster) that mirrors a production setup for logs:

  • OpenSearch Cluster: The data store.
  • OpenSearch Dashboards: The UI.
  • Nginx: A sample application generating logs.
  • Fluent Bit: The lightweight log shipper.

All configurations will be declarative, using Helm charts and Kubernetes Jobs, making it fully automated and version-controllable—true SRE style.

Prerequisites:

  • Kubernetes cluster (Minikube recommended for local dev)
  • kubectl installed and configured
  • helm installed

1. Initialize Minikube

Ensure your Minikube has enough resources for OpenSearch:

minikube start --cpus 4 --memory 8192 --driver docker

2. Prepare Helm Repositories

helm repo add opensearch https://opensearch-project.github.io/helm-charts/
helm repo add fluent https://fluent.github.io/helm-charts
helm repo update

3. Deploy OpenSearch & Dashboards (with a Password)

We’ll use a values.yaml file to set up a single-node OpenSearch cluster with a secure admin password and OpenSearch Dashboards.

Save the following as opensearch-values.yaml:

# opensearch-values.yaml
singleNode: true
persistence:
  enabled: false # For a lab, we'll keep it stateless. Set to true for production with PVs.
extraEnvs:
  - name: OPENSEARCH_INITIAL_ADMIN_PASSWORD
    value: YourStrongPassword123! # <<< CHANGE THIS TO A STRONG PASSWORD

Now deploy:

helm install my-os opensearch/opensearch -f opensearch-values.yaml
helm install my-dashboards opensearch/opensearch-dashboards

Note: OpenSearch will take a few minutes to start as it initializes its JVM and security plugin. Keep an eye on kubectl get pods -w.

4. Deploy Sample Nginx App

This is our log source:

kubectl create deployment nginx-server --image=nginx
kubectl expose deployment nginx-server --port=80

5. Deploy Fluent Bit (The Log Shipper)

Fluent Bit will scrape Nginx logs and send them to OpenSearch. We’ll use a fluent-bit-values.yaml to handle the configuration declaratively and resolve common SRE headaches (DNS, auth, mapping issues).

Save the following as fluent-bit-values.yaml:

# fluent-bit-values.yaml
config:
  service: |
    [SERVICE]
        Daemon          Off
        Flush           1
        Log_Level       info
        Parsers_File    parsers.conf
        HTTP_Server     On
        HTTP_Listen     0.0.0.0
        HTTP_Port       2020
        Health_Check    On

  inputs: |
    [INPUT]
        Name           tail
        Path           /var/log/containers/*.log
        multiline.parser docker, cri
        Tag            kube.*
        Mem_Buf_Limit  5MB
        Skip_Long_Lines On

  filters: |
    [FILTER]
        Name                kubernetes
        Match               kube.*
        Kube_URL            https://kubernetes.default.svc:443
        Kube_CA_File        /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
        Kube_Token_File     /var/run/secrets/kubernetes.io/serviceaccount/token
        Kube_Tag_Prefix     kube.var.log.containers.
        Merge_Log           On
        Merge_Log_Key       log_processed
        Keep_Log            Off
        # These help prevent mapping conflicts from inconsistent pod labels
        Labels              Off
        Annotations         Off

  outputs: |
    [OUTPUT]
        Name            es
        Match           *
        Host            my-os-opensearch # This is the Kubernetes Service name for OpenSearch
        Port            9200
        HTTP_User       admin
        HTTP_Passwd     YourStrongPassword123! # <<< USE THE SAME PASSWORD AS ABOVE
        Logstash_Format On
        Logstash_Prefix nginx-logs
        tls             On
        tls.verify      Off
        Suppress_Type_Name On # Crucial for OpenSearch 2.x
        Trace_Error        On # For better debugging in Fluent Bit logs

Deploy Fluent Bit:

helm install fluent-bit fluent/fluent-bit -f fluent-bit-values.yaml
The Log Shipper: Why Fluent Bit?

In a Kubernetes ecosystem, logs are ephemeral—when a pod dies, its logs go with it. To prevent this, we need a DaemonSet that acts as a “log vacuum.” We chose Fluent Bit because it’s the lightweight, high-performance cousin of Fluentd. It has a tiny memory footprint (crucial when you’re running it on every node in a cluster) and handles the “Log Pipeline” in three distinct stages:

  • Input (Tail): It watches the raw .log files created by the Kubernetes container engine on the node’s disk.
  • Filters (Kubernetes): This is the “SRE secret sauce.” Fluent Bit talks to the Kubernetes API to enrich your logs with metadata like the Pod Name, Namespace, and Labels. We’ve also added logic here to “clean” the logs to prevent mapping conflicts.
  • Output (OpenSearch): It batches the logs and ships them securely via HTTPS to our OpenSearch cluster.
The Anatomy of a Fluent Bit Pipeline

When you look at a Fluent Bit configuration, it’s organized into distinct sections. Each one has a specific job in the “Ingestion Lifecycle.” Understanding these is the key to onboarding any new service into OpenSearch.

  • [SERVICE] (The Global Brain): This section defines the engine’s behavior. It controls how often data is “flushed” (sent) to the destination, where the internal logs go, and whether to enable a health-check server. For SREs, this is where we tune performance and monitoring for the log shipper itself.
  • [INPUT] (The Collector): This is the “vacuum cleaner.” It tells Fluent Bit where to get data. In Kubernetes, we usually use the tail input to follow the log files generated by the container runtime (CRI/Docker). You can have multiple inputs—one for system logs, one for app logs, and even one for metrics.
  • [FILTER] (The Processor): This is where the magic happens. Filters allow you to modify data in flight.
    • The Kubernetes Filter is the most popular; it reaches out to the K8s API to tag your logs with pod names and namespaces.
    • You can also use filters to drop sensitive data (PII), parse strings into JSON, or “flatten” complex labels to avoid the mapping conflicts we saw earlier.
  • [OUTPUT] (The Destination): This defines the “Exit” for your data. In our case, it’s the es (Elasticsearch/OpenSearch) plugin. This section handles the connection details, authentication, and index naming conventions.

6. Automate ISM Policy & Index Pattern with a Kubernetes Job

This is where the SRE magic happens. We’ll use a Job to apply our ISM policy (7-day retention) and create the nginx-logs-* index pattern in Dashboards, all via API calls.

Save the following as observability-setup-job.yaml:

# observability-setup-job.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: opensearch-sre-setup-script
data:
  setup.sh: |
    #!/bin/bash
    set -euo pipefail

    OPENSEARCH_USER="admin"
    OPENSEARCH_PASSWORD="YourStrongPassword123!" # <<< USE THE SAME PASSWORD
    DASHBOARDS_URL="http://my-dashboards-opensearch-dashboards.default.svc.cluster.local:5601"
    OPENSEARCH_URL="https://my-os-opensearch.default.svc.cluster.local:9200"

    echo "Waiting for OpenSearch Dashboards to be available..."
    until curl -s -u "$OPENSEARCH_USER:$OPENSEARCH_PASSWORD" -k "$DASHBOARDS_URL/api/status" | grep -q "available"; do
      sleep 5
    done
    echo "OpenSearch Dashboards is available."

    echo "Creating ISM policy: nginx_retention..."
    curl -X PUT "$OPENSEARCH_URL/_plugins/_ism/policies/nginx_retention" \
      -u "$OPENSEARCH_USER:$OPENSEARCH_PASSWORD" -k -H "Content-Type: application/json" \
      -d '{
            "policy": {
              "description": "Delete logs after 7 days",
              "default_state": "hot",
              "states": [
                {
                  "name": "hot",
                  "actions": [],
                  "transitions": [
                    {
                      "state_name": "delete",
                      "conditions": { "min_index_age": "7d" }
                    }
                  ]
                },
                {
                  "name": "delete",
                  "actions": [ { "delete": {} } ],
                  "transitions": []
                }
              ],
              "ism_template": [
                {
                  "index_patterns": ["nginx-logs-*"],
                  "priority": 100
                }
              ]
            }
          }'
    echo "ISM policy created."

    echo "Creating OpenSearch Dashboards Index Pattern: nginx-logs-pattern..."
    curl -X POST "$DASHBOARDS_URL/api/saved_objects/index-pattern/nginx-logs-pattern" \
      -u "$OPENSEARCH_USER:$OPENSEARCH_PASSWORD" -H "osd-xsrf: true" -H "Content-Type: application/json" \
      -d '{
            "attributes": {
              "title": "nginx-logs-*",
              "timeFieldName": "@timestamp"
            }
          }'
    echo "Index pattern created."

---
apiVersion: batch/v1
kind: Job
metadata:
  name: opensearch-observability-setup
spec:
  template:
    spec:
      containers:
      - name: setup-runner
        image: curlimages/curl:latest # A lightweight image with curl and bash
        command: ["bash", "/scripts/setup.sh"]
        volumeMounts:
        - name: setup-script
          mountPath: /scripts
      volumes:
      - name: setup-script
        configMap:
          name: opensearch-sre-setup-script
          defaultMode: 0744 # Make the script executable
      restartPolicy: OnFailure

Apply the setup Job:

kubectl apply -f observability-setup-job.yaml
Understanding the Automation: ISM & Index Patterns

Before we fire off the job, let’s talk about what’s happening under the hood. We aren’t just pushing config; we are defining the lifecycle and visibility of our data.

  • ISM (Index State Management): In a production cluster, logs are a “growing fire.” If you don’t manage them, they will eventually eat your disk and crash your nodes. By defining an ISM Policy, we automate the “Hot-to-Delete” lifecycle. In this setup, we’re telling OpenSearch: “Keep these logs fresh for 7 days, then delete them automatically.” No manual cleanup, no 3 AM disk-space alerts.
  • Index Patterns: While the Index is where the data lives, the Index Pattern is the “Lens” used by OpenSearch Dashboards to see it. By creating this via code, we ensure that as soon as the stack is up, your Discover tab is ready to go. We’re essentially “gluing” all those daily daily nginx-logs-* into one continuous timeline so you can query across multiple days without lifting a finger.

7. Generate Some Nginx Traffic

To see logs, your Nginx server needs visitors:

kubectl run load-gen --image=busybox --restart=Never -- /bin/sh -c "while true; do wget -qO- http://nginx-server; sleep 2; done"

8. Access OpenSearch Dashboards

Port-forward Dashboards to your local machine:

kubectl port-forward svc/my-dashboards-opensearch-dashboards 5601:5601

Then, open https://localhost:5601 in your browser. Log in with admin and your chosen password. Go to the “Discover” tab, select nginx-logs-* in the dropdown, set your time range (e.g., “Last 15 minutes”), and watch your logs flow in!

Conclusion

By following this automated approach, you’ve built a robust, observable, and easily reproducible log aggregation stack with OpenSearch. You’ve tackled critical SRE challenges like security, data ingestion, and lifecycle management—all as code. This hands-on experience forms a solid foundation for further exploration into metrics, traces, and advanced alerting in your cloud-native environments.

What other OpenSearch challenges will you automate next?

See the whole code at my repository

La entrada OpenSearch for SREs: The Open-Source Observability Powerhouse on Kubernetes se publicó primero en CloudArch.

]]>
https://cloudarch.es/opensearch-for-sres/feed/ 0 767
Introduction to KEDA: Event-Driven Autoscaling for Kubernetes https://cloudarch.es/introduction-to-keda/ https://cloudarch.es/introduction-to-keda/#respond Sun, 07 Dec 2025 20:12:58 +0000 https://cloudarch.es/?p=752 When building applications in Kubernetes, one of the most important challenges is efficiently managing workloads. You want your services to […]

La entrada Introduction to KEDA: Event-Driven Autoscaling for Kubernetes se publicó primero en CloudArch.

]]>
When building applications in Kubernetes, one of the most important challenges is efficiently managing workloads. You want your services to scale up when demand spikes, and scale down (even to zero) when idle, to save resources and reduce costs. This is especially critical in production environments where traffic can be unpredictable.

This is where KEDA (Kubernetes Event-Driven Autoscaling) comes in. KEDA is a lightweight, open-source component that integrates with Kubernetes to provide event-driven autoscaling. Unlike traditional Horizontal Pod Autoscalers (HPAs) that only scale based on CPU or memory usage, KEDA can scale your workloads based on external metrics or events, such as:

  • Messages in a queue (Redis, RabbitMQ, Azure Service Bus, etc.)
  • Jobs waiting in Kafka topics
  • Custom metrics from Prometheus
  • Database triggers or cloud events

With KEDA, your applications can react instantly to real-world workloads without over-provisioning resources. It allows you to run microservices cost-effectively while maintaining responsiveness.

What we are building

In this guide, we’ll build a realistic event-driven microservice that processes jobs from a Redis queue. The scenario mirrors what many production systems face:

  • A backend service receives tasks (e.g., image processing, notifications, or data ingestion) and pushes them into a Redis queue.
  • Worker pods consume tasks from the queue.
  • KEDA monitors the queue and scales the number of worker pods up or down depending on the number of pending tasks.

By the end of this guide, you will have:

  • A working Minikube cluster with KEDA installed
  • A Redis-backed job queue
  • A worker deployment that automatically scales according to queue length
  • Complete YAML manifests and a test workflow to simulate real production traffic

This project is valuable because it lets you experience a real-world autoscaling scenario. Most production systems have unpredictable workloads, and learning how KEDA responds to events gives you a deep understanding of:

  • Event-driven design patterns
  • How Kubernetes interacts with external triggers
  • Autoscaling strategies beyond CPU/memory metrics

Even if your real application is more complex (with multiple queues or different event sources), this example provides a solid foundation to implement production-ready, cost-efficient, and resilient services.

What you’ll learn:

  • Start a 3‑node Minikube cluster (if not started)
  • Install Helm (if needed)
  • Install KEDA with Helm
  • Deploy Redis (simple single‑replica) and a worker Deployment that consumes jobs
  • Create a ScaledObject that uses the Redis List scaler
  • Create a producer CronJob that pushes items into the Redis list to simulate load
  • Observe scaling behavior and clean up

Prerequisites

Make sure you have the following locally installed and working on your machine:

  • kubectl (compatible with your Minikube Kubernetes version)
  • minikube (we will use it to run a 3‑node local cluster)
  • helm (Helm 3)
  • docker (or your container runtime used by Minikube driver)

If any of those are missing, install them first. For example on Debian/Ubuntu: sudo snap install kubectl –classic (or use your distro’s package manager). I will not assume any particular OS beyond these tools being available.

Start a 3-node Minikube cluster

# Start minikube with 3 nodes (one control-plane + 2 workers), adjust memory/
CPUs if needed
$ minikube start --nodes=3 --driver=docker --memory=4096 --cpus=2

# Verify nodes come up
$ kubectl get nodes

Notes: –nodes=3 creates 1 control-plane and 2 worker nodes by default. If you already have the cluster running, skip the minikube start step. Make sure kubectl context points to your minikube cluster: kubectl config current-context .

Install Helm

This command will install Helm for you in your local machine. Please, refer to this other post if you wanna learn more about it.

# Add KEDA Helm repo and update
helm repo add kedacore https://kedacore.github.io/charts
helm repo update

# Install KEDA into namespace 'keda'
helm install keda kedacore/keda --namespace keda --create-namespace

# Wait for KEDA pods to become ready
kubectl -n keda get pods

Why Helm? Helm is the easiest official way to install KEDA and its CRDs. KEDA installs CRDs that are required for ScaledObject and ScaledJob resources.

What we’re going to deploy (high level)

  • redis-deployment.yaml — a simple Redis single‑pod deployment + service to be able to access it from other pods
  • worker-deployment.yaml — a small Python worker Deployment that polls job-queue (a Redis list) and processes items, initially it will have 0 replics since it will imitate a job that doesn’t need to be running always
  • producer-cronjob.yaml — a CronJob that pushes messages into Redis periodically to generate load
  • scaledobject-redis.yaml — KEDA ScaledObject definition that scales worker based on Redis list length

I’ll explain each file before showing it. All manifests are designed to run in the default namespace for simplicity. You can change namespace fields if you prefer.

Redis Manifest

# redis-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: redis
  labels:
    app: redis
spec:
  replicas: 1
  selector:
    matchLabels:
      app: redis
  template:
    metadata:
      labels:
        app: redis
    spec:
      containers:
      - name: redis
        image: redis:7.0-alpine
        ports:
        - containerPort: 6379
        resources:
          requests:
            cpu: "100m"
            memory: "128Mi"
          limits:
            cpu: "250m"
            memory: "256Mi"
---
apiVersion: v1
kind: Service
metadata:
  name: redis
  labels:
    app: redis
spec:
  ports:
  - port: 6379
    targetPort: 6379
    protocol: TCP
  selector:
    app: redis
  type: ClusterIP

As we should know already, we are gonna deploy this in our cluster with:

kubectl apply -f redis-deployment.yaml

To be sure the redis service is client, we can test the connectivity with a redis client:

kubectl run -i --tty redis-client --image=redis:7.0-alpine --restart=Never
-- sh
# inside pod shell run: redis-cli -h redis ping
# should reply PONG

Worker code + Deployment

We’ll create a tiny Python worker that continuously polls a Redis list (named job-queue ) with BRPOP and “processes” messages (here, just sleep and echo ). Just to replicate some functionality of reading from Redis. The purpouse of this guide is not the logic of the code itself but it’s realibility on production environments.

In real life your worker would do meaningful job processing.

# worker.py
import time
import os
import redis


REDIS_HOST = os.getenv('REDIS_HOST', 'redis')
REDIS_PORT = int(os.getenv('REDIS_PORT', '6379'))
LIST_NAME = os.getenv('LIST_NAME', 'job-queue')


r = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
print('Worker started, connecting to', REDIS_HOST)


while True:
    try:
        # BRPOP blocks until an item is available
        item = r.brpop(LIST_NAME, timeout=5)
        if item:
            # item is (list_name, value)
            value = item[1]
            print('Processing', value)
            # simulate processing
            time.sleep(2)
            print('Done', value)
        else:
            # nothing to do, sleep to avoid tight loop
            time.sleep(1)
    except Exception as e:
        print('Worker error:', e)
        time.sleep(2)
# Dockerfile.worker
FROM python:3.11-slim
WORKDIR /app
COPY worker.py .
RUN pip install --no-cache-dir redis
CMD ["python","/app/worker.py"]

By default if we create the image, it will be in our local registry and Minikube won’t see it. To solve this problem we can enable an addon called registry which will allow us to create a registry in Minikube.

minikube addons enable registry

We need to do a port-forward to forward the data to the registry:

kubectl port-forward -n kube-system service/registry 5000:80

And now, in other terminal to not terminate our tunnel, we need to build and push the image so it can be used within the cluster:

docker build -t keda-worker:latest -f Dockerfile.worker .
docker tag keda-worker:latest localhost:5000/keda-worker:latest
docker push localhost:5000/keda-worker:latest

And now we need to create the deployment of our service using the brand new image:

# worker-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: keda-worker
  labels:
    app: keda-worker
spec:
  replicas: 0 # start with 0 so we can see KEDA scale up from zero
  selector:
    matchLabels:
      app: keda-worker
  template:
    metadata:
      labels:
        app: keda-worker
    spec:
      containers:
      - name: worker
        image: localhost:5000/keda-worker:latest
        imagePullPolicy: Always
        env:
        - name: REDIS_HOST
          value: "redis"
        - name: REDIS_PORT
          value: "6379"
        - name: LIST_NAME
          value: "job-queue"
        resources:
          requests:
            cpu: "50m"
            memory: "64Mi"
          limits:
            cpu: "200m"
            memory: "256Mi"

Please, pay attention that we stated that we want a total of 0 replicas since the design of this is to only run when Keda allows it, saving us computing time and resources in our cluster.

kubectl apply -f worker-deployment.yaml

Create the KEDA ScaledObject for Redis List

Now that we have Redis ready to have message and our service ready to start reading those messages when it’s needed, we need to create a ScaledObject which will tell our service when it’s time to work, create instances and do its job.

# scaledobject-redis.yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: keda-worker-scaledobject
  labels:
    app: keda-worker
spec:
  scaleTargetRef:
    name: keda-worker
  pollingInterval: 5 # how often KEDA checks Redis (seconds)
  cooldownPeriod: 30 # how long to wait after scale down before next check
  minReplicaCount: 0
  maxReplicaCount: 10
  triggers:
  - type: redis
    metadata:
      address: "redis.default.svc:6379"
      listName: "job-queue"
      listLength: "5" # scale target: number of items -> triggers scale
      activationListLength: "1" # minimum backlog to activate scaling

Fields explained:

  • scaleTargetRef.name — the Deployment the ScaledObject controls (kedaworker).
  • pollingInterval — how often KEDA will query Redis.
  • minReplicaCount /maxReplicaCount — limits for scaling.
  • triggers — an array of trigger definitions. Here we use the redis scaler.
  • listLength is the target backlog size that will cause KEDA to adjust replicas (KEDA converts this into HPA metrics internally).
  • activationListLength prevents scaling up until backlog is at least that value.

Please, pay special attention at triggers.type since that value is unique for the kind of job we are doing. Keda provides us with an extensive list of different triggers we can use depending on what we want to observe in order to scale our services. You can have a complete list at their official site.

Keda and Redis are on different namespaces on our cluster, when specifying the address make sure you are pointing correctly to Redis. If you don’t use default namespace as I did, it will be different than redis.default.svc:6379

kubectl apply -f scaledobject-redis.yaml

Producer Cronjon to generate load

Finally, we are going to create a cronjob that will just generate some load to replicate a real user, and will just push a message to redis every minute so we can emulate the whole flow automatically and don’t hit any button neither waiting for an user.

# producer-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
  name: job-producer
spec:
  schedule: "*/1 * * * *"
  jobTemplate:
    spec:
      template:
        spec:
          restartPolicy: Never
          containers:
          - name: producer
            image: redis:7.0-alpine
            command: [ "sh", "-c" ]
            args:
            - |
              now=$(date +%s)
              echo "Producing job-$now"
              redis-cli -h redis rpush job-queue job-$now
              sleep 1
kubectl apply -f producer-cronjob.yaml

Observe KEDA scaling in action

For this purpose I have splited my terminal into 3 sections, so I can see everything at one glance. First I have to sections of the same size: right and left.

On the right size I will be running the logs of the Scaled Object to see how it’s being triggered everytime it sees there is 1 or more messages in the queue

kubectl logs -f -n keda keda-operator-f948b6c4-ln9ch

And then the left side I will have it splitted into two parts again. In the upper part I will be checking how many messages do I have on the list

watch 'kubectl exec -it deploy/redis -- redis-cli llen job-queue'

And at the botton I will be watching the number of pods from my deplotments

watch 'kubectl get deploy'

There we will be able to see how every minute the list has a new message, a pod of the worker is being created and the message is deleted. Reading all the logs in the right side.

By this way, we could have a service that only will be running when it’s needed. But also, this can be replicated and configured to just scale up or down services based on some triggers to ensure we are always giving the desired availability and reliability.

Clean up resources

kubectl delete -f producer-cronjob.yaml
kubectl delete -f scaledobject-redis.yaml
kubectl delete -f worker-deployment.yaml
kubectl delete -f redis-deployment.yaml
helm uninstall keda -n keda
kubectl delete namespace keda

You can see and download all the code on GitHub https://github.com/JoaquinJimenezGarcia/LearningKeda

La entrada Introduction to KEDA: Event-Driven Autoscaling for Kubernetes se publicó primero en CloudArch.

]]>
https://cloudarch.es/introduction-to-keda/feed/ 0 752
Monitoring Docker with Prometheus: Gain Full Visibility into Your Containers https://cloudarch.es/monitoring-docker-with-prometheus/ https://cloudarch.es/monitoring-docker-with-prometheus/#respond Wed, 03 Sep 2025 16:54:29 +0000 https://cloudarch.es/?p=711 Introduction As organizations increasingly rely on containerized applications, ensuring their performance, stability, and reliability becomes a critical task. Docker makes […]

La entrada Monitoring Docker with Prometheus: Gain Full Visibility into Your Containers se publicó primero en CloudArch.

]]>

Introduction

As organizations increasingly rely on containerized applications, ensuring their performance, stability, and reliability becomes a critical task. Docker makes it easy to package and deploy applications, but without proper monitoring, you may miss vital insights into how your containers are behaving. Issues such as resource exhaustion, unexpected crashes, or networking bottlenecks can quickly escalate if left unnoticed.

This is where Prometheus, an open-source monitoring and alerting toolkit, comes into play. Prometheus is designed to collect, store, and query time-series metrics, making it an excellent fit for containerized environments. When paired with Docker, it gives you the ability to:

  • Track resource usage (CPU, memory, disk, network) of your containers in real time.
  • Detect and troubleshoot performance bottlenecks before they affect users.
  • Enable alerts for abnormal behavior or failures.
  • Provide historical data to help understand trends and plan for scaling.
  • Integrate seamlessly with visualization tools such as Grafana for clear, actionable dashboards.

By setting up Prometheus to monitor Docker, you establish a foundation for observability that not only improves day-to-day operations but also builds confidence in your system’s resilience. In this post, we will walk through how to configure Prometheus to collect Docker metrics and show you how this setup can be the backbone of a reliable monitoring strategy.

Getting our Prometheus up and running

Nowadays, getting a Prometheus instance up and running is simpler than it sounds. Thanks to Docker itself we can get our instance deployed and ready to be queried in seconds. Let’s see an example of a docker-compose file.

# docker-compose.yml
version: "3.8"

services:
  prometheus:
    image: prom/prometheus:latest
    container_name: prometheus
    restart: unless-stopped
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - prometheus_data:/prometheus

volumes:
  prometheus_data:

We can create also a small configuration for starting collecting Prometheus own metrics

# prometheus.yml
global:
  scrape_interval: 15s

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

Now if we run docker compose up -d we will have our instance serving traffic on port 9090

$ docker ps

CONTAINER ID   IMAGE                    COMMAND                  CREATED         STATUS         PORTS                                         NAMES
6416f7540370   prom/prometheus:latest   "/bin/prometheus --c…"   9 seconds ago   Up 9 seconds   0.0.0.0:9090->9090/tcp, [::]:9090->9090/tcp   prometheus

Preparing Docker for exporting the metrics

Now it’s important to understand that we are gonna monitor docker itself, not the applications running with Docker.

Docker has a new feature to expose metrics on Prometheus-like format of the Docker server without being force to expose the whole Docker API, which is a win on security. To do that, we need to enable metrics-addr: on the deamon.json.

Thanks to that, Prometheus will be able to read and store our metrics.

# /etc/docker/daemon.json
{
  "metrics-addr": "0.0.0.0:9323"
}

After adding that, we will need to restart docker.

Now that we are exposing the metrics, we would be able to see them on the specified port and /metrics path

$ curl localhost:9323/metrics

# HELP builder_builds_failed_total Number of failed image builds
# TYPE builder_builds_failed_total counter
builder_builds_failed_total{reason="build_canceled"} 0
builder_builds_failed_total{reason="build_target_not_reachable_error"} 0
builder_builds_failed_total{reason="command_not_supported_error"} 0
builder_builds_failed_total{reason="dockerfile_empty_error"} 0
builder_builds_failed_total{reason="dockerfile_syntax_error"} 0
builder_builds_failed_total{reason="error_processing_commands_error"} 0
builder_builds_failed_total{reason="missing_onbuild_arguments_error"} 0
builder_builds_failed_total{reason="unknown_instruction_error"} 0
# HELP builder_builds_triggered_total Number of triggered image builds
# TYPE builder_builds_triggered_total counter
builder_builds_triggered_total 0
# HELP engine_daemon_container_actions_seconds The number of seconds it takes to process each container action
# TYPE engine_daemon_container_actions_seconds histogram
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.005"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.01"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.025"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.05"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.1"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.25"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.5"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="1"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="2.5"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="5"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="10"} 1

Telling Prometheus where to scrape

Now that we have Prometheus up and running and Docker exporting its metrics, it’s time to tell Prometheus were to look and scrape for the metrics, so later we can navigate through them.

In order to tell Prometheus where are the metrics, we need to modify the prometheus.yml file that we created during the first steps. We need to create a new job and specify the address. Our file should look like this now:

# prometheus.yml

global:
  scrape_interval: 15s

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]
  - job_name: 'DockerStats'
    static_configs:
      - targets: ['172.17.0.1:9323']

Note 1: Even on the curl command we specified the /metrics path, here it’s not needed. By default, Prometheus will scrape on that path.

Note 2: Please, be careful of the target. 127.0.0.1 or localhost won’t work. Metrics are exposed on the docker bridge network which you can get from network interface docker0

ip a | grep docker0
7: docker0: mtu 1500 qdisc noqueue state DOWN group default
inet 172.17.0.1/16 brd 172.17.255.255 scope global docker0

Now we cand go to our Prometheus instance and we will see our job scraping the metrics correctly on /targets path.

Next Steps

That enables us for a vast options to explore our metrics and get to know better the state of our Docker service. As for example, we could use Prometheus as a Dataset on Grafana to visualize all the stats and a more beautiful way, and even create alerts based on thresholds.

La entrada Monitoring Docker with Prometheus: Gain Full Visibility into Your Containers se publicó primero en CloudArch.

]]>
https://cloudarch.es/monitoring-docker-with-prometheus/feed/ 0 711
🩺 Docker HEALTHCHECK: Is Your App Really Alive Inside the Container? https://cloudarch.es/docker-healthcheck-guide/ https://cloudarch.es/docker-healthcheck-guide/#respond Thu, 05 Jun 2025 17:29:30 +0000 https://cloudarch.es/?p=674 Running a container doesn’t mean your app is running fine. It might look like everything’s green from the outside… but […]

La entrada 🩺 Docker HEALTHCHECK: Is Your App Really Alive Inside the Container? se publicó primero en CloudArch.

]]>
Running a container doesn’t mean your app is running fine. It might look like everything’s green from the outside… but inside? Your app could be frozen, stuck, or completely dead 🧊💀

Welcome to the world of Docker HEALTHCHECK — a super underrated feature that can make or break your reliability game. Today we’ll dive into:

✅ Why HEALTHCHECK is essential
⚠ Real risks of skipping it
⚙ How Docker HEALTHCHECK Works
🛠 How to add it to your Dockerfiles
👀 Two practical test cases (healthy vs unhealthy)

❗ Why You Should Care

Let’s be honest. We often celebrate when our container is “up and running” — but that just means the process inside hasn’t crashed. It doesn’t tell us if:

  • The web server is responding 🕸
  • The database is reachable 📉
  • Your app logic is frozen in a loop 🔁

Without a healthcheck, Docker assumes everything is okay. That’s dangerous in production, but also in dev: it gives you a false sense of security.

Healthchecks add real visibility — if your app doesn’t behave as expected, Docker will mark it as unhealthy, and tools like Docker Swarm or Kubernetes can act accordingly (restarts, scaling, etc.).

⚙ How Docker HEALTHCHECK Works — Under the Hood

When you add a HEALTHCHECK instruction in your Dockerfile, you’re telling the Docker engine to periodically run a command inside the container to determine its health status. Here’s how it works step by step:


🧱 1. The HEALTHCHECK Instruction

Example:

HEALTHCHECK --interval=10s --timeout=3s --retries=3 \
CMD curl -f http://localhost:5000/health || exit 1

You’re defining:

OptionMeaning
CMDThe actual command to run inside the container. It must exit with 0 for healthy, non-zero for unhealthy.
--intervalHow often to run the health check (default: 30s).
--timeoutHow long to wait before the command is considered failed (default: 30s).
--retriesNumber of consecutive failures before the container is marked unhealthy (default: 3).

🧠 2. Docker Monitors Using a Background Healthcheck Manager

When you start a container that has a HEALTHCHECK, Docker spawns a lightweight internal timer per container. This timer schedules and executes the CMD at the interval you define.

It’s all handled by the Docker daemon, which adds a health state entry to the container’s metadata.

🧪 3. Exit Codes Determine Health

Docker executes the healthcheck command inside the container, and uses its exit code to decide the result:

Exit CodeMeaning
0Healthy ✅
1Unhealthy ❌
>1Unhealthy ❌

CMD not found or fails to run? Still counts as unhealthy.

Docker tracks the consecutive failures, and once the retry limit is reached, the container is marked as unhealthy.

🔄 4. Status Stored in Container Metadata

You can view this with:

docker inspect --format='{{json .State.Health}}' [container_name] | jq

It shows:

  • Status: starting, healthy, or unhealthy
  • FailingStreak: how many times it failed consecutively
  • Log: recent healthcheck attempts with timestamps and outputs

Docker updates this metadata in real-time, and you can consume it via:

  • CLI (docker ps, docker inspect)
  • Docker Remote API (/containers/id/json)
  • Orchestration tools (like Swarm or Kubernetes)

🪄 5. No Magic, Just Smart Logic

Docker doesn’t inject anything magical into your container. It simply:

  • Executes the given command using the container’s existing binaries (like curl, wget, etc.)
  • Waits for the result
  • Updates internal health state

But this tiny mechanism becomes powerful when combined with:

  • Restart policies (--restart=on-failure)
  • Health-based load balancers (Swarm, K8s, Traefik)
  • Alerting systems (via Docker events or logs)

💡 A Note About “Starting”

After the container boots, healthchecks begin after a default grace period of 0s (can be configured). During this period, the container status shows as:

"Status": "starting"

Once the first successful check is done, status becomes healthy. If it fails N times, it becomes unhealthy.

🚫 What Healthchecks DON’T Do

  • ❌ They do not stop or restart containers by themselves
  • ❌ They don’t directly affect container networking or DNS
  • ❌ They don’t send alerts unless you wire them to an external system

🔧 Adding a HEALTHCHECK to Your Dockerfile

It’s simple! Here’s the syntax:

HEALTHCHECK --interval=10s --timeout=3s --retries=3 CMD curl -f http://localhost:5000/health || exit 1

This checks every 10 seconds if the /health endpoint returns a success. If it fails 3 times in a row, the container becomes unhealthy.

🧪 Let’s Test It in Action

We’ll create two test containers:

✅ Healthy App

This one includes a proper /health endpoint that always returns 200 OK.

Dockerfile:

FROM python:3.11-slim
ENV DEBIAN_FRONTEND=noninteractive
WORKDIR /app
COPY app.py .
# Install curl
RUN apt-get update && \
apt-get install -y curl && \
apt-get clean && \
rm -rf /var/lib/apt/lists/*
RUN pip install flask
EXPOSE 5000
HEALTHCHECK --interval=10s CMD curl -f http://127.0.0.1:5000/health || exit 1
CMD ["python", "app.py"]

app.py:

from flask import Flask
app = Flask(__name__)

@app.route('/')
def home():
return "All good!"

@app.route('/health')
def health():
return "OK", 200

app.run(host="0.0.0.0", port=5000)

👉 Build and run:

docker build -t healthy-app .
docker run -d --name
healthtest healthy-app
docker inspect --format='{{.State.Health.Status}}'
healthtest
$ docker ps
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
55e9b6f148a9 healthcheck_test "python app.py" 18 seconds ago Up 18 seconds (healthy) 0.0.0.0:5000->5000/tcp, [::]:5000->5000/tcp healthtest

🎉 You’ll get: healthy

❌ Unhealthy App

Now let’s break the /health endpoint.

Modified app.py:

@app.route('/health')
def health():
return "Error", 500

Build and run again:

docker build -t unhealthy-app .
docker run -d --name broken unhealthy-app
docker inspect --format='{{.State.Health.Status}}' broken

💥 Result: unhealthy

You’ll also see the logs showing failed healthcheck attempts:

$ docker inspect broken | jq '.[].State.Health.Log'
[
{
"Start": "2025-06-05T19:26:07.323262742+02:00",
"End": "2025-06-05T19:26:07.367595028+02:00",
"ExitCode": 1,
"Output": " % Total % Received % Xferd Average Speed Time Time Time Current\n Dload Upload Total Spent Left Speed\n\r 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\r 0 5 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\ncurl: (22) The requested URL returned error: 500\n"
},
{
"Start": "2025-06-05T19:26:17.369661511+02:00",
"End": "2025-06-05T19:26:17.408770486+02:00",
"ExitCode": 1,
"Output": " % Total % Received % Xferd Average Speed Time Time Time Current\n Dload Upload Total Spent Left Speed\n\r 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\r 0 5 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\ncurl: (22) The requested URL returned error: 500\n"
},
{
"Start": "2025-06-05T19:26:27.409488914+02:00",
"End": "2025-06-05T19:26:27.450101106+02:00",
"ExitCode": 1,
"Output": " % Total % Received % Xferd Average Speed Time Time Time Current\n Dload Upload Total Spent Left Speed\n\r 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\r 0 5 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\ncurl: (22) The requested URL returned error: 500\n"
},
{
"Start": "2025-06-05T19:26:37.450803223+02:00",
"End": "2025-06-05T19:26:37.492805511+02:00",
"ExitCode": 1,
"Output": " % Total % Received % Xferd Average Speed Time Time Time Current\n Dload Upload Total Spent Left Speed\n\r 0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\r 0 5 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0\ncurl: (22) The requested URL returned error: 500\n"
}
]

👁 What’s the Impact?

ScenarioBehavior
No HEALTHCHECKDocker marks container as healthy by default
HEALTHCHECK passesContainer state = healthy ✅
HEALTHCHECK failsContainer state = unhealthy 🚨

Why it matters:

  • Your orchestration tools (Swarm, Kubernetes, etc.) rely on this signal
  • You can detect failing containers early during development
  • It helps your CI/CD pipeline make smart decisions

🧠 Final Thoughts

A HEALTHCHECK is like a pulse check for your app ❤️‍🩹
Just because a container runs doesn’t mean your service is okay.

Whether you’re in local development or scaling in production, a tiny HEALTHCHECK line in your Dockerfile can save you hours of debugging and nights of firefighting.

So go ahead — make your containers honest.

Docker HEALTHCHECK is:

  • A built-in mechanism that runs periodic commands inside containers
  • Based entirely on the exit status of your script or command
  • Tracked by the Docker daemon, with results exposed via CLI & API
  • Powerful when combined with orchestration, restarts, and alerts

📚 Bonus tip: Want to auto-restart unhealthy containers?
Add this when running your container:

docker run --restart=on-failure ...

La entrada 🩺 Docker HEALTHCHECK: Is Your App Really Alive Inside the Container? se publicó primero en CloudArch.

]]>
https://cloudarch.es/docker-healthcheck-guide/feed/ 0 674
🐳 Docker Monitoring: Keeping an Eye on Your Containers from the Start https://cloudarch.es/docker-monitoring-keeping-an-eye-on-your-containers-from-the-start/ https://cloudarch.es/docker-monitoring-keeping-an-eye-on-your-containers-from-the-start/#respond Wed, 04 Jun 2025 20:51:50 +0000 https://cloudarch.es/?p=670 Monitoring Docker containers is crucial—not just in production, but right from the development phase. Early monitoring helps catch performance issues, […]

La entrada 🐳 Docker Monitoring: Keeping an Eye on Your Containers from the Start se publicó primero en CloudArch.

]]>
Monitoring Docker containers is crucial—not just in production, but right from the development phase. Early monitoring helps catch performance issues, resource bottlenecks, and unexpected behaviors before they escalate. While tools like Grafana and Prometheus are powerful, they might be overkill for smaller projects or initial stages. Let’s explore some lightweight alternatives to keep your containers in check without the overhead.

🛠 The Power of Built-in Tools

Docker provides built-in commands that offer valuable insights into container performance:

docker stats

This command displays real-time metrics for your running containers, including CPU usage, memoryconsumption, and network I/O.

docker logs


Access the logs of a container to monitor its output and diagnose issues.


These commands are straightforward and require no additional setup, making them ideal for quick checks during development.

🧪 Testing with docker.io/spkane/train-os:latest

To see these tools in action, let’s use the docker.io/spkane/train-os:latest image, which simulates system stress and is perfect for testing monitoring setups.

Run the container:

$ docker container run --rm -d --name stress docker.io/spkane/train-os:latest stress -v --cpu 2 --io 1 --vm 2 --vm-bytes 128M --timeout 60s

Unable to find image 'spkane/train-os:latest' locally
latest: Pulling from spkane/train-os
d4df0db66c89: Pull complete
19c5d5a1e2b2: Pull complete
2b25593057c7: Pull complete
0355d914b0bb: Pull complete
Digest: sha256:5acc35b4325d348c8ce6843f6751f62de6e83e518f94f5abe29d0f3ac0fb54be
Status: Downloaded newer image for spkane/train-os:latest
45e7f21918af3000a67d8f78bdfc6601d059160af9429304fca616b75e6036ac

Monitor with docker stats:

$ docker container stats stress --no-stream
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS
b75e1302b035 stress 429.59% 119.9MiB / 31.07GiB 0.38% 5.82kB / 126B 0B / 0B 6

You’ll observe metrics like CPU and memory usage updating in real-time. We are using `–no-stream` to just have a brief output of the current state. Otherwise, it will be running being updated the values each few seconds.

View logs:

$ docker logs stress
stress: info: [1] dispatching hogs: 2 cpu, 1 io, 2 vm, 0 hdd
stress: dbug: [1] using backoff sleep of 15000us
stress: dbug: [1] setting timeout to 60s
stress: dbug: [1] --> hogcpu worker 2 [7] forked
stress: dbug: [1] --> hogio worker 1 [8] forked
stress: dbug: [1] --> hogvm worker 2 [9] forked

🔍 Accessing Metrics via Docker API

For more advanced monitoring or integration with custom tools, you can access container stats directly through the Docker API:

$ curl --no-buffer -X GET --unix-socket /var/run/docker.sock http://docker/containers/stress/stats | head -n 1 | jq

% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0{
"name": "/stress",
"id": "54370040079f3f7c3c6fd8608968050569b1141412c255949a0a72161f4a326a",
"read": "2025-06-04T19:45:21.220812324Z",
"preread": "0001-01-01T00:00:00Z",
"pids_stats": {
"current": 6,
"limit": 37968
},
"blkio_stats": {
"io_service_bytes_recursive": [
{
"major": 259,
"minor": 0,
"op": "read",
"value": 0
},
{
"major": 259,
"minor": 0,
"op": "write",
"value": 0
}
],

This command fetches real-time statistics for the stress container in JSON format, which can be parsed and utilized by various monitoring solutions. That helps us to build our own monitoring solutions too, so if we run a budget environment or want to have control over all our stack we can easily control how our containers behave.

Note that curl is not making a TCP/IP call, we are directly hearing over the unix socket exposed for docker --unix-socket /var/run/docker.sock. That socket exports throught the Docker API /stats/ all needed parameters.

🧭 Choosing the Right Monitoring Approach

ScenarioRecommended Approach
Development & Testingdocker stats and docker logs
Custom IntegrationsDocker API via curl
Production & Large DeploymentsGrafana, Prometheus, etc.

For small-scale applications or during development, Docker’s built-in tools are often sufficient. They provide immediate insights without the complexity of setting up external monitoring systems. However, as your application scales, integrating more robust solutions like Grafana and Prometheus becomes beneficial for long-term monitoring and alerting.

🚀 Conclusion

Monitoring doesn’t have to be complex. Starting with Docker’s native tools allows for quick and effective oversight of your containers. As your needs grow, you can seamlessly transition to more comprehensive solutions. Remember, the key is to implement monitoring early to ensure smooth and efficient container operations.

La entrada 🐳 Docker Monitoring: Keeping an Eye on Your Containers from the Start se publicó primero en CloudArch.

]]>
https://cloudarch.es/docker-monitoring-keeping-an-eye-on-your-containers-from-the-start/feed/ 0 670
What is Helm? The Kubernetes Package Manager Explained for Beginners 🚀 https://cloudarch.es/what-is-helm-the-kubernetes-package-manager/ https://cloudarch.es/what-is-helm-the-kubernetes-package-manager/#respond Tue, 08 Apr 2025 10:30:27 +0000 https://cloudarch.es/?p=650 Welcome to the first post in our Helm series! Whether you’re a DevOps engineer, cloud enthusiast, or just getting started […]

La entrada What is Helm? The Kubernetes Package Manager Explained for Beginners 🚀 se publicó primero en CloudArch.

]]>

Welcome to the first post in our Helm series! Whether you’re a DevOps engineer, cloud enthusiast, or just getting started with Kubernetes, you’re about to discover a powerful tool that can make your life way easier: Helm.

In this post, we’ll dive into:

  • What is Helm?
  • What problems does it solve?
  • Key benefits of using Helm
  • How to install Helm (in just a few steps!)

Let’s jump right in!


What is Helm? 🧭

Imagine managing a complex application in Kubernetes — dozens of YAML files, multiple services, deployments, secrets, configs, and everything in between. It’s like assembling IKEA furniture without the manual.

Helm is here to be your manual.

It’s the package manager for Kubernetes — just like apt is for Ubuntu or yum is for CentOS, but made for the cloud-native world.

With Helm, you can:

  • Package your Kubernetes YAMLs into reusable templates (called Charts)
  • Deploy complex applications with a single command
  • Easily manage upgrades, rollbacks, and configurations

In short: Helm makes deploying to Kubernetes faster, simpler, and less error-prone.


What Problems Does Helm Solve? ⚠

Kubernetes is powerful, but managing it manually is like juggling flaming swords. Here are some real headaches Helm helps with:

1. Too many YAML files

Applications often require multiple resources: Deployments, Services, ConfigMaps, Ingresses, etc. Helm bundles them into one chart.

2. Hard-coded values

Editing the same YAML files over and over to change values (like image tags or environment settings)? Helm allows dynamic values with templates and variables.

3. No easy way to upgrade or rollback

With Helm, you can upgrade an app version and roll back instantly if something goes wrong.

4. App reuse across environments

Want the same app in dev, staging, and prod with just different configs? Helm makes it effortless.


Benefits of Using Helm ✨

Helm isn’t just a nice-to-have — it’s a must-have for scalable, maintainable Kubernetes environments.

Here’s why:

• Saves Time ⏱

Automate and simplify deployments with one-liners.

• Improves Consistency

Avoid human error and deploy the same app across environments with confidence.

• Version Control Friendly

Helm charts can live in Git, making your infrastructure-as-code even more powerful.

• Easily Shareable

Share your charts with teammates or open source them for the community.

• Built-in Rollbacks

One bad deployment? Roll it back like it never happened.


How to Install Helm 🛠

Ready to get started? Installing Helm is quick and painless.

Step 1: Download the Helm Binary

For macOS:

brew install helm

For Linux:

curl https://raw.githubusercontent.com/helm/helm/main/scripts/get-helm-3 | bash

For Windows:

Use Chocolatey or Scoop:

choco install kubernetes-helm

Step 2: Verify the Installation

helm version

You should see something like:

version.BuildInfo{Version:"v3.x.x", GitCommit:"...", ...}

Boom — you’re ready to helm your ship! ⛵


What’s Next?

In the next post, we’ll break down what a Helm Chart really is, how it’s structured, and how to create your very first one.

If you’re enjoying this series, don’t forget to:

  • Subscribe to the blog
  • Share this post with your DevOps squad
  • Leave a comment if you have questions or want a topic covered!

Until next time — keep it cloud-native! ☁


La entrada What is Helm? The Kubernetes Package Manager Explained for Beginners 🚀 se publicó primero en CloudArch.

]]>
https://cloudarch.es/what-is-helm-the-kubernetes-package-manager/feed/ 0 650
Improve docker build speed https://cloudarch.es/improve-docker-build-speed/ https://cloudarch.es/improve-docker-build-speed/#comments Thu, 29 Aug 2024 19:19:33 +0000 https://cloudarch.es/?p=589 Nowadays we have powerfull machines which can compile large and heavy files in record time. However is not always the […]

La entrada Improve docker build speed se publicó primero en CloudArch.

]]>
Nowadays we have powerfull machines which can compile large and heavy files in record time. However is not always the case when we have resources enough and we need to dedicate just a few for compiling our Docker images. That’s why in this post we are gonna cover a few topics to improve docker build speed, and make it faster.

Taking into consideration that every Docker image is based into layers which can be cache’d for later re-use them and decrease the build time, we need to remember that if we modify a layer in our Dockerfile, the consequent layers will need to be rebuilt again, increasing the time. That’s why we can have two approaches.

Writing the more dynamic layers at the end

To improve docker build speed, we need to make sure we can reuse the maximum number of layers cache’d, so the rebuild time will be the minimul. For example, if we need to install some dependencies or updates some packeges, the latest we do it, the better. Look at the following example:

FROM python:3.11

RUN pip install gunicorn
COPY . /app
WORKDIR /app
RUN pip install -r requirements.txt
ENV PORT 8080

CMD exec gunicorn --bind :$PORT --workers 1 --threads 8 app:app

In that case, we are installing a dependency, doing some steps, installing the requirements and then setting up an environment variable. Some of those layers will be rebuild without a real need to do so, like for example, the ENV PORT 8080 that’s a layer that will be rebuild. In this concrete case is only one and very light, but in bigger Dockerfiles it could mean much more.

A way to force more layers to be reused, could be moving the installation to the very end, right before executing CMD. In that way, we are making sure we are re-using most of the layers from time to time and saving time.

FROM python:3.11

COPY . /app
WORKDIR /app

RUN pip install gunicorn
RUN pip install -r requirements.txt
ENV PORT 8080

CMD exec gunicorn --bind :$PORT --workers 1 --threads 8 app:app

Now we would be saving some compiling time.

Directory Caching

By doing directory caching, Docker is able to “remember” where something is located and use it in case it exists to build faster the image. It can bind a special layer in build time with your requirements and unmount it before the snapshot is made. This is often used to handle directories where tools like Linux software installers (apt, apk, dnf, etc) and language dependency managers (npm, bundler, pip, etc) are.

To use this feature, we need to make sure that we have buildkit enabled by running export DOCKER_BUILDKIT=1

Following the previous Dockerfile, if we add a new dependency to the requirements.txt file, it will need to download again all the dependecies, but why if we already have done that in a previous build? This inefficiency results from the builder seeing that we have made a change that impacts this layer and therefore completely re-creates the layer, so we lose that cache, even though we had it stored in the image layer.

In order to improve this situation and save time during compilation, we can mount the directory where pip typically saves the dependencies and mount it as a cache layer to make Docker look at that directory before downloading anything and only download the new dependencies. To do that we must add a little modification to the pip install line. But also we need to add a new line at the top of the file.

#syntax=docker/dockerfile:1

FROM python:3.11

COPY . /app
WORKDIR /app

RUN pip install gunicorn

RUN --mount=type=cache,target=/root/.cache pip install -r requirements.txt

ENV PORT 8080

CMD exec gunicorn --bind :$PORT --workers 1 --threads 8 app:app

The first header line we added is needed to tell Docker we are going to use a newer version of the Dockerfile frontent, which provides us with access to BuildKit’s new features. However the second line is actually mounting a caching layer into the container at /root/.cache for the duration of this step to build. It will remove the content of the directory from the resulting image.

In this way, it will check in consecutive builds which dependencies are already installed to save us some time to download again the same.


Did you like what you read? Don’t forget to read more articles about Docker in our blog.

La entrada Improve docker build speed se publicó primero en CloudArch.

]]>
https://cloudarch.es/improve-docker-build-speed/feed/ 1 589
Make our Docker images lighter https://cloudarch.es/make-our-docker-images-lighter/ https://cloudarch.es/make-our-docker-images-lighter/#comments Sun, 18 Aug 2024 10:12:59 +0000 https://cloudarch.es/?p=581 Frequently, after we create our Docker image using Dockerfile, we get images heavier than we expected and this ends with […]

La entrada Make our Docker images lighter se publicó primero en CloudArch.

]]>
Frequently, after we create our Docker image using Dockerfile, we get images heavier than we expected and this ends with fulling our file system and storage. So, how can we make our Docker containers images lighter?

We are not going to define and talk about what exactly are the layers but as a very basic notion, Docker images are built by layers and these layers are the instructions written in the Docker file. Every instruction (line in a Dockerfile) is a layer and it has some weight depending on what’s doing.

# One layer, which also contains layers since it's an image
FROM docker.io/fedora
# Another layer
RUN dnf install -y httpd
# Another layer
CMD ["/usr/sbin/httpd", "-DFOREGROUND"]

Layers are immutable and additive, meaning that once a Layer has done its job, it cannot be deleted or modified.

The main goal here would be to use a very minimal base image, that only contains what we need. Nothing else. Usually you won’t find a base image that contains everything, so probably you may end up by adding some extra layers like installing some packages like in the previous example.

Optimizing the layers to make our Docker images lighter

If we build the previous image, we would see how Docker is building layer by layer and storing in into the cache for its previous usage.

terrorsys@Andromeda:~/TestLayers$ docker build -t testlayers .
[+] Building 39.3s (6/6) FINISHED                                                         docker:default
 => [internal] load build definition from Dockerfile                                                0.0s
 => => transferring dockerfile: 124B                                                                0.0s
 => [internal] load metadata for docker.io/library/fedora:latest                                    3.2s
 => [internal] load .dockerignore                                                                   0.0s
 => => transferring context: 2B                                                                     0.0s
 => [1/2] FROM docker.io/library/fedora:latest@sha256:5ce8497aeea599bf6b54ab3979133923d82aaa4f6ca5  4.7s
 => => resolve docker.io/library/fedora:latest@sha256:5ce8497aeea599bf6b54ab3979133923d82aaa4f6ca5  0.0s
 => => sha256:5e22da79803c567fceb0e255f1168977259525a4279cb518016a60df025412fb 2.00kB / 2.00kB      0.0s
 => => sha256:d4df0db66c89d7e6225ce9d3597a045fb95c020f3174af1830df88a37a871db8 80.12MB / 80.12MB    3.7s
 => => sha256:5ce8497aeea599bf6b54ab3979133923d82aaa4f6ca5ced1812611b197c79eb0 1.25kB / 1.25kB      0.0s
 => => sha256:a0f4dffd30e0af6e53f57533e79a9e32699d37d8e850132ff89f612d6ea8a300 529B / 529B          0.0s
 => => extracting sha256:d4df0db66c89d7e6225ce9d3597a045fb95c020f3174af1830df88a37a871db8           1.0s
 => [2/2] RUN dnf install -y httpd                                                                 30.5s
 => exporting to image                                                                              0.7s
 => => exporting layers                                                                             0.7s
 => => writing image sha256:28bff713c5bc28c097fbae424ce9f4bb87226596736f04f7eb215aa92dbb4078        0.0s 
 => => naming to docker.io/library/testlayers                                                       0.0s 

Finally we can see how the image is built and inspect the layers to see their weight.

terrorsys@Andromeda:~/TestLayers$ docker images | grep testlayer
testlayers                    latest    28bff713c5bc   8 seconds ago    386MB


terrorsys@Andromeda:~/TestLayers$ docker image history testlayers
IMAGE          CREATED          CREATED BY                                      SIZE      COMMENT
28bff713c5bc   16 seconds ago   CMD ["/usr/sbin/httpd" "-DFOREGROUND"]          0B        buildkit.dockerfile.v0
<missing>      16 seconds ago   RUN /bin/sh -c dnf install -y httpd # buildk…   164MB     buildkit.dockerfile.v0
<missing>      3 months ago     /bin/sh -c #(nop)  CMD ["/bin/bash"]            0B        
<missing>      3 months ago     /bin/sh -c #(nop) ADD file:b8701dca3d7c8dad1…   222MB     
<missing>      11 months ago    /bin/sh -c #(nop)  ENV DISTTAG=f40container …   0B        
<missing>      3 years ago      /bin/sh -c #(nop)  LABEL maintainer=Clement …   0B 

Basically, we can see there layers coming from the base image, and the layers we added. But how about if we decrease the size of it to make the image lighter?

As we can see, since we are using “dnf install”, we are also caching so many other data we don’t really need. Accordingly on Fedora configuration, we can simply remove it by running “dnf clean all”.

However, as we saw recently, layers are additive and immutable, meaning that once the layer has been written and cached, we cannot modify it. It won’t matter if we add a new layer with that step, it won’t make any change.

Nevertheless, we can squash these three steps in an unique layer, so when the step finishes and save on cache, the size will be much lower. We can do that by modifying our Dockerfile.

FROM docker.io/fedora
RUN dnf install -y httpd && \
        dnf clean all
CMD ["/usr/sbin/httpd", "-DFOREGROUND"]

Basically we can use “&&” to say “run this second command only if the first one was successfully done” and then by “\” we indicate a line jump, just to see our Dockerfile much more clear.

Once we have created again our image, we can see now that the size it’s 303 MB instead of 386 MB and the layer went from 164 MB to 81.3 MB. This may not seem like a huge improvement but take into consideration this two scenarios:

  • If we have multiple layers and we clean all of them, the final result will be a much lighter image. In a production environment where we cannot add space as we need, this is crucial if we multiply this for many different images. We can save our disk.
  • In environments where the network is congested, we need to reduce the size of our objects to allow a faster traffic.

Using the right base image

Another point here is the base image we are using to create our own. This one runs a Fedora-based image. But since we only want to run Apache, we can find in the Docker registry another lighter base image which already contains Apache: https://hub.docker.com/r/nimmis/alpine-apache This image for example is an Alpine (much lighter than Fedora) which already has the desired package.

So our Dockerfile would look like the following:

FROM nimmis/alpine-apache
CMD ["/usr/sbin/httpd", "-DFOREGROUND"]

Once we compiled it, we can see that the size of our image went from 303 MB to only 20.7 MB. This is much lighter size, which will improve our disk usage and network traffic to pull/push images. It’s ideal for production environments or other critical environments where we prioritize network and disk.

terrorsys@Andromeda:~/TestLayers$ docker images | grep testlayer
testlayers                    latest    311bb5eeeb68   2 years ago    20.7MB

If you liked this post, don’t forget to comment and read other posts regarding Docker.

La entrada Make our Docker images lighter se publicó primero en CloudArch.

]]>
https://cloudarch.es/make-our-docker-images-lighter/feed/ 1 581
How to make Docker images more secure https://cloudarch.es/how-to-make-docker-images-more-secure/ https://cloudarch.es/how-to-make-docker-images-more-secure/#respond Mon, 12 Aug 2024 20:37:24 +0000 https://cloudarch.es/?p=578 Nowadays many of us work in daily basis with Docker, and we create our own Docker images with Dockerfile. However, […]

La entrada How to make Docker images more secure se publicó primero en CloudArch.

]]>
Nowadays many of us work in daily basis with Docker, and we create our own Docker images with Dockerfile. However, do you know how to make Docker images more secure?

When we are writing our Dockerfile, Docker is using by default the root user to run the commands declared to create every layer in our image. Also, other times he copy and paste other Dockerfile templates which implicitly are using the root user by declaring this line:

USER root

Even this is redundant because Docker already uses it as default is not a good practice to keep that user. Instead, we should have our own user created for our specific purpose with only the needed permissions.

Why this is not a good practice

Even the Docker containers have certain level of isolation, we cannot forget Docker containers are still sharing the same kernel with the host, so using the root user in the wrong hands could end in a disaster.

The root user is not intended for ordinary tasks and should not be used for running our apps.

How to make Docker images more secure

The best practice to follow is to create a new user and a new group for our service, and assignt to it the right permissions at system level.

In order to create the user, we can run the following layers on our Dockerfile:

# Create a custom user with UID 1234 and GID 1234
RUN groupadd -g 1234 customgroup && \
    useradd -m -u 1234 -g customgroup customuser
 
# Switch to the custom user
USER customuser

Did you like this post? Don’t forget to read other related posts, leave your comment and ask for more content!


La entrada How to make Docker images more secure se publicó primero en CloudArch.

]]>
https://cloudarch.es/how-to-make-docker-images-more-secure/feed/ 0 578
How to deploy your Kubernetes app – Part II https://cloudarch.es/how-to-deploy-your-kubernetes-app-part-ii/ https://cloudarch.es/how-to-deploy-your-kubernetes-app-part-ii/#comments Sun, 04 Aug 2024 22:42:02 +0000 https://cloudarch.es/?p=569 In a previous post we were learning how to install and configure ArgoCD to learn how to deploy your Kubernetes […]

La entrada How to deploy your Kubernetes app – Part II se publicó primero en CloudArch.

]]>
In a previous post we were learning how to install and configure ArgoCD to learn how to deploy your Kubernetes app.

In this new post we will be seeing how to integrate our first Kubernetes deployment by default and learn the basics to start automating our services and workloads.

Before starting we must have something to deploy. For this lab I have prepared a default deployment that will create 3 pods with a Nginx container in each. You can take a look or using it by this link: https://github.com/JoaquinJimenezGarcia/argocd-deploy-test but in case you want to create your own repo basically it looks like this:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: nginx-deployment
  labels:
    app: nginx
spec:
  replicas: 3
  selector:
    matchLabels:
      app: nginx
  template:
    metadata:
      labels:
        app: nginx
    spec:
      containers:
      - name: nginx
        image: nginx:1.14.2
        ports:
        - containerPort: 80

Now that we have our ArgoCD up and running, we will see that it’s empty. However in the upper left corner we would see a button labeled “New App”. We must click there and a new window will pop up with some info we must complete.

However these are the most important fields we need to fill:

  • Application Name: this is the name that will return our ArgoCD to refeer to our resources. We can assign here the value that we want, but following the name of the deployment I gave it the same one: argocd-deploy-test
  • Project Name: the project we are going to use in ArgoCD, since we only have “default” this is the value it must have, but we can create multiple projects with multiple resources
  • Sync Policy: if we want to sync manually or automatically, we would mark “automatically” since we want to be updated if we update the YAML files on Github
  • Repository URL: where the YAML files live, if you are using directly the one I have created, it’s https://github.com/JoaquinJimenezGarcia/argocd-deploy-test
  • Path: the path to deploy our pods, by default I left it as “.”
  • Cluster URL: in which cluster we want to deploy, since we only have the cluster where ArgoCD is installed, we can leave it as https://kubernetes.default.svc
  • Namespace: the namespace where the pods would be deployed, as a good practice we should have multiple namespaces, but as per this lab we are going to be uing only “default” namespace

Now we have all these fields ready, we need to click on create app and ArgoCD autimatically will go to our repo, read the yaml files and execute them.

Because this is only a small YAML file, we should see just a few seconds later our app completeley deployed

If we want to double check, we can go to our cluster and examine the pods looking to see if the pod names match with the ones are deployed there:

And as we can see, the results matches, so our app is perfectly deployed.

So as we marked the auto sync, now it will be reading our Git each 3 minutes to detect changes and apply them. So if for example we push a PR changing the replicas from 3 to 5, it would be updated automatically without more human interaction.

La entrada How to deploy your Kubernetes app – Part II se publicó primero en CloudArch.

]]>
https://cloudarch.es/how-to-deploy-your-kubernetes-app-part-ii/feed/ 2 569