[{"Value":"","Discard":false,"Expires":9999999999}]
La entrada OpenSearch for SREs: The Open-Source Observability Powerhouse on Kubernetes se publicó primero en CloudArch.
]]>In the dynamic world of cloud-native infrastructure, robust observability is not just a nice-to-have; it’s a foundational requirement for any Site Reliability Engineer (SRE). For years, Elasticsearch dominated the landscape for log aggregation and search. However, a significant licensing change by Elastic in early 2021 created a void, prompting AWS to fork the last Apache 2.0 licensed version and launch OpenSearch.
What is OpenSearch? At its core, OpenSearch is a distributed, RESTful search and analytics engine built on Apache Lucene. It’s designed for high-volume data ingestion and rapid querying across massive datasets. Think of it as a specialized database for semi-structured data like logs, metrics, and traces.
Why is its open-source nature important for SREs? The “open” in OpenSearch isn’t just a marketing buzzword; it’s critical for SREs:
For an SRE, OpenSearch provides the backbone for what’s often called the “OSD Stack” (OpenSearch, OpenSearch Dashboards, and Data Prepper/Fluent Bit), empowering proactive monitoring, rapid incident response, and deep operational insights.
To wield OpenSearch effectively, you need to understand its fundamental building blocks:
cluster_manager (formerly “master”) nodes (brains for cluster state), others are data nodes (muscles for storing and searching data), and some can be ingest nodes (for pre-processing data). SREs design clusters with dedicated nodes for stability and scale.application-logs-2026.02.08).The best way to understand OpenSearch is to run it. We’ll set up a mini-stack on Minikube (or any Kubernetes cluster) that mirrors a production setup for logs:
All configurations will be declarative, using Helm charts and Kubernetes Jobs, making it fully automated and version-controllable—true SRE style.
kubectl installed and configuredhelm installedEnsure your Minikube has enough resources for OpenSearch:
minikube start --cpus 4 --memory 8192 --driver dockerhelm repo add opensearch https://opensearch-project.github.io/helm-charts/
helm repo add fluent https://fluent.github.io/helm-charts
helm repo updateWe’ll use a values.yaml file to set up a single-node OpenSearch cluster with a secure admin password and OpenSearch Dashboards.
Save the following as opensearch-values.yaml:
# opensearch-values.yaml
singleNode: true
persistence:
enabled: false # For a lab, we'll keep it stateless. Set to true for production with PVs.
extraEnvs:
- name: OPENSEARCH_INITIAL_ADMIN_PASSWORD
value: YourStrongPassword123! # <<< CHANGE THIS TO A STRONG PASSWORDNow deploy:
helm install my-os opensearch/opensearch -f opensearch-values.yaml
helm install my-dashboards opensearch/opensearch-dashboardsNote: OpenSearch will take a few minutes to start as it initializes its JVM and security plugin. Keep an eye on kubectl get pods -w.
This is our log source:
kubectl create deployment nginx-server --image=nginx
kubectl expose deployment nginx-server --port=80Fluent Bit will scrape Nginx logs and send them to OpenSearch. We’ll use a fluent-bit-values.yaml to handle the configuration declaratively and resolve common SRE headaches (DNS, auth, mapping issues).
Save the following as fluent-bit-values.yaml:
# fluent-bit-values.yaml
config:
service: |
[SERVICE]
Daemon Off
Flush 1
Log_Level info
Parsers_File parsers.conf
HTTP_Server On
HTTP_Listen 0.0.0.0
HTTP_Port 2020
Health_Check On
inputs: |
[INPUT]
Name tail
Path /var/log/containers/*.log
multiline.parser docker, cri
Tag kube.*
Mem_Buf_Limit 5MB
Skip_Long_Lines On
filters: |
[FILTER]
Name kubernetes
Match kube.*
Kube_URL https://kubernetes.default.svc:443
Kube_CA_File /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
Kube_Token_File /var/run/secrets/kubernetes.io/serviceaccount/token
Kube_Tag_Prefix kube.var.log.containers.
Merge_Log On
Merge_Log_Key log_processed
Keep_Log Off
# These help prevent mapping conflicts from inconsistent pod labels
Labels Off
Annotations Off
outputs: |
[OUTPUT]
Name es
Match *
Host my-os-opensearch # This is the Kubernetes Service name for OpenSearch
Port 9200
HTTP_User admin
HTTP_Passwd YourStrongPassword123! # <<< USE THE SAME PASSWORD AS ABOVE
Logstash_Format On
Logstash_Prefix nginx-logs
tls On
tls.verify Off
Suppress_Type_Name On # Crucial for OpenSearch 2.x
Trace_Error On # For better debugging in Fluent Bit logsDeploy Fluent Bit:
helm install fluent-bit fluent/fluent-bit -f fluent-bit-values.yamlIn a Kubernetes ecosystem, logs are ephemeral—when a pod dies, its logs go with it. To prevent this, we need a DaemonSet that acts as a “log vacuum.” We chose Fluent Bit because it’s the lightweight, high-performance cousin of Fluentd. It has a tiny memory footprint (crucial when you’re running it on every node in a cluster) and handles the “Log Pipeline” in three distinct stages:
.log files created by the Kubernetes container engine on the node’s disk.When you look at a Fluent Bit configuration, it’s organized into distinct sections. Each one has a specific job in the “Ingestion Lifecycle.” Understanding these is the key to onboarding any new service into OpenSearch.
[SERVICE] (The Global Brain): This section defines the engine’s behavior. It controls how often data is “flushed” (sent) to the destination, where the internal logs go, and whether to enable a health-check server. For SREs, this is where we tune performance and monitoring for the log shipper itself.[INPUT] (The Collector): This is the “vacuum cleaner.” It tells Fluent Bit where to get data. In Kubernetes, we usually use the tail input to follow the log files generated by the container runtime (CRI/Docker). You can have multiple inputs—one for system logs, one for app logs, and even one for metrics.[FILTER] (The Processor): This is where the magic happens. Filters allow you to modify data in flight.[OUTPUT] (The Destination): This defines the “Exit” for your data. In our case, it’s the es (Elasticsearch/OpenSearch) plugin. This section handles the connection details, authentication, and index naming conventions.This is where the SRE magic happens. We’ll use a Job to apply our ISM policy (7-day retention) and create the nginx-logs-* index pattern in Dashboards, all via API calls.
Save the following as observability-setup-job.yaml:
# observability-setup-job.yaml
apiVersion: v1
kind: ConfigMap
metadata:
name: opensearch-sre-setup-script
data:
setup.sh: |
#!/bin/bash
set -euo pipefail
OPENSEARCH_USER="admin"
OPENSEARCH_PASSWORD="YourStrongPassword123!" # <<< USE THE SAME PASSWORD
DASHBOARDS_URL="http://my-dashboards-opensearch-dashboards.default.svc.cluster.local:5601"
OPENSEARCH_URL="https://my-os-opensearch.default.svc.cluster.local:9200"
echo "Waiting for OpenSearch Dashboards to be available..."
until curl -s -u "$OPENSEARCH_USER:$OPENSEARCH_PASSWORD" -k "$DASHBOARDS_URL/api/status" | grep -q "available"; do
sleep 5
done
echo "OpenSearch Dashboards is available."
echo "Creating ISM policy: nginx_retention..."
curl -X PUT "$OPENSEARCH_URL/_plugins/_ism/policies/nginx_retention" \
-u "$OPENSEARCH_USER:$OPENSEARCH_PASSWORD" -k -H "Content-Type: application/json" \
-d '{
"policy": {
"description": "Delete logs after 7 days",
"default_state": "hot",
"states": [
{
"name": "hot",
"actions": [],
"transitions": [
{
"state_name": "delete",
"conditions": { "min_index_age": "7d" }
}
]
},
{
"name": "delete",
"actions": [ { "delete": {} } ],
"transitions": []
}
],
"ism_template": [
{
"index_patterns": ["nginx-logs-*"],
"priority": 100
}
]
}
}'
echo "ISM policy created."
echo "Creating OpenSearch Dashboards Index Pattern: nginx-logs-pattern..."
curl -X POST "$DASHBOARDS_URL/api/saved_objects/index-pattern/nginx-logs-pattern" \
-u "$OPENSEARCH_USER:$OPENSEARCH_PASSWORD" -H "osd-xsrf: true" -H "Content-Type: application/json" \
-d '{
"attributes": {
"title": "nginx-logs-*",
"timeFieldName": "@timestamp"
}
}'
echo "Index pattern created."
---
apiVersion: batch/v1
kind: Job
metadata:
name: opensearch-observability-setup
spec:
template:
spec:
containers:
- name: setup-runner
image: curlimages/curl:latest # A lightweight image with curl and bash
command: ["bash", "/scripts/setup.sh"]
volumeMounts:
- name: setup-script
mountPath: /scripts
volumes:
- name: setup-script
configMap:
name: opensearch-sre-setup-script
defaultMode: 0744 # Make the script executable
restartPolicy: OnFailureApply the setup Job:
kubectl apply -f observability-setup-job.yamlBefore we fire off the job, let’s talk about what’s happening under the hood. We aren’t just pushing config; we are defining the lifecycle and visibility of our data.
nginx-logs-* into one continuous timeline so you can query across multiple days without lifting a finger.To see logs, your Nginx server needs visitors:
kubectl run load-gen --image=busybox --restart=Never -- /bin/sh -c "while true; do wget -qO- http://nginx-server; sleep 2; done"Port-forward Dashboards to your local machine:
kubectl port-forward svc/my-dashboards-opensearch-dashboards 5601:5601Then, open https://localhost:5601 in your browser. Log in with admin and your chosen password. Go to the “Discover” tab, select nginx-logs-* in the dropdown, set your time range (e.g., “Last 15 minutes”), and watch your logs flow in!
By following this automated approach, you’ve built a robust, observable, and easily reproducible log aggregation stack with OpenSearch. You’ve tackled critical SRE challenges like security, data ingestion, and lifecycle management—all as code. This hands-on experience forms a solid foundation for further exploration into metrics, traces, and advanced alerting in your cloud-native environments.
What other OpenSearch challenges will you automate next?
See the whole code at my repository
La entrada OpenSearch for SREs: The Open-Source Observability Powerhouse on Kubernetes se publicó primero en CloudArch.
]]>La entrada Introduction to KEDA: Event-Driven Autoscaling for Kubernetes se publicó primero en CloudArch.
]]>This is where KEDA (Kubernetes Event-Driven Autoscaling) comes in. KEDA is a lightweight, open-source component that integrates with Kubernetes to provide event-driven autoscaling. Unlike traditional Horizontal Pod Autoscalers (HPAs) that only scale based on CPU or memory usage, KEDA can scale your workloads based on external metrics or events, such as:
With KEDA, your applications can react instantly to real-world workloads without over-provisioning resources. It allows you to run microservices cost-effectively while maintaining responsiveness.
In this guide, we’ll build a realistic event-driven microservice that processes jobs from a Redis queue. The scenario mirrors what many production systems face:
By the end of this guide, you will have:
This project is valuable because it lets you experience a real-world autoscaling scenario. Most production systems have unpredictable workloads, and learning how KEDA responds to events gives you a deep understanding of:
Even if your real application is more complex (with multiple queues or different event sources), this example provides a solid foundation to implement production-ready, cost-efficient, and resilient services.
What you’ll learn:
Make sure you have the following locally installed and working on your machine:
If any of those are missing, install them first. For example on Debian/Ubuntu: sudo snap install kubectl –classic (or use your distro’s package manager). I will not assume any particular OS beyond these tools being available.
# Start minikube with 3 nodes (one control-plane + 2 workers), adjust memory/
CPUs if needed
$ minikube start --nodes=3 --driver=docker --memory=4096 --cpus=2
# Verify nodes come up
$ kubectl get nodesNotes: –nodes=3 creates 1 control-plane and 2 worker nodes by default. If you already have the cluster running, skip the minikube start step. Make sure kubectl context points to your minikube cluster: kubectl config current-context .
This command will install Helm for you in your local machine. Please, refer to this other post if you wanna learn more about it.
# Add KEDA Helm repo and update
helm repo add kedacore https://kedacore.github.io/charts
helm repo update
# Install KEDA into namespace 'keda'
helm install keda kedacore/keda --namespace keda --create-namespace
# Wait for KEDA pods to become ready
kubectl -n keda get podsWhy Helm? Helm is the easiest official way to install KEDA and its CRDs. KEDA installs CRDs that are required for ScaledObject and ScaledJob resources.
I’ll explain each file before showing it. All manifests are designed to run in the default namespace for simplicity. You can change namespace fields if you prefer.
# redis-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: redis
labels:
app: redis
spec:
replicas: 1
selector:
matchLabels:
app: redis
template:
metadata:
labels:
app: redis
spec:
containers:
- name: redis
image: redis:7.0-alpine
ports:
- containerPort: 6379
resources:
requests:
cpu: "100m"
memory: "128Mi"
limits:
cpu: "250m"
memory: "256Mi"
---
apiVersion: v1
kind: Service
metadata:
name: redis
labels:
app: redis
spec:
ports:
- port: 6379
targetPort: 6379
protocol: TCP
selector:
app: redis
type: ClusterIP
As we should know already, we are gonna deploy this in our cluster with:
kubectl apply -f redis-deployment.yamlTo be sure the redis service is client, we can test the connectivity with a redis client:
kubectl run -i --tty redis-client --image=redis:7.0-alpine --restart=Never
-- sh
# inside pod shell run: redis-cli -h redis ping
# should reply PONGWe’ll create a tiny Python worker that continuously polls a Redis list (named job-queue ) with BRPOP and “processes” messages (here, just sleep and echo ). Just to replicate some functionality of reading from Redis. The purpouse of this guide is not the logic of the code itself but it’s realibility on production environments.
In real life your worker would do meaningful job processing.
# worker.py
import time
import os
import redis
REDIS_HOST = os.getenv('REDIS_HOST', 'redis')
REDIS_PORT = int(os.getenv('REDIS_PORT', '6379'))
LIST_NAME = os.getenv('LIST_NAME', 'job-queue')
r = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
print('Worker started, connecting to', REDIS_HOST)
while True:
try:
# BRPOP blocks until an item is available
item = r.brpop(LIST_NAME, timeout=5)
if item:
# item is (list_name, value)
value = item[1]
print('Processing', value)
# simulate processing
time.sleep(2)
print('Done', value)
else:
# nothing to do, sleep to avoid tight loop
time.sleep(1)
except Exception as e:
print('Worker error:', e)
time.sleep(2)# Dockerfile.worker
FROM python:3.11-slim
WORKDIR /app
COPY worker.py .
RUN pip install --no-cache-dir redis
CMD ["python","/app/worker.py"]By default if we create the image, it will be in our local registry and Minikube won’t see it. To solve this problem we can enable an addon called registry which will allow us to create a registry in Minikube.
minikube addons enable registryWe need to do a port-forward to forward the data to the registry:
kubectl port-forward -n kube-system service/registry 5000:80And now, in other terminal to not terminate our tunnel, we need to build and push the image so it can be used within the cluster:
docker build -t keda-worker:latest -f Dockerfile.worker .
docker tag keda-worker:latest localhost:5000/keda-worker:latest
docker push localhost:5000/keda-worker:latestAnd now we need to create the deployment of our service using the brand new image:
# worker-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: keda-worker
labels:
app: keda-worker
spec:
replicas: 0 # start with 0 so we can see KEDA scale up from zero
selector:
matchLabels:
app: keda-worker
template:
metadata:
labels:
app: keda-worker
spec:
containers:
- name: worker
image: localhost:5000/keda-worker:latest
imagePullPolicy: Always
env:
- name: REDIS_HOST
value: "redis"
- name: REDIS_PORT
value: "6379"
- name: LIST_NAME
value: "job-queue"
resources:
requests:
cpu: "50m"
memory: "64Mi"
limits:
cpu: "200m"
memory: "256Mi"Please, pay attention that we stated that we want a total of 0 replicas since the design of this is to only run when Keda allows it, saving us computing time and resources in our cluster.
kubectl apply -f worker-deployment.yamlNow that we have Redis ready to have message and our service ready to start reading those messages when it’s needed, we need to create a ScaledObject which will tell our service when it’s time to work, create instances and do its job.
# scaledobject-redis.yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: keda-worker-scaledobject
labels:
app: keda-worker
spec:
scaleTargetRef:
name: keda-worker
pollingInterval: 5 # how often KEDA checks Redis (seconds)
cooldownPeriod: 30 # how long to wait after scale down before next check
minReplicaCount: 0
maxReplicaCount: 10
triggers:
- type: redis
metadata:
address: "redis.default.svc:6379"
listName: "job-queue"
listLength: "5" # scale target: number of items -> triggers scale
activationListLength: "1" # minimum backlog to activate scaling
Fields explained:
Please, pay special attention at triggers.type since that value is unique for the kind of job we are doing. Keda provides us with an extensive list of different triggers we can use depending on what we want to observe in order to scale our services. You can have a complete list at their official site.
Keda and Redis are on different namespaces on our cluster, when specifying the address make sure you are pointing correctly to Redis. If you don’t use
defaultnamespace as I did, it will be different thanredis.default.svc:6379
kubectl apply -f scaledobject-redis.yamlFinally, we are going to create a cronjob that will just generate some load to replicate a real user, and will just push a message to redis every minute so we can emulate the whole flow automatically and don’t hit any button neither waiting for an user.
# producer-cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: job-producer
spec:
schedule: "*/1 * * * *"
jobTemplate:
spec:
template:
spec:
restartPolicy: Never
containers:
- name: producer
image: redis:7.0-alpine
command: [ "sh", "-c" ]
args:
- |
now=$(date +%s)
echo "Producing job-$now"
redis-cli -h redis rpush job-queue job-$now
sleep 1kubectl apply -f producer-cronjob.yamlFor this purpose I have splited my terminal into 3 sections, so I can see everything at one glance. First I have to sections of the same size: right and left.
On the right size I will be running the logs of the Scaled Object to see how it’s being triggered everytime it sees there is 1 or more messages in the queue
kubectl logs -f -n keda keda-operator-f948b6c4-ln9chAnd then the left side I will have it splitted into two parts again. In the upper part I will be checking how many messages do I have on the list
watch 'kubectl exec -it deploy/redis -- redis-cli llen job-queue'And at the botton I will be watching the number of pods from my deplotments
watch 'kubectl get deploy'
There we will be able to see how every minute the list has a new message, a pod of the worker is being created and the message is deleted. Reading all the logs in the right side.
By this way, we could have a service that only will be running when it’s needed. But also, this can be replicated and configured to just scale up or down services based on some triggers to ensure we are always giving the desired availability and reliability.
kubectl delete -f producer-cronjob.yaml
kubectl delete -f scaledobject-redis.yaml
kubectl delete -f worker-deployment.yaml
kubectl delete -f redis-deployment.yaml
helm uninstall keda -n keda
kubectl delete namespace kedaYou can see and download all the code on GitHub https://github.com/JoaquinJimenezGarcia/LearningKeda
La entrada Introduction to KEDA: Event-Driven Autoscaling for Kubernetes se publicó primero en CloudArch.
]]>La entrada Monitoring Docker with Prometheus: Gain Full Visibility into Your Containers se publicó primero en CloudArch.
]]>As organizations increasingly rely on containerized applications, ensuring their performance, stability, and reliability becomes a critical task. Docker makes it easy to package and deploy applications, but without proper monitoring, you may miss vital insights into how your containers are behaving. Issues such as resource exhaustion, unexpected crashes, or networking bottlenecks can quickly escalate if left unnoticed.
This is where Prometheus, an open-source monitoring and alerting toolkit, comes into play. Prometheus is designed to collect, store, and query time-series metrics, making it an excellent fit for containerized environments. When paired with Docker, it gives you the ability to:
By setting up Prometheus to monitor Docker, you establish a foundation for observability that not only improves day-to-day operations but also builds confidence in your system’s resilience. In this post, we will walk through how to configure Prometheus to collect Docker metrics and show you how this setup can be the backbone of a reliable monitoring strategy.
Nowadays, getting a Prometheus instance up and running is simpler than it sounds. Thanks to Docker itself we can get our instance deployed and ready to be queried in seconds. Let’s see an example of a docker-compose file.
# docker-compose.yml
version: "3.8"
services:
prometheus:
image: prom/prometheus:latest
container_name: prometheus
restart: unless-stopped
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
- prometheus_data:/prometheus
volumes:
prometheus_data:We can create also a small configuration for starting collecting Prometheus own metrics
# prometheus.yml
global:
scrape_interval: 15s
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]Now if we run docker compose up -d we will have our instance serving traffic on port 9090
$ docker ps
CONTAINER ID IMAGE COMMAND CREATED STATUS PORTS NAMES
6416f7540370 prom/prometheus:latest "/bin/prometheus --c…" 9 seconds ago Up 9 seconds 0.0.0.0:9090->9090/tcp, [::]:9090->9090/tcp prometheus
Now it’s important to understand that we are gonna monitor docker itself, not the applications running with Docker.
Docker has a new feature to expose metrics on Prometheus-like format of the Docker server without being force to expose the whole Docker API, which is a win on security. To do that, we need to enable metrics-addr: on the deamon.json.
Thanks to that, Prometheus will be able to read and store our metrics.
# /etc/docker/daemon.json
{
"metrics-addr": "0.0.0.0:9323"
}After adding that, we will need to restart docker.
Now that we are exposing the metrics, we would be able to see them on the specified port and /metrics path
$ curl localhost:9323/metrics
# HELP builder_builds_failed_total Number of failed image builds
# TYPE builder_builds_failed_total counter
builder_builds_failed_total{reason="build_canceled"} 0
builder_builds_failed_total{reason="build_target_not_reachable_error"} 0
builder_builds_failed_total{reason="command_not_supported_error"} 0
builder_builds_failed_total{reason="dockerfile_empty_error"} 0
builder_builds_failed_total{reason="dockerfile_syntax_error"} 0
builder_builds_failed_total{reason="error_processing_commands_error"} 0
builder_builds_failed_total{reason="missing_onbuild_arguments_error"} 0
builder_builds_failed_total{reason="unknown_instruction_error"} 0
# HELP builder_builds_triggered_total Number of triggered image builds
# TYPE builder_builds_triggered_total counter
builder_builds_triggered_total 0
# HELP engine_daemon_container_actions_seconds The number of seconds it takes to process each container action
# TYPE engine_daemon_container_actions_seconds histogram
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.005"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.01"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.025"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.05"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.1"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.25"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="0.5"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="1"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="2.5"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="5"} 1
engine_daemon_container_actions_seconds_bucket{action="changes",le="10"} 1Now that we have Prometheus up and running and Docker exporting its metrics, it’s time to tell Prometheus were to look and scrape for the metrics, so later we can navigate through them.
In order to tell Prometheus where are the metrics, we need to modify the prometheus.yml file that we created during the first steps. We need to create a new job and specify the address. Our file should look like this now:
# prometheus.yml
global:
scrape_interval: 15s
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
- job_name: 'DockerStats'
static_configs:
- targets: ['172.17.0.1:9323']Note 1: Even on the
curlcommand we specified the/metricspath, here it’s not needed. By default, Prometheus will scrape on that path.
Note 2: Please, be careful of the target. 127.0.0.1 or
localhostwon’t work. Metrics are exposed on the docker bridge network which you can get from network interfacedocker0ip a | grep docker0 7: docker0: mtu 1500 qdisc noqueue state DOWN group default inet 172.17.0.1/16 brd 172.17.255.255 scope global docker0
Now we cand go to our Prometheus instance and we will see our job scraping the metrics correctly on /targets path.

That enables us for a vast options to explore our metrics and get to know better the state of our Docker service. As for example, we could use Prometheus as a Dataset on Grafana to visualize all the stats and a more beautiful way, and even create alerts based on thresholds.
La entrada Monitoring Docker with Prometheus: Gain Full Visibility into Your Containers se publicó primero en CloudArch.
]]>La entrada <h1>🔍 Observability: The Superpower Behind Healthy Servers 🚀</h1> se publicó primero en CloudArch.
]]>You might think it’s “just for DevOps”, but in reality, these practices protect your users, your revenue, and yes… even your weekend sleep 
What Is Observability (and Why Should You Care)?Let’s break it down:
Imagine your servers are a spaceship.
Monitoring is your dashboard—gauges, lights, speed indicators.
Observability is the system logs, black box, and mission control data that explain why the ship shakes when you press a button.
And alerting is the alarm that yells: “
Engine overheating!”
Without these, you’re flying blind. With them, you’re in control. 
The 3 Pillars of ObservabilityObservability is powered by:
— Numbers that reflect system performance (CPU, RAM, latency, etc.)
— Time-stamped records of system events.
— Data that follows the path of requests across services (crucial in microservices).Together, they give you deep visibility into how your system behaves—not just in a single spot, but across your entire stack.
Why It MattersLet’s get real.
Without observability:
With observability:
You detect issues earlier
You fix them faster
You prevent them from happening again
You reduce stress for your team and downtime for your users
Real Business Impact (with Data!)
Return on InvestmentA 2023 Observability Forecast by New Relic showed that:
“We were able to go from 12 hours of downtime a month to almost zero.”
— DevOps Manager, financial sector
Outage Cost Reduction
Without observability:
Average outage cost = $9.83M/year
With full-stack observability:
Reduced to $6.17M/year
That’s a savings of $3.66M annually… just by having the right insights! 
Faster Recovery = Happier Users
William Hill improved MTTR by 80%
Seven Network maintained 100% uptime during peak streaming
BlackLine cut cloud spend by $16M/year
Developer Experience“We spend $80K/month on observability to protect $15M/year in revenue. One missed SLA costs us $250K.”
— Reddit /r/devops user
The Role of Smart AlertsMonitoring is great, but alerting is what protects you from waking up to angry clients (or worse, a dead business). 
But not all alerts are created equal.
Bad alerting = noisy Slack channels and alert fatigue
Good alerting = smart, context-aware signals that only fire when something really needs attention
Combine alerts with automation (like restarting services or scaling infrastructure) and you’re moving toward self-healing systems 
A Real ExampleImagine your WordPress site is sluggish on mobile. 
Monitoring says all systems are “green”.
But using traces, you discover that a mobile-specific JS file fails to load, causing timeouts.
Without observability? You’d be in the dark.
With it? You fix it in 5 minutes—before users even notice.
In Conclusion: Why Observability Is EssentialIt’s not just about logs and dashboards.
It’s about trust, speed, resilience, and business success.
Catch problems early
Troubleshoot faster
Optimize cost and performance
Keep your team and customers happy
Innovate without fearObservability is your infrastructure’s early warning system, diagnosis tool, and performance coach—all in one. 

Want to See It in Action?Check out our live Grafana Demo Dashboard where we simulate a WordPress-based Linux server running real-time fake data. Perfect for learning, showing clients, or testing dashboards. 

Final WordsIn the world of cloud-native infrastructure, ignorance is never bliss.
Investing in monitoring, observability, and alerting isn’t a nice-to-have…
…it’s the foundation of a stable, scalable, and successful system. 
Until next time—stay observable, stay reliable, and may your error budgets be low! 
La entrada <h1>🔍 Observability: The Superpower Behind Healthy Servers 🚀</h1> se publicó primero en CloudArch.
]]>La entrada 🐳 Docker Monitoring: Keeping an Eye on Your Containers from the Start se publicó primero en CloudArch.
]]>
The Power of Built-in ToolsDocker provides built-in commands that offer valuable insights into container performance:
docker stats
This command displays real-time metrics for your running containers, including CPU usage, memoryconsumption, and network I/O.
docker logs
Access the logs of a container to monitor its output and diagnose issues.
These commands are straightforward and require no additional setup, making them ideal for quick checks during development.
Testing with docker.io/spkane/train-os:latestTo see these tools in action, let’s use the docker.io/spkane/train-os:latest image, which simulates system stress and is perfect for testing monitoring setups.
Run the container:
$ docker container run --rm -d --name stress docker.io/spkane/train-os:latest stress -v --cpu 2 --io 1 --vm 2 --vm-bytes 128M --timeout 60s
Unable to find image 'spkane/train-os:latest' locally
latest: Pulling from spkane/train-os
d4df0db66c89: Pull complete
19c5d5a1e2b2: Pull complete
2b25593057c7: Pull complete
0355d914b0bb: Pull complete
Digest: sha256:5acc35b4325d348c8ce6843f6751f62de6e83e518f94f5abe29d0f3ac0fb54be
Status: Downloaded newer image for spkane/train-os:latest
45e7f21918af3000a67d8f78bdfc6601d059160af9429304fca616b75e6036acMonitor with docker stats:
$ docker container stats stress --no-stream
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS
b75e1302b035 stress 429.59% 119.9MiB / 31.07GiB 0.38% 5.82kB / 126B 0B / 0B 6You’ll observe metrics like CPU and memory usage updating in real-time. We are using `–no-stream` to just have a brief output of the current state. Otherwise, it will be running being updated the values each few seconds.
View logs:
$ docker logs stress
stress: info: [1] dispatching hogs: 2 cpu, 1 io, 2 vm, 0 hdd
stress: dbug: [1] using backoff sleep of 15000us
stress: dbug: [1] setting timeout to 60s
stress: dbug: [1] --> hogcpu worker 2 [7] forked
stress: dbug: [1] --> hogio worker 1 [8] forked
stress: dbug: [1] --> hogvm worker 2 [9] forked
Accessing Metrics via Docker APIFor more advanced monitoring or integration with custom tools, you can access container stats directly through the Docker API:
$ curl --no-buffer -X GET --unix-socket /var/run/docker.sock http://docker/containers/stress/stats | head -n 1 | jq
% Total % Received % Xferd Average Speed Time Time Time Current
Dload Upload Total Spent Left Speed
0 0 0 0 0 0 0 0 --:--:-- --:--:-- --:--:-- 0{
"name": "/stress",
"id": "54370040079f3f7c3c6fd8608968050569b1141412c255949a0a72161f4a326a",
"read": "2025-06-04T19:45:21.220812324Z",
"preread": "0001-01-01T00:00:00Z",
"pids_stats": {
"current": 6,
"limit": 37968
},
"blkio_stats": {
"io_service_bytes_recursive": [
{
"major": 259,
"minor": 0,
"op": "read",
"value": 0
},
{
"major": 259,
"minor": 0,
"op": "write",
"value": 0
}
],This command fetches real-time statistics for the stress container in JSON format, which can be parsed and utilized by various monitoring solutions. That helps us to build our own monitoring solutions too, so if we run a budget environment or want to have control over all our stack we can easily control how our containers behave.
Note that curl is not making a TCP/IP call, we are directly hearing over the unix socket exposed for docker --unix-socket /var/run/docker.sock. That socket exports throught the Docker API /stats/ all needed parameters.
Choosing the Right Monitoring Approach| Scenario | Recommended Approach |
|---|---|
| Development & Testing | docker stats and docker logs |
| Custom Integrations | Docker API via curl |
| Production & Large Deployments | Grafana, Prometheus, etc. |
For small-scale applications or during development, Docker’s built-in tools are often sufficient. They provide immediate insights without the complexity of setting up external monitoring systems. However, as your application scales, integrating more robust solutions like Grafana and Prometheus becomes beneficial for long-term monitoring and alerting.
ConclusionMonitoring doesn’t have to be complex. Starting with Docker’s native tools allows for quick and effective oversight of your containers. As your needs grow, you can seamlessly transition to more comprehensive solutions. Remember, the key is to implement monitoring early to ensure smooth and efficient container operations.
La entrada 🐳 Docker Monitoring: Keeping an Eye on Your Containers from the Start se publicó primero en CloudArch.
]]>La entrada Mastering Prometheus and Alertmanager for Monitoring and Alerting se publicó primero en CloudArch.
]]>This guide will cover everything you need to start using Prometheus and Alertmanager, even if you are a complete beginner.
Prometheus is a robust monitoring tool designed for cloud-native environments. It collects metrics from configured targets at given intervals, evaluates rule expressions, and triggers alerts when thresholds are breached.
Key Features:
Alertmanager handles alerts generated by Prometheus, deduplicates them, groups them, and routes them to various receivers like email, Slack, or PagerDuty.
Key Features:
prometheus.yml configuration file:global:
scrape_interval: 15s # Default scrape interval
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']docker run -d --name=prometheus \
-p 9090:9090 \
-v $(pwd)/prometheus.yml:/etc/prometheus/prometheus.yml \
prom/prometheusalertmanager.yml configuration file:global:
resolve_timeout: 5m
route:
receiver: 'email-alert'
receivers:
- name: 'email-alert'
email_configs:
- to: 'your-email@example.com'
from: 'alertmanager@example.com'
smarthost: 'smtp.example.com:587'
auth_username: 'your-username'
auth_password: 'your-password'docker run -d --name=alertmanager \
-p 9093:9093 \
-v $(pwd)/alertmanager.yml:/etc/alertmanager/alertmanager.yml \
prom/alertmanagerModify prometheus.yml:
alerting:
alertmanagers:
- static_configs:
- targets: ['localhost:9093']
rule_files:
- 'alert_rules.yml'
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']Create an alert_rules.yml file:
groups:
- name: example-alert
rules:
- alert: InstanceDown
expr: up == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Instance {{ $labels.instance }} is down"
description: "No response from {{ $labels.instance }} for over 1 minute."Reload Prometheus:
curl -X POST http://localhost:9090/-/reloaddocker run -d -p 9100:9100 prom/node-exporterprometheus.yml:yamlCopy codescrape_configs: - job_name: 'node' static_configs: - targets: ['localhost:9100']--storage.tsdb.retention.time=15dPrometheus allows custom metrics via client libraries:
Install the library:
pip install prometheus_clientCreate a simple exporter:
from prometheus_client import start_http_server, Gauge
import random
import time
# Define a gauge metric
my_gauge = Gauge('random_number', 'A random number generator')
if __name__ == "__main__":
start_http_server(8000)
while True:
my_gauge.set(random.randint(0, 100))
time.sleep(5)Add it to prometheus.yml:
scrape_configs:
- job_name: 'custom-metrics'
static_configs:
- targets: ['localhost:8000']Use the Node Exporter to monitor Linux system metrics such as CPU, memory, disk usage, and network.
docker run -d -p 9100:9100 prom/node-exporterscrape_configs:
- job_name: 'node'
static_configs:
- targets: ['localhost:9100']node_cpu_seconds_totalnode_memory_Active_bytesalert_rules.yml:groups:
- name: high_cpu_usage
rules:
- alert: HighCPUUsage
expr: 100 - (avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 2m
labels:
severity: warning
annotations:
summary: "High CPU usage detected on {{ $labels.instance }}"
description: "CPU usage is above 80% for more than 2 minutes."curl -X POST http://localhost:9090/-/reloadBy following this guide, you can set up a powerful monitoring and alerting system with Prometheus and Alertmanager. Start small, experiment with metrics and alerts, and refine as you scale.
Feel free to share your questions or experiences in the comments below!
La entrada Mastering Prometheus and Alertmanager for Monitoring and Alerting se publicó primero en CloudArch.
]]>