Shashikant shah

Saturday, 4 April 2026

Introduction of Prometheus ?













Prometheus is an open-source monitoring and alerting toolkit designed for recording real-time metrics in a time-series database (TSDB), built especially for cloud-native and container-based environments like Kubernetes.

It collects and stores metrics as time series data (i.e., values with timestamps), supports powerful querying via PromQL, and integrates with Grafana and Alertmanager.


Key Features of Prometheus:














Feature

Description

Time Series Storage

Stores metrics as time series with labels

Pull-based model

Prometheus scrapes metrics from targets, unlike push-based systems

PromQL

Built-in query language for filtering, calculations, and alert conditions

Service Discovery

Automatically finds services via Kubernetes, Consul, EC2, etc.

Visualization

Built-in graph UI; best used with Grafana dashboards

Alerting

Define alerts based on thresholds; send notifications via Alertmanager


 Why Use Prometheus?

  • Open source and widely used
  • Lightweight and easy to install
  • Perfect for microservices and containerized apps
  • Strong community and support ecosystem
  • Compatible with exporters for system, database, application monitoring

Prometheus Components?

Prometheus has several core components that work together to collect, store, query, and alert based on metrics.

Component

Description

Prometheus Server

Core component that collects (scrapes), stores, and queries metrics

Exporters

Services or agents that expose metrics in Prometheus format

Alerting Rules

PromQL-based rules to define alert conditions

Alertmanager

Manages alerts – grouping, silencing, routing to email, Slack, etc.

Pushgateway

Allows short-lived jobs to push metrics to Prometheus

Service Discovery

Auto-discovers scrape targets (e.g., Kubernetes pods, EC2 instances)

Prometheus UI

Built-in web interface to run queries, check targets, and view alerts

Grafana

External tool for beautiful dashboards and visualizations using Prometheus data

Prometheus vs. Other Monitoring Tools.

Feature

Prometheus

Graphite

Elastic Stack

Datadog

Architecture

Pull-based

Push & Pull

Pull

Agent-based

Primary Data Model

Time series

Time series

Logs

Custom metrics

Query Language

PromQL

Custom DSL

kibana query

Custom UI

Horizontal Scaling

Supported

Limited

Supported

Fully managed service

Open source

Yes

Yes

Yes

No


Basic Terminologies in Prometheus

basic terminologies in Prometheus explained simply, with practical context for better understanding:

Term

Meaning

Time Series

Prometheus stores metric values with timestamps, so it creates time series data.

Metric

Metric is a measured value collected from a system, like CPU, memory, or request count. (e.g., http_requests_total)

Label

Key-value pair to differentiate time series

Job

Logical group of scrape targets

Instance

A single scrape target (usually host:port)

Target

Actual endpoint Prometheus scrapes metrics from

Scraping

Prometheus pulling data from targets at regular intervals

PromQL

Prometheus Query Language to analyze and fetch data

Exporter

Service or agent that exposes metrics in Prometheus format

Recording Rule

Precomputed PromQL query result stored as a new time series

Alerting Rule

PromQL expression that triggers alerts when conditions are met

Alertmanager

Handles alerts: grouping, deduping, routing to email, Slack, etc.

Pushgateway

Allows short-lived jobs to push metrics to Prometheus

Retention

Duration Prometheus keeps time-series data (e.g., 15 days)

TSDB

Time Series Database used internally by Prometheus

Service Discovery

Automatically finds targets (Kubernetes, Consul, EC2, etc.)

Histogram

Metric type that counts observations in configurable buckets

Summary

Metric type similar to histogram but provides quantile estimation

Gauge

Metric that can go up or down (e.g., memory usage, temperature)

Counter

Monotonically increasing metric (e.g., number of requests)

Label Set

Complete set of labels attached to a metric time series

Expression Browser

Prometheus web UI for querying and visualizing time series data


Architecture of Prometheus.



1. Time Series

  • What: A series of metric values tracked over time.
  • Example: CPU usage of a server every 10 seconds.

2. Metric

  • What: The actual measurement name.
  • Example: http_requests_total (total HTTP requests)

3. Label

  • What: Key-value pair to give more information about a metric.
  • Example:
    http_requests_total{method="GET", status="200"}
    method and status are labels.

4. Job

  • What: A group of similar targets.
  • Example: All Node Exporters can be under job node.

5. Instance

  • What: A single target with address (host:port).
  • Example: 10.1.2.3:9100 for one Node Exporter.

6. Target

  • What: The actual endpoint Prometheus collects data from.
  • Includes: job, instance, labels.

7. Scraping

  • What: The process where Prometheus collects data from a target.

8. PromQL

  • What: Prometheus Query Language.
  • Use: To filter, calculate, and display data.
  • Example:
    rate(http_requests_total[5m])

9. Exporter

  • What: A tool that exposes metrics in a Prometheus-readable format.
  • Example:
    • node_exporter for Linux server metrics
    • mysqld_exporter for MySQL metrics

10. Recording Rule

  • What: Saves the result of a query as a new metric.
  • Why: Reduces query load and speeds up dashboards.

11. Alerting Rule

  • What: A rule that defines when an alert should fire.
  • Example:
    Alert when CPU usage > 90% for 5 minutes.

12. Alertmanager

  • What: Manages alerts – sends them via Email, Slack, etc.
  • Also handles:
    • Grouping
    • Silencing
    • Routing

13. Pushgateway

  • What: Allows short-lived jobs to send data to Prometheus.
  • Why needed: Those jobs finish before Prometheus can scrape them.

14. Retention

  • What: How long Prometheus stores data.
  • Default: 15 days (can be customized)

15. TSDB (Time Series DB)

  • What: Internal database where Prometheus stores its data.
  • Supports: Fast reads, writes, and compression.

16. Service Discovery

  • What: Automatically detects targets (like Kubernetes pods, EC2).
  • Benefit: No need to manually add each new server.

17. Histogram

  • What: Metric type that counts values in buckets.
  • Used for: Request duration, response sizes.

18. Summary

  • What: Like histogram but shows percentiles (e.g., 95th percentile).
  • Used for: Latency measurement.

19. Gauge

  • What: Metric that goes up and down.
  • Examples: Memory usage, temperature.

20. Counter

  • What: Only increases.
  • Examples: Total HTTP requests, total errors.

21. Label Set

  • What: All labels assigned to a time series.
  • Helps to: Identify and group metrics.

22. Expression Browser

  • What: Built-in UI in Prometheus to test and run queries.
  • Use: For debugging or checking real-time metrics.

Prometheus ki Limitations:


1. Long-Term Storage nahi hai.

Prometheus apna data local TSDB me store karta hai. Ye months ya years tak metrics store karne ke liye design nahi hua hai.
Solution: Thanos, Grafana Mimir, VictoriaMetrics

2. Scalability Limited hai.

Ek single Prometheus server ki capacity limited hoti hai.

Agar:

  • Bahut saare servers ho
  • Millions of metrics ho
  • Bahut zyada scrape targets ho

to CPU, RAM aur Disk usage bahut badh jata hai.

Solution: Multiple Prometheus + Thanos/Mimir

3. Built-in High Availability (HA) nahi hai

Agar Prometheus server down ho gaya to monitoring bhi ruk jayegi.

Iske liye manually multiple Prometheus servers configure karne padte hain.

4. High Cardinality Problem

Agar labels me unique values bahut zyada hain, jaise:

  • user_id
  • session_id
  • request_id

to Prometheus millions of time series bana deta hai.

Isse:

  • RAM bahut consume hoti hai.
  • CPU usage badhta hai.
  • Queries slow ho jati hain.
  • Kabhi-kabhi Prometheus crash bhi ho sakta hai.

Ye Prometheus ki sabse badi limitation mani jati hai.

5. Sirf Metrics Store karta hai

Prometheus sirf metrics collect karta hai.

Ye:

  • Logs
  • Traces
  • Events

store nahi karta.

Complete observability ke liye Loki, Tempo ya Jaeger jaise tools use kiye jate hain.

6. Pull Model par kaam karta hai

Prometheus target se metrics scrape karta hai.

Agar target firewall ya NAT ke piche ho aur Prometheus us tak na pahunch sake, to monitoring mushkil ho jati hai.

7. Complex Queries Slow ho sakti hain

Agar:

  • Bahut bada data ho
  • Long time range ho
  • High-cardinality metrics ho

to PromQL queries slow chal sakti hain aur zyada resources consume karti hain.


Tools used with Prometheus

1. Long-Term Storage & High Availability

Ye tools Prometheus ki storage aur scalability ki limitation ko solve karte hain.

  • Thanos (Sabse popular)
  • Grafana Mimir
  • VictoriaMetrics
  • Cortex (Ab kam use hota hai, Mimir ne kaafi had tak replace kar diya)

2. Visualization (Dashboard)

Prometheus khud basic UI deta hai, dashboards ke liye:

  • Grafana (Industry Standard)
  • Kibana (Agar Elasticsearch use ho)
  • Chronograf (InfluxDB ecosystem)

3. Alerting

Prometheus ke alerts manage karne ke liye:

  • Alertmanager (Official)
  • PagerDuty
  • Opsgenie
  • Slack
  • Microsoft Teams
  • Email
  • Webhook

4. Log Monitoring

Prometheus sirf metrics collect karta hai, logs ke liye:

  • Grafana Loki
  • Elasticsearch
  • OpenSearch
  • Splunk
  • Graylog

5. Distributed Tracing

Microservices tracing ke liye:

  • Grafana Tempo
  • Jaeger
  • Zipkin

6. OpenTelemetry

Aajkal OpenTelemetry bahut popular hai.

  • OpenTelemetry Collector
  • OpenTelemetry SDK

Ye metrics, logs aur traces ko collect karke Prometheus, Loki, Tempo ya doosre backends ko bhej sakta hai.


7. Kubernetes Monitoring

Kubernetes me commonly:

  • kube-state-metrics
  • Node Exporter
  • cAdvisor
  • Prometheus Operator
  • kube-prometheus-stack

8. Exporters

Prometheus directly applications se data nahi leta. Exporters use hote hain.

Examples:

  • Node Exporter
  • Blackbox Exporter
  • SNMP Exporter
  • JMX Exporter
  • PostgreSQL Exporter
  • MySQL Exporter
  • Redis Exporter
  • NGINX Exporter
  • HAProxy Exporter
  • Kafka Exporter
  • RabbitMQ Exporter

9. Service Discovery

Targets automatically discover karne ke liye:

  • Kubernetes Service Discovery
  • Consul
  • DNS Service Discovery
  • EC2 Service Discovery
  • Azure Service Discovery
  • GCE Service Discovery

10. Push-Based Metrics

Agar pull model possible na ho:

  • Pushgateway
  • OpenTelemetry Collector
  • Prometheus Agent

Enterprise Monitoring Stack (Most Common)

Applications / Servers

        |

   Exporters / OpenTelemetry

        |

    Prometheus

        |

   +-------------+

   | Alertmanager|

   +-------------+

        |

     Grafana

        |

+--------------------------+

| Thanos / Mimir /         |

| VictoriaMetrics          |

+--------------------------+

        |

  Long-Term Storage

        |

Logs --> Loki

Traces --> Tempo / Jaeger


Prometheus ke sath kaun-kaun se tools use hote hain.

  • Monitoring: Prometheus

  • Visualization: Grafana

  • Alerting: Alertmanager

  • Long-Term Storage & HA: Thanos, Grafana Mimir, VictoriaMetrics

  • Logs: Loki, Elasticsearch/OpenSearch

  • Tracing: Tempo, Jaeger

  • Telemetry Collection: OpenTelemetry Collector

  • Kubernetes: Prometheus Operator, kube-state-metrics, Node Exporter

  • Exporters: Node, PostgreSQL, MySQL, Kafka, Redis, NGINX, HAProxy, JMX, Blackbox Exporter 

No comments:

Post a Comment