This project is a containerized uptime, observability, and alerting platform designed to simulate how modern engineering teams monitor services in production environments.
The platform allows users to register public website URLs through an API and automatically provides:
- Continuous uptime monitoring
- Response time and latency tracking
- Prometheus-based metrics collection
- Grafana dashboard visualization
- Real-time alerting through Alertmanager and Slack
- Centralized observability for monitored services
Unlike traditional manually configured monitoring setups, this platform follows a Platform Engineering approach, where monitoring capabilities are standardized, reusable, and automated through a self-service workflow.
The project demonstrates how modern DevOps and Platform Engineering practices are used to build scalable internal monitoring systems with a strong focus on:
- Automation
- Observability
- Standardization
- Reusability
- Scalability
- Incident response
The platform provides:
- Automated uptime monitoring
- Prometheus-based observability
- Grafana visualization dashboards
- Real-time alert routing
- Self-service monitoring onboarding
- Standardized metrics and alerts
- Containerized deployment architecture
In traditional environments, monitoring is often configured manually for each service, leading to inconsistent observability and operational overhead.
This platform addresses that challenge by providing:
- A reusable monitoring workflow
- Standardized metrics and alert structures
- Automated service onboarding
- Centralized observability and alerting
- Production-style container orchestration
👉 The result is a monitoring platform that reflects real-world DevOps and Site Reliability Engineering (SRE) practices.
The diagram below illustrates the high-level architecture of the monitoring platform, including the application layer, observability stack, and alerting workflow.
The platform is composed of two primary layers:
- Monitoring & Service Layer
- Shared Observability Layer
Together, these components provide automated uptime monitoring, metrics collection, visualization, and real-time alerting.
This layer is responsible for service onboarding, monitoring execution, metrics generation, and persistence.
Provides a self-service interface for registering URLs to be monitored.
Handles monitoring workflows and application logic, including:
- URL registration and validation
- HTTP availability checks
- Response time measurement
- Metrics exposure through
/metrics - Persistence of monitoring results
Executes automated monitoring checks at scheduled intervals to ensure continuous service validation.
Stores:
- Registered monitoring targets
- Historical uptime results
- Response time data
- Monitoring timestamps
This enables historical analysis and uptime tracking.
Represents external public websites and services monitored by the platform.
This layer centralizes metrics collection, visualization, alert evaluation, and notification delivery.
Responsible for:
- Scraping application metrics
- Collecting infrastructure metrics from Node Exporter
- Storing time-series data
- Evaluating alert rules
Provides real-time dashboards for:
- Uptime visibility
- Response time analysis
- Infrastructure monitoring
- Operational observability
Processes and routes alerts generated by Prometheus, including grouping and notification management.
Delivers real-time operational alerts and incident notifications.
Exposes host-level infrastructure metrics, including:
- CPU utilization
- Memory usage
- Disk activity
- Network statistics
The platform follows the following monitoring lifecycle:
- A monitoring target is registered through the API
- The scheduler executes periodic uptime checks
- Metrics are exposed through the application metrics endpoint
- Prometheus scrapes and stores monitoring metrics
- Alert rules evaluate service health and performance thresholds
- Alertmanager routes alerts to Slack
- Grafana visualizes operational metrics and monitoring trends
The platform architecture was designed with the following objectives:
- Automated service monitoring
- Centralized observability
- Reusable monitoring workflows
- Standardized metrics and alerting
- Containerized deployment
- Production-style operational visibility
The platform was designed using Platform Engineering principles to provide standardized, reusable, and automated monitoring workflows.
Instead of configuring monitoring independently for each service, the system centralizes observability through a consistent onboarding and monitoring model.
Monitoring targets can be registered dynamically through the API without requiring manual infrastructure or observability configuration.
Once a URL is onboarded, the platform automatically begins:
- Availability monitoring
- Response time collection
- Metrics generation
- Alert evaluation
The monitoring lifecycle is fully automated:
- Scheduled uptime checks
- Continuous metrics exposure
- Alert evaluation through Prometheus
- Incident routing through Alertmanager
- Real-time Slack notifications
This reduces operational overhead and ensures consistent monitoring behavior across services.
All monitored services follow a unified observability structure, including:
- Consistent metric naming
- Shared alerting policies
- Reusable dashboard patterns
- Centralized monitoring workflows
This improves scalability, maintainability, and operational consistency.
The platform architecture supports onboarding multiple services without requiring additional monitoring configuration.
The monitoring pipeline remains reusable across:
- Different target URLs
- Multiple environments
- Additional monitored services
The platform simplifies service onboarding by allowing developers to immediately gain:
- Uptime monitoring
- Performance visibility
- Metrics collection
- Automated alerting
without directly managing observability tooling.
- A monitoring target is registered through the API
- The backend validates and stores the target
- Scheduled monitoring checks execute automatically
- Response status and latency are recorded
- Monitoring data is persisted in PostgreSQL
- Metrics are exposed through
/metrics
- Prometheus scrapes application and infrastructure metrics
- Metrics are stored as time-series data
- Alert rules continuously evaluate service health
- Grafana visualizes operational and performance metrics
- Alertmanager processes and routes incidents
- Slack receives real-time operational notifications
The platform was built to:
- Centralize monitoring and observability
- Standardize metrics and alerting workflows
- Automate uptime and performance monitoring
- Provide reusable monitoring infrastructure
- Simulate production-style observability practices
- Improve operational visibility and incident response
Modern distributed systems require continuous operational visibility to maintain reliability and performance.
Key operational concerns include:
- Service availability
- Latency and response behavior
- Failure detection
- Alert responsiveness
- Infrastructure visibility
This platform addresses these concerns through a unified monitoring and observability architecture combining:
- Monitoring
- Metrics collection
- Visualization
- Alerting
- Incident notification
- Continuous uptime monitoring
- Response time measurement
- Automated background checks
- On-demand monitoring execution
- Historical monitoring persistence
- Prometheus-based metrics collection
- Custom application instrumentation
- Histogram-based latency tracking
- Grafana dashboard visualization
- Real-time alert evaluation
- Slack-based incident notifications
- Infrastructure monitoring through Node Exporter
- Self-service service onboarding
- Standardized observability workflows
- Automated monitoring lifecycle
- Reusable monitoring architecture
- Containerized deployment model
- Scalable metric and alert design
| Category | Technology |
|---|---|
| Backend API | Node.js, Express |
| Database | PostgreSQL |
| Monitoring | Prometheus |
| Visualization | Grafana |
| Alerting | Alertmanager, Slack |
| Metrics | Prometheus Client |
| Containerization | Docker, Docker Compose |
| Scheduling | Node-Cron |
| System Metrics | Node Exporter |
| Version Control | Git & GitHub |
| Observability | Prometheus Ecosystem |
| Platform Engineering | Self-Service Monitoring Architecture |
The platform currently supports:
- Automated uptime monitoring
- Response time and latency tracking
- Prometheus metrics exposure through
/metrics - Centralized observability with Grafana dashboards
- Real-time alert routing through Alertmanager
- Slack-based operational notifications
- Infrastructure-level monitoring via Node Exporter
- Persistent monitoring history using PostgreSQL
- Containerized multi-service deployment
Future improvements may include:
- Multi-user onboarding
- Authentication and access control
- Kubernetes-based deployment
- Distributed tracing with OpenTelemetry
- AI-assisted anomaly detection
- Multi-environment monitoring support
- Advanced dashboard provisioning
The repository is organized into modular components separating application logic, monitoring infrastructure, observability configuration, and deployment orchestration.
Web-uptime-Monitoring-Platform/
│
├── app/ # Monitoring Application Layer
│ ├── src/
│ │ ├── controllers/ # API request handlers
│ │ ├── services/ # Monitoring and business logic
│ │ ├── routes/ # API route definitions
│ │ ├── metrics/ # Prometheus instrumentation
│ │ ├── scheduler/ # Automated background monitoring
│ │ └── db/ # Database connectivity and queries
│ │
│ ├── server.js # Application entry point
│ ├── package.json # Application dependencies
│ └── Dockerfile # Container build configuration
│
├── database/
│ └── init.sql # Database schema initialization
│
├── monitoring/ # Observability & Alerting Layer
│ ├── prometheus/
│ │ ├── prometheus.yml # Metrics scrape configuration
│ │ └── alert.rules.yml # Prometheus alert rules
│ │
│ ├── alertmanager/
│ │ └── alertmanager.yml # Alert routing configuration
│ │
│ └── grafana/
│ └── provisioning/ # Dashboard and datasource provisioning
│
├── Images/ # Documentation assets and screenshots
│
├── docker-compose.yml # Multi-container orchestration
│
└── README.md # Project documentationThis project demonstrates how modern observability platforms can be engineered using standardized monitoring, centralized metrics collection, automated alerting, and reusable infrastructure patterns.
The platform combines:
- Continuous uptime monitoring
- Centralized observability
- Automated incident notification
- Infrastructure monitoring
- Reusable monitoring workflows
into a unified monitoring and alerting architecture.
Through this project, the platform implements:
- Automated monitoring workflows
- Standardized observability patterns
- Reusable alerting infrastructure
- Self-service monitoring onboarding
- Centralized operational visibility
- Production-style monitoring architecture
👉 The result is a scalable and observable monitoring platform aligned with modern DevOps, SRE, and Platform Engineering practices.
Establish the centralized observability layer responsible for:
- Metrics collection
- Infrastructure monitoring
- Alert evaluation
- Visualization
- Incident notification
This layer serves as the shared monitoring foundation for all services onboarded into the platform.
Rather than configuring observability independently for each service, the platform centralizes monitoring through reusable infrastructure components.
The observability stack is designed to provide:
- Standardized metrics collection
- Shared alerting policies
- Reusable dashboards
- Centralized operational visibility
This approach reflects modern DevOps, SRE, and Platform Engineering practices.
Create the required project directories and configuration files:
mkdir -p monitoring/prometheus
mkdir -p monitoring/alertmanager
mkdir -p monitoring/grafana
mkdir -p database
mkdir -p Images
touch docker-compose.yml
# Prometheus configuration
touch monitoring/prometheus/prometheus.yml
touch monitoring/prometheus/alert.rules.yml
# Alertmanager configuration
touch monitoring/alertmanager/alertmanager.yml
# Database initialization
touch database/init.sqlEdit the docker-compose.yml file:
services:
prometheus:
image: prom/prometheus
ports:
- "9091:9090"
volumes:
- ./monitoring/prometheus/prometheus.yml:/etc/prometheus/prometheus.yml
- ./monitoring/prometheus/alert.rules.yml:/etc/prometheus/alert.rules.yml
command:
- "--config.file=/etc/prometheus/prometheus.yml"
grafana:
image: grafana/grafana
ports:
- "4000:3000"
alertmanager:
image: prom/alertmanager
ports:
- "9094:9093"
volumes:
- ./monitoring/alertmanager/alertmanager.yml:/etc/alertmanager/alertmanager.yml
command:
- "--config.file=/etc/alertmanager/alertmanager.yml"
node-exporter:
image: prom/node-exporterThe observability stack was configured using Docker Compose to simplify service orchestration and container networking.
Key considerations:
- No fixed
container_namevalues were used - Internal container networking is used for service communication
- Node Exporter remains internal to avoid host-level port conflicts
- Volume mounts provide persistent configuration management
Create:
monitoring/prometheus/prometheus.ymlglobal:
scrape_interval: 5s
rule_files:
- "alert.rules.yml"
alerting:
alertmanagers:
- static_configs:
- targets:
- "alertmanager:9093"
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["prometheus:9090"]
- job_name: "node-exporter"
static_configs:
- targets: ["node-exporter:9100"]Prometheus acts as the centralized metrics engine responsible for:
- Metrics scraping
- Time-series storage
- Alert rule evaluation
- Metrics querying
Create:
monitoring/prometheus/alert.rules.ymlgroups:
- name: system-alerts
rules:
- alert: HighCPUUsage
expr: 100 - (avg by(instance)(rate(node_cpu_seconds_total{mode="idle"}[1m])) * 100) > 80
for: 1m
labels:
severity: warning
annotations:
summary: "High CPU usage detected"This establishes the initial alerting policy for infrastructure-level monitoring.
Create:
monitoring/alertmanager/alertmanager.ymlglobal:
resolve_timeout: 5m
route:
receiver: "default"
receivers:
- name: "default"Alertmanager is responsible for:
- Alert routing
- Alert grouping
- Notification management
- Incident delivery workflows
Slack integration is implemented in later stages.
docker compose up -d
docker compose psAccess:
http://localhost:9091/targets
Expected targets:
prometheus→ UPnode-exporter→ UP
Access:
http://localhost:4000
Default credentials:
admin / admin
Access:
http://localhost:9094
Infrastructure metrics were successfully exposed internally through Node Exporter and verified through Prometheus target discovery.
Collected host metrics include:
- CPU utilization
- Memory usage
- Disk activity
- Network statistics
The observability platform is considered operational when:
- Prometheus targets are healthy ✅
- Grafana is accessible ✅
- Alertmanager is operational ✅
- Infrastructure metrics are successfully collected ✅
Common setup issues during observability deployment included:
- YAML indentation errors
- Incorrect configuration mount paths
- Container networking issues
- Port conflicts between services
At this stage, the platform provides a reusable observability foundation consisting of:
- Centralized metrics collection
- Infrastructure monitoring
- Alert evaluation
- Dashboard visualization
- Containerized observability services
This observability layer becomes the shared monitoring backbone for all services onboarded into the platform.
Develop the monitoring application responsible for:
- Service onboarding
- Availability validation
- Response time measurement
- Metrics generation
- Prometheus instrumentation
This service acts as the operational core of the monitoring platform.
The monitoring application provides a self-service workflow for onboarding monitoring targets dynamically through an API-driven interface.
Once a target is registered, the platform automatically:
- Executes uptime checks
- Measures latency
- Generates observability metrics
- Exposes monitoring data to Prometheus
This removes the need for manually configuring monitoring per service.
The application layer was organized into modular components separating routing, business logic, metrics instrumentation, and service orchestration.
app/
├── src/
│ ├── controllers/
│ │ └── monitorController.js
│ ├── routes/
│ │ └── monitorRoutes.js
│ ├── services/
│ │ └── monitorService.js
│ ├── metrics/
│ │ └── metrics.js
│ ├── scheduler/
│ └── db/
│
├── server.js
├── package.json
└── Dockerfilenpm init -y
npm install express axios prom-clientInstalled packages provide:
- Express → API framework
- Axios → HTTP request execution
- prom-client → Prometheus metrics instrumentation
Create:
src/metrics/metrics.jsThe application exposes standardized Prometheus metrics including:
http_requests_totalfailed_checks_totalresponse_time_secondsuptime_status
Metrics are exposed through the /metrics endpoint and scraped by Prometheus.
The metrics implementation follows a structured labeling approach using:
serviceurlstatus
This enables:
- Per-service filtering
- Uptime analysis
- Response-time visibility
- Alert evaluation
- Operational observability
Create:
src/services/monitorService.jsThe monitoring service is responsible for:
- Sending HTTP requests to target services
- Measuring response latency
- Determining service availability
- Updating Prometheus counters and histograms
- Persisting monitoring results
Create:
src/controllers/monitorController.jsResponsibilities include:
- Request validation
- Monitoring workflow execution
- Structured API response handling
Create:
src/routes/monitorRoutes.jsPOST /monitor{
"url": "https://example.com"
}{
"url": "https://example.com",
"status": "UP",
"responseTime": 0.123
}Create:
server.jsThe server layer is responsible for:
- Initializing the Express application
- Registering API routes
- Enabling JSON request parsing
- Exposing Prometheus metrics through
/metrics
GET /metricsThe endpoint exposes:
- Custom monitoring metrics
- Prometheus-compatible application metrics
- Default Node.js runtime metrics
Prometheus continuously scrapes this endpoint for observability data collection.
Access:
http://localhost:3000/metrics
Create:
app/DockerfileContainerization ensures:
- Consistent runtime behavior
- Simplified deployment
- Integration with the observability stack
- Reproducible development environments
node server.jscurl -X POST http://localhost:3000/monitor \
-H "Content-Type: application/json" \
-d '{"url":"https://google.com"}'The platform stores monitoring targets and historical monitoring results in PostgreSQL for persistence and historical analysis.
SELECT * FROM monitored_urls;
SELECT * FROM check_results;The monitoring application layer is considered operational when:
- Monitoring targets can be registered ✅
- Availability checks execute successfully ✅
- Response times are measured correctly ✅
- Metrics are exposed through
/metrics✅ - Prometheus successfully scrapes application metrics ✅
- Monitoring data persists in PostgreSQL ✅
Common implementation issues included:
- Missing JSON body parsing
- Invalid URL handling
- Incorrect metric registration
- Missing
/metricsexposure - Docker networking misconfiguration
- Prometheus scrape configuration errors
At this stage, the platform provides a fully operational monitoring application capable of:
- Dynamic monitoring target onboarding
- Automated uptime validation
- Response time measurement
- Prometheus metrics generation
- Persistent monitoring history
- Integration with the centralized observability stack
This transforms the platform from a static observability setup into an API-driven monitoring system.
Introduce persistent storage into the monitoring platform to enable:
- Service registration persistence
- Historical uptime tracking
- Response-time history retention
- Stateful monitoring operations
This transforms the platform from a stateless monitoring API into a persistent observability system.
Modern monitoring platforms require historical visibility into service behavior over time.
Instead of performing temporary checks only, the platform now persists:
- Registered monitoring targets
- Monitoring execution history
- Availability states
- Performance measurements
This enables long-term operational analysis and service observability.
Create:
database/init.sqlCREATE TABLE monitored_urls (
id SERIAL PRIMARY KEY,
url TEXT UNIQUE NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
CREATE TABLE check_results (
id SERIAL PRIMARY KEY,
url_id INTEGER REFERENCES monitored_urls(id) ON DELETE CASCADE,
status VARCHAR(10),
response_time FLOAT,
checked_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);The persistence layer is divided into two primary tables:
Stores registered monitoring targets.
Stores historical monitoring executions including:
- Availability state
- Response latency
- Execution timestamp
This structure enables historical uptime analysis and future dashboard expansion.
npm install pgThe pg package enables communication between the Node.js application and PostgreSQL.
Create:
src/db/index.jsconst { Pool } = require('pg');
const pool = new Pool({
user: 'postgres',
host: 'db',
database: 'monitoring',
password: 'postgres',
port: 5432,
});
module.exports = pool;Create:
src/db/queries.jsThe query layer handles:
- Monitoring target registration
- URL lookup and reuse
- Monitoring result persistence
const pool = require('./index');
// Register or retrieve URL
async function getOrCreateUrl(url) {
const result = await pool.query(
`INSERT INTO monitored_urls (url)
VALUES ($1)
ON CONFLICT (url) DO UPDATE SET url = EXCLUDED.url
RETURNING id`,
[url]
);
return result.rows[0].id;
}
// Store monitoring result
async function saveCheckResult(urlId, status, responseTime) {
await pool.query(
`INSERT INTO check_results (url_id, status, response_time)
VALUES ($1, $2, $3)`,
[urlId, status, responseTime]
);
}
module.exports = {
getOrCreateUrl,
saveCheckResult,
};Update:
src/services/monitorService.jsAdditional responsibilities introduced:
- Register monitoring targets automatically
- Persist uptime results after every check
- Store response latency history
- Maintain monitoring state over time
This integration ensures all monitoring activity becomes traceable and queryable.
Update:
docker-compose.ymldb:
image: postgres:15
restart: always
environment:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: monitoring
volumes:
- ./database/init.sql:/docker-entrypoint-initdb.d/init.sql
ports:
- "5432:5432"Ensure the monitoring application depends on PostgreSQL.
app:
build: ./app
ports:
- "3000:3000"
depends_on:
- dbThis guarantees the database service initializes before the application starts.
docker compose up -d --buildThis rebuilds the application and initializes the PostgreSQL database schema automatically.
curl -X POST http://localhost:3000/monitor \
-H "Content-Type: application/json" \
-d '{"url":"https://google.com"}'Access PostgreSQL through Docker Compose:
docker compose exec db psql -U postgres -d monitoringSELECT * FROM monitored_urls;SELECT * FROM check_results;The database now stores:
- Registered monitoring targets
- Historical uptime records
- Response latency history
- Failure states
- Monitoring timestamps
The persistence layer is considered operational when:
- Monitoring targets are stored successfully ✅
- Monitoring results persist correctly ✅
- Historical records remain queryable ✅
- PostgreSQL initializes automatically ✅
- API functionality remains operational ✅
Common implementation issues included:
- Using
localhostinstead of Docker service name (db) - Missing database initialization mount
- Failure to rebuild containers after schema changes
- Missing application dependency ordering
- PostgreSQL connection timing issues
At this stage, the platform now supports:
- Persistent monitoring target registration
- Historical uptime tracking
- Stateful monitoring operations
- Long-term observability analysis
- Structured monitoring data retention
This transforms the platform into a production-style monitoring system with persistent operational visibility.
Introduce automated background monitoring to enable:
- Continuous uptime validation
- Scheduled monitoring execution
- Automated observability updates
- Real-time monitoring without manual API requests
This transforms the platform from an on-demand monitoring system into an automated monitoring service.
Production monitoring platforms continuously execute health checks in the background without requiring manual interaction.
Instead of relying on API-triggered checks only, the platform now:
- Continuously monitors registered targets
- Executes scheduled uptime validation
- Automatically updates metrics
- Continuously persists monitoring history
This introduces the platform’s automation layer.
Inside the application source directory:
mkdir -p src/scheduler
touch src/scheduler/scheduler.jsCreate:
src/scheduler/scheduler.jsThe scheduler is responsible for:
- Retrieving registered monitoring targets
- Executing monitoring checks automatically
- Running monitoring cycles continuously
- Updating monitoring history and metrics
const { monitorUrl } = require('../services/monitorService');
const pool = require('../db');
async function runMonitoringCycle() {
try {
const res = await pool.query('SELECT url FROM monitored_urls');
const urls = res.rows;
for (const row of urls) {
const url = row.url;
console.log(`Checking: ${url}`);
await monitorUrl(url);
}
} catch (error) {
console.error('Scheduler error:', error.message);
}
}
function startScheduler() {
console.log('🚀 Background monitoring started...');
// Execute every 30 seconds
setInterval(runMonitoringCycle, 30000);
}
module.exports = { startScheduler };Update:
server.jsconst { startScheduler } = require('./src/scheduler/scheduler');app.listen(PORT, () => {
console.log(`Server running on port ${PORT}`);
startScheduler();
});This ensures continuous monitoring begins automatically when the application starts.
Run the application:
node server.jsThe scheduler now starts automatically alongside the API service.
Once the application starts, monitoring cycles begin automatically.
🚀 Background monitoring started...
Checking: https://google.com
Checking: https://invalid-url-test-123.comThe scheduler continuously executes monitoring checks at the configured interval.
Wait for multiple monitoring cycles to complete, then access PostgreSQL:
docker compose exec db psql -U postgres -d monitoringRun:
SELECT * FROM check_results;You should observe new monitoring records being inserted automatically over time.
This confirms:
- Continuous monitoring execution
- Automated persistence
- Historical monitoring growth
Verify metrics generation:
curl http://localhost:3000/metricsThe metrics endpoint should continuously update without requiring manual monitoring requests.
This confirms:
- Prometheus metrics are updating automatically
- Background monitoring is operational
- Observability data is continuously generated
The automated monitoring lifecycle now operates as follows:
- Scheduler retrieves registered targets
- Monitoring service executes health checks
- Results are stored in PostgreSQL
- Metrics are updated automatically
- Prometheus scrapes updated observability data
- Grafana visualizes real-time monitoring information
This establishes a production-style automated monitoring pipeline.
The automation layer is considered operational when:
- Scheduler starts automatically ✅
- Registered URLs are monitored continuously ✅
- Monitoring results persist automatically ✅
- Metrics update without manual API execution ✅
- Monitoring cycles execute repeatedly at intervals ✅
Common implementation issues included:
- Forgetting to initialize
startScheduler() - Incorrect scheduler import path
- Database connection timing issues
- Failure to restart the application after scheduler integration
- Incorrect monitoring service function references
At this stage, the platform now supports:
- Fully automated background monitoring
- Continuous uptime validation
- Automated observability generation
- Real-time monitoring execution
- Persistent monitoring history growth
This transforms the platform into a continuously operating monitoring system aligned with real-world observability workflows.
Enhance the observability layer by introducing structured Prometheus labels and richer monitoring metrics.
This enables:
- Improved Prometheus querying
- Better Grafana visualization
- More granular monitoring visibility
- Clear distinction between healthy and failed monitoring states
The platform now moves beyond basic metrics into production-style observability instrumentation.
Modern observability systems rely heavily on labeled metrics for filtering, aggregation, and operational analysis.
Instead of collecting generic counters only, the platform now generates metrics enriched with contextual labels such as:
statusurlservice
This improves:
- Query precision
- Dashboard quality
- Alert targeting
- Operational visibility
Update:
src/metrics/metrics.jsThe metrics layer was enhanced to support structured Prometheus labels and richer observability data.
const client = require('prom-client');
const register = new client.Registry();
client.collectDefaultMetrics({ register });
// Total monitoring requests
const httpRequestsTotal = new client.Counter({
name: 'http_requests_total',
help: 'Total number of monitoring requests',
labelNames: ['status'],
});
// Failed monitoring checks
const failedChecksTotal = new client.Counter({
name: 'failed_checks_total',
help: 'Total number of failed uptime checks',
});
// Response time histogram
const responseTimeHistogram = new client.Histogram({
name: 'response_time_seconds',
help: 'Response time of monitored URLs',
labelNames: ['status'],
buckets: [0.1, 0.5, 1, 2, 5],
});
// Current uptime state
const uptimeStatus = new client.Gauge({
name: 'uptime_status',
help: 'Current uptime status (1 = up, 0 = down)',
});
register.registerMetric(httpRequestsTotal);
register.registerMetric(failedChecksTotal);
register.registerMetric(responseTimeHistogram);
register.registerMetric(uptimeStatus);
module.exports = {
register,
httpRequestsTotal,
failedChecksTotal,
responseTimeHistogram,
uptimeStatus,
};The enhanced metrics layer now tracks:
Tracks total monitoring executions categorized by status.
Tracks failed uptime validations.
Captures latency distributions through histogram buckets.
Represents the current service availability state.
Update:
src/services/monitorService.jsThe monitoring service was updated to generate labeled metrics dynamically during monitoring execution.
const axios = require('axios');
const {
httpRequestsTotal,
failedChecksTotal,
responseTimeHistogram,
uptimeStatus,
} = require('../metrics/metrics');
const { getOrCreateUrl, saveCheckResult } = require('../db/queries');
async function checkUrl(url) {
const start = Date.now();
let status = 'DOWN';
let duration = 0;
try {
await axios.get(url);
duration = (Date.now() - start) / 1000;
status = 'UP';
httpRequestsTotal.labels('up').inc();
responseTimeHistogram.labels('up').observe(duration);
uptimeStatus.set(1);
} catch (error) {
duration = (Date.now() - start) / 1000;
httpRequestsTotal.labels('down').inc();
failedChecksTotal.inc();
responseTimeHistogram.labels('down').observe(duration);
uptimeStatus.set(0);
}
// Save to database
const urlId = await getOrCreateUrl(url);
await saveCheckResult(urlId, status, duration);
return {
url,
status,
responseTime: duration,
};
}
module.exports = { checkUrl };Restart the application to apply the new metrics instrumentation.
node server.jsExecute a monitoring request:
curl -X POST http://localhost:3000/monitor \
-H "Content-Type: application/json" \
-d '{"url":"https://google.com"}'This generates labeled Prometheus metrics automatically.
Access the metrics endpoint:
curl http://localhost:3000/metricsThe metrics output should now include labeled observability data.
The platform now provides:
- Status-aware metrics (
up/down) - Latency distribution visibility
- Structured histogram buckets
- Query-ready Prometheus data
- Improved Grafana dashboard compatibility
This significantly improves monitoring visibility and operational analysis.
The enhanced observability layer is considered operational when:
- Metrics expose labels correctly ✅
- UP and DOWN states are tracked independently ✅
- Histogram buckets capture response-time distributions ✅
- Metrics update continuously through scheduler execution ✅
- Prometheus can query labeled metrics successfully ✅
Common implementation issues included:
- Forgetting to use
.labels() - Incorrect metric registration
- Missing application restart after instrumentation changes
- Misconfigured histogram labels
- Inconsistent metric naming
At this stage, the platform now supports:
- Structured Prometheus instrumentation
- Labeled observability metrics
- Enhanced monitoring visibility
- Query-optimized metrics
- Dashboard-ready telemetry
This establishes a more production-aligned observability foundation capable of supporting advanced alerting and visualization workflows.
Implement an automated alerting pipeline to transform the monitoring platform from passive observability into an active incident response system.
This enables:
- Automatic downtime detection
- Real-time alert generation
- Alert routing and notification delivery
- Faster operational response to incidents
👉 This layer introduces production-style alert management and operational visibility.
In modern platform environments:
- Monitoring alone is insufficient
- Systems must automatically detect failures and notify operators
- Alerting must be centralized, standardized, and reusable
👉 This layer introduces:
- Prometheus Alert Rules → define incident conditions
- Alertmanager Routing → process and distribute alerts
- Slack Notifications → provide real-time operational awareness
Update:
monitoring/prometheus/alert.rules.ymlgroups:
- name: system-alerts
rules:
- alert: HighCPUUsage
expr: 100 - (avg by(instance)(rate(node_cpu_seconds_total{mode="idle"}[1m])) * 100) > 80
for: 1m
labels:
severity: warning
annotations:
summary: "High CPU usage detected"
- name: uptime-alerts
rules:
- alert: WebsiteDown
expr: uptime_status == 0
for: 1m
labels:
severity: critical
service: uptime-monitor
annotations:
summary: "Website DOWN"
description: "Website {{ $labels.url }} is not reachable."
- alert: HighResponseTime
expr: response_time_seconds > 2
for: 1m
labels:
severity: warning
service: uptime-monitor
annotations:
summary: "High Response Time"
description: "Website {{ $labels.url }} is responding slowly."👉 These rules define standardized alert conditions for uptime and performance monitoring.
Update:
monitoring/prometheus/prometheus.ymlEnsure:
rule_files:
- "alert.rules.yml"👉 Prometheus automatically evaluates alert conditions using these rules.
Update:
monitoring/alertmanager/alertmanager.ymlglobal:
resolve_timeout: 5m
route:
receiver: "slack-notifications"
receivers:
- name: "slack-notifications"
slack_configs:
- api_url: "${SLACK_WEBHOOK_URL}"
channel: "#uptime-monitoring-alerts"
send_resolved: true
title: |
{{ if eq .Status "resolved" }}
✅ Website RESOLVED
{{ else }}
🚨 Website DOWN
{{ end }}
text: |
*{{ if eq .Status "resolved" }}✅ Website Monitoring RESOLVED{{ else }}🚨 Website Monitoring Alert{{ end }}*
Alert: {{ .CommonLabels.alertname }}
Status: {{ if eq .Status "resolved" }}🟢 RESOLVED{{ else }}🔴 DOWN{{ end }}
Severity: {{ .CommonLabels.severity }}
Website: {{ .CommonLabels.url }}
Message:
{{ if eq .Status "resolved" }}
Website is back to normal
{{ else }}
Website is not reachable
{{ end }}
Time:
{{ .StartsAt }}👉 Alertmanager routes alerts from Prometheus directly to Slack for operational visibility.
docker compose down
docker compose up -d👉 Restarting ensures Prometheus and Alertmanager reload updated configurations.
Run a monitoring request against an invalid URL:
curl -X POST http://localhost:3000/monitor \
-H "Content-Type: application/json" \
-d '{"url":"https://google-invalid.com"}'👉 You can also allow the background scheduler to detect failures automatically.
Open:
http://localhost:9091/alerts
Expected behavior:
WebsiteDown→ FIRING 🔴HighResponseTime→ depends on latency- Alert transitions to INACTIVE after recovery
👉 Prometheus successfully detected service downtime and transitioned the alert into the FIRING state.
👉 After service recovery, the alert automatically transitioned back to the INACTIVE state.
Alertmanager automatically sends incident notifications to Slack.
👉 Alertmanager successfully routed downtime notifications to Slack for real-time operational visibility.
👉 Recovery notifications were automatically delivered after the monitored service returned to a healthy state.
- Prometheus alert rules are configured correctly ✅
- Downtime incidents are automatically detected ✅
- Alerts transition through lifecycle states correctly ✅
- Alertmanager successfully routes alerts ✅
- Slack receives real-time incident notifications ✅
- Recovery notifications are automatically delivered ✅
- Incorrect metric names in alert expressions ❌
- Slack webhook not configured properly ❌
- Forgetting
send_resolved: true❌ - Not restarting Prometheus or Alertmanager ❌
- Using incorrect alert thresholds ❌
You have successfully implemented the Incident Response Layer, enabling:
- Automated failure detection
- Real-time alert generation
- Centralized alert routing
- Slack-based operational notifications
- Full incident lifecycle visibility
👉 The platform now supports production-style observability and incident management workflows.
Implement a visualization layer to transform raw monitoring metrics into real-time operational insights.
This enables:
- Real-time visibility into service health
- Performance trend analysis
- Faster troubleshooting and incident investigation
- Operational observability through dashboards
👉 This layer converts Prometheus metrics into actionable visual intelligence.
In modern platform environments:
- Raw metrics alone are insufficient
- Engineering teams require centralized and reusable dashboards
- Visualization improves operational awareness and incident response
👉 This layer introduces:
- Grafana Dashboards → visualize monitoring metrics
- Prometheus Data Source → query observability data
- Reusable Visualization Layer → shared across monitored services
Open:
http://localhost:4000
Login credentials:
Username: admin
Password: admin
👉 Grafana provides the centralized visualization interface for platform observability.
- Navigate to:
Connections → Data Sources
- Click:
Add data source
- Select:
Prometheus
URL: http://prometheus:9090
- Click:
Save & Test
Expected result:
Successfully queried the Prometheus API.
👉 Grafana is now connected to Prometheus and can query monitoring metrics in real time.
Navigate to:
Dashboards → New Dashboard → Add Visualization
👉 Multiple panels will be created to visualize uptime, traffic, failures, and response time metrics.
uptime_status
Time Series
👉 Displays service availability status over time:
1→ Service UP0→ Service DOWN
👉 Provides immediate visibility into service availability and uptime behavior.
http_requests_total
Time Series
👉 Displays monitoring request activity and traffic trends over time.
👉 Helps identify monitoring activity patterns and request growth over time.
failed_checks_total
Time Series
👉 Displays accumulated monitoring failures detected by the platform.
👉 Helps operators quickly identify recurring service failures and instability trends.
response_time_seconds_sum
Time Series
👉 Displays response latency trends for monitored services.
👉 Provides visibility into application responsiveness and performance degradation.
Enhance dashboard presentation for operational usability:
Uptime Monitoring Platform
- Service Availability
- Monitoring Request Volume
- Failure Detection
- Response Time Trend
- Green → Healthy / UP
- Red → Failed / DOWN
- Yellow → Warning / High Latency
👉 Consistent visual design improves readability and operational response efficiency.
Trigger a monitoring failure:
curl -X POST http://localhost:3000/monitor \
-H "Content-Type: application/json" \
-d '{"url":"https://google-invalid.com"}'Observe dashboard behavior:
- Uptime status changes
- Failure count increases
- Response metrics update
- Alerts begin firing
👉 Grafana updates automatically as Prometheus receives new metrics.
- Grafana successfully connects to Prometheus ✅
- Monitoring metrics are visualized correctly ✅
- Uptime behavior is visible in real time ✅
- Request and failure metrics update automatically ✅
- Dashboard reflects live operational state ✅
- Incorrect Prometheus URL configuration ❌
- Empty dashboards caused by missing metrics ❌
- No traffic generated for visualization ❌
- Using overly complex PromQL queries too early ❌
- Forgetting to save dashboard panels ❌
You have successfully implemented the Visualization Layer, enabling:
- Real-time operational visibility
- Centralized observability dashboards
- Performance and uptime analysis
- Human-readable monitoring insights
👉 The platform now provides full-stack observability across metrics, alerts, and visualization workflows.
Transform the monitoring system into a self-service platform where developers can onboard services dynamically without manual observability configuration.
This enables:
- Dynamic service registration
- Automated monitoring workflows
- Continuous background checks
- Fully integrated metrics, dashboards, and alerting
👉 The platform now behaves like an internal developer platform for observability.
In traditional monitoring systems:
- Engineers manually configure monitoring for each service
- Alerting and dashboards require additional setup
In a Platform Engineering model:
-
Developers simply register their service
-
The platform automatically provides:
- Monitoring
- Metrics collection
- Alerting
- Visualization
👉 This task introduces the core self-service experience expected in modern engineering platforms.
User manually triggered monitoring checks through the API
User registers URL → platform continuously monitors automatically
👉 Monitoring is now fully automated after onboarding.
At this stage, the platform already includes:
- Persistent database storage ✅
- Background scheduler ✅
- Metrics collection ✅
- Alerting pipeline ✅
👉 The scheduler will now continuously monitor every registered URL automatically.
Update:
app/src/scheduler/scheduler.jsconst cron = require("node-cron");
const { getAllUrls } = require("../db/queries");
const { monitorUrl } = require("../services/monitorService");
cron.schedule("* * * * *", async () => {
console.log("Running scheduled monitoring...");
try {
const urls = await getAllUrls();
for (const urlObj of urls) {
await monitorUrl(urlObj.url);
}
} catch (error) {
console.error("Scheduler error:", error.message);
}
});👉 The scheduler now acts as the platform automation engine.
👉 Registered services are now monitored automatically at scheduled intervals.
Update:
app/src/db/queries.jsasync function getAllUrls() {
const result = await pool.query("SELECT url FROM monitored_urls");
return result.rows;
}👉 This allows the scheduler to dynamically retrieve all onboarded services.
Update:
app/src/controllers/monitorController.jsconst { getOrCreateUrl } = require("../db/queries");
async function registerUrl(req, res) {
const { url } = req.body;
try {
await getOrCreateUrl(url);
res.json({
message: "URL registered successfully",
url,
});
} catch (error) {
res.status(500).json({
error: error.message,
});
}
}
module.exports = { registerUrl };👉 The API now performs service onboarding instead of manual monitoring execution.
Update:
app/src/routes/monitorRoutes.jsrouter.post("/monitor", registerUrl);👉 This endpoint now behaves as a self-service onboarding API.
Register a new service:
curl -X POST http://localhost:3000/monitor \
-H "Content-Type: application/json" \
-d '{"url":"https://google.com"}'Expected response:
{
"message": "URL registered successfully",
"url": "https://google.com"
}👉 The platform successfully onboarded the service dynamically.
After registration:
- URL stored in PostgreSQL ✅
- Scheduler detects registered service ✅
- Continuous monitoring begins automatically ✅
- Metrics update continuously ✅
- Alerts trigger automatically when failures occur ✅
- Grafana dashboards reflect live status ✅
👉 The monitoring lifecycle is now fully automated.
👉 Prometheus metrics continuously update without requiring manual API execution.
- Services can be registered dynamically ✅
- Scheduler continuously monitors onboarded URLs ✅
- Monitoring runs automatically without manual API triggers ✅
- Metrics update continuously in Prometheus ✅
- Alerting and dashboards continue functioning correctly ✅
- Forgetting to install
node-cron❌ - Scheduler not executing correctly ❌
- API still performing manual checks ❌
- Database queries not returning URLs ❌
- Not rebuilding containers after updates ❌
You have successfully implemented the Self-Service Monitoring Layer, enabling:
- Dynamic service onboarding
- Fully automated monitoring workflows
- Continuous observability pipelines
- Reusable platform monitoring capabilities
👉 The system now operates like a real internal Platform Engineering solution rather than a standalone monitoring application.
Standardize the observability platform to ensure:
- Consistent metric structures
- Reusable alerting patterns
- Scalable monitoring architecture
- Predictable observability workflows
👉 This transforms the project from a functional monitoring system into a reusable and production-aligned observability platform.
In modern engineering organizations:
- Multiple services are monitored simultaneously
- Metrics must follow standardized conventions
- Alerting pipelines must remain predictable
- Observability systems must scale cleanly
Without standardization:
- Metrics become inconsistent
- Dashboards become difficult to maintain
- Alert routing becomes unreliable
- Platform scalability becomes limited
👉 This task introduces standardized observability design principles commonly used in production environments.
Metrics and alerts function correctly but lack standardized structure
All observability components follow reusable and scalable conventions
👉 The platform now supports cleaner scaling and operational consistency.
To improve scalability and observability consistency, all metrics now follow a unified labeling structure.
| Label | Purpose |
|---|---|
service |
Identifies monitored platform/service |
url |
Identifies monitored endpoint |
status |
Indicates monitoring result (up / down) |
👉 This enables:
- Multi-service observability
- Label-based filtering
- Granular dashboard queries
- Structured alert routing
- Scalable monitoring patterns
Update:
app/src/services/monitorService.jshttpRequestsTotal.inc({
service: "uptime-monitor",
url,
status: status.toLowerCase(),
});responseTimeHistogram.observe(
{ service: "uptime-monitor", url, status: status.toLowerCase() },
duration
);failedChecksTotal.inc({
service: "uptime-monitor",
url,
});setActiveWebsiteStatus(url, status, "uptime-monitor");👉 Metrics now follow a production-style labeling strategy.
👉 The metrics output now includes:
serviceurlstatus
allowing structured observability queries across monitored services.
Update:
monitoring/prometheus/alert.rules.ymlAll alerts now include standardized labels for filtering, grouping, and routing.
- alert: WebsiteDown
expr: uptime_status{service="uptime-monitor"} == 0
for: 30s
labels:
severity: critical
service: uptime-monitor
annotations:
summary: "Website Down"
description: "Service {{ $labels.url }} is not reachable"- alert: HighResponseTime
expr: response_time_seconds{service="uptime-monitor"} > 2
for: 30s
labels:
severity: warning
service: uptime-monitor
annotations:
summary: "High Response Time"
description: "Service {{ $labels.url }} is responding slowly"👉 Alerts now support consistent filtering and routing behavior.
👉 Demonstrates:
WebsiteDownalert firing correctly- URL-specific detection
- Standardized service labeling
- Real-time incident visibility
👉 Demonstrates:
- Automatic recovery detection
- Alert lifecycle transition
- Healthy state restoration
👉 Shows:
- Centralized alert visibility
- Structured alert grouping
- Standardized alert organization
http://localhost:9091/alerts
👉 Alerts now include:
- Consistent labels
- Standardized naming
- Service-based filtering
- Clear operational visibility
To maintain operational consistency, alerts now follow a predictable naming convention.
| Alert Category | Standardized Name |
|---|---|
| Downtime | WebsiteDown |
| Recovery | WebsiteUp |
| Performance | HighResponseTime |
👉 This improves:
- Readability
- Incident response workflows
- Automation compatibility
- Operational consistency
Update Alertmanager routing configuration:
matchers:
- service="uptime-monitor"👉 This enables:
- Service-aware alert routing
- Scalable alert grouping
- Cleaner incident management
- Future multi-service expansion
http_requests_total{service="uptime-monitor"}
👉 Confirms:
- Metrics are filterable by service
- URL-based monitoring is functioning
- Status labels are applied consistently
http://localhost:3000/metrics
👉 Confirms:
- Structured metric output
- Standardized labels
- Production-style observability formatting
- Metrics include standardized labels (
service,url,status) ✅ - Alerts use consistent labeling structure ✅
- Alert naming conventions are standardized ✅
- Prometheus queries support service-level filtering ✅
- Alertmanager routing aligns with standardized labels ✅
- Platform observability structure is reusable and scalable ✅
- Adding labels in code but not updating metric definitions ❌
- Missing
servicelabels in alert rules ❌ - Inconsistent alert naming conventions ❌
- Forgetting to restart services after updates ❌
- Overusing high-cardinality labels unnecessarily ❌
You have successfully implemented the Standardization & Reusability Layer, enabling:
- Consistent observability architecture
- Scalable monitoring workflows
- Reusable platform monitoring patterns
- Cleaner integration across metrics, alerts, and dashboards
👉 The platform now reflects production-style observability engineering practices commonly used in modern Platform Engineering environments.
Containerize and orchestrate the entire monitoring platform using Docker Compose, enabling:
- Consistent environment setup
- Reliable service-to-service communication
- Persistent data storage
- Simplified deployment workflow
👉 This transforms the system into a fully portable and production-like monitoring platform.
In modern production environments:
- Applications run inside isolated containers
- Services communicate through internal container networking
- Infrastructure must remain reproducible across environments
- Persistent storage must survive container restarts
👉 Containerization ensures the monitoring platform behaves like a real-world deployment architecture.
Services run independently and require manual coordination
Entire platform runs through a unified and reproducible containerized stack
👉 The platform is now portable, scalable, and deployment-ready.
Update:
docker-compose.ymlThe platform now includes:
- App → Node.js Monitoring API
- Database → PostgreSQL
- Prometheus → Metrics collection
- Grafana → Visualization layer
- Alertmanager → Alert routing
- Node Exporter → Host-level metrics
services:
app:
build: ./app
ports:
- "3000:3000"
environment:
DB_HOST: db
DB_PORT: 5432
DB_USER: postgres
DB_PASSWORD: postgres
DB_NAME: monitoring
PORT: 3000
depends_on:
db:
condition: service_healthy
restart: unless-stopped
db:
image: postgres:15
restart: always
environment:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: monitoring
volumes:
- postgres_data:/var/lib/postgresql/data
- ./database/init.sql:/docker-entrypoint-initdb.d/init.sql
ports:
- "5432:5432"
healthcheck:
test: ["CMD-SHELL", "pg_isready -U postgres"]
interval: 5s
timeout: 5s
retries: 5
prometheus:
image: prom/prometheus
ports:
- "9091:9090"
volumes:
- ./monitoring/prometheus:/etc/prometheus
command:
- "--config.file=/etc/prometheus/prometheus.yml"
depends_on:
- app
restart: unless-stopped
grafana:
image: grafana/grafana
ports:
- "4000:3000"
volumes:
- grafana_data:/var/lib/grafana
depends_on:
- prometheus
restart: unless-stopped
alertmanager:
image: prom/alertmanager
ports:
- "9094:9093"
volumes:
- ./monitoring/alertmanager:/etc/alertmanager
env_file:
- ./app/.env
command:
- "--config.file=/etc/alertmanager/alertmanager.yml"
restart: unless-stopped
node-exporter:
image: prom/node-exporter
command:
- '--path.rootfs=/host'
- '--path.procfs=/host/proc'
- '--path.sysfs=/host/sys'
- '--collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($|/)'
pid: host
restart: unless-stopped
volumes:
- '/:/host:ro,rslave'
- '/proc:/host/proc:ro'
- '/sys:/host/sys:ro'
volumes:
postgres_data:
grafana_data:Docker Compose automatically provides DNS-based internal networking.
👉 Services communicate using container service names:
| Service | Internal Hostname |
|---|---|
| PostgreSQL | db |
| Monitoring App | app |
| Prometheus | prometheus |
| Alertmanager | alertmanager |
Update:
app/.envDB_HOST=db👉 Inside containers, localhost refers to the current container itself.
Using:
localhost
would break database connectivity between services.
Update:
monitoring/prometheus/prometheus.ymlscrape_configs:
- job_name: "node-app"
metrics_path: /metrics
static_configs:
- targets: ["app:3000"]
labels:
service: uptime-monitor👉 This ensures:
- Prometheus scrapes the correct application container
- Metrics include standardized service labels
- Monitoring remains consistent across environments
Run:
docker compose down -v
docker compose up --build -d👉 This ensures:
- Fresh environment initialization
- Updated configurations applied
- Containers rebuilt correctly
- Volumes recreated cleanly
Run:
docker compose ps👉 Confirms:
- All platform containers are healthy
- Services are running successfully
- Port mappings are configured correctly
Open:
http://localhost:3000/metrics
👉 Confirms:
- Metrics are exposed successfully
- Prometheus-compatible output is available
- Monitoring application is functioning correctly
Open:
http://localhost:9091/targets
👉 Confirms:
- Prometheus successfully scrapes all targets
- Monitoring jobs are healthy
- Internal service networking is functioning
Open:
http://localhost:4000
Login:
Username: admin
Password: admin
👉 Confirms:
- Grafana successfully connects to Prometheus
- Dashboards display monitoring data
- Visualization layer is operational
- All services run successfully through Docker Compose ✅
- Application connects to PostgreSQL using service name (
db) ✅ - Prometheus successfully scrapes application metrics ✅
- Grafana visualizes monitoring data correctly ✅
- Alertmanager continues routing alerts successfully ✅
- Persistent storage survives container restarts ✅
- Using
localhostinside containers ❌ - Missing database health checks ❌
- Forgetting persistent volumes ❌
- Incorrect Prometheus scrape targets ❌
- Not rebuilding containers after configuration changes ❌
You have successfully implemented the Containerized Platform Deployment Layer, enabling:
- Full platform orchestration
- Consistent deployment environments
- Reliable service communication
- Persistent observability infrastructure
- Production-style deployment architecture
👉 Your monitoring platform now reflects a modern containerized observability stack used in real-world engineering environments.
The platform consists of multiple integrated observability components working together:
- Node.js Monitoring API → handles service registration and monitoring
- PostgreSQL → stores monitored URLs and historical check results
- Prometheus → collects and stores metrics
- Alertmanager → manages and routes alerts
- Grafana → visualizes monitoring data
- Node Exporter → exposes host-level system metrics
The system follows a layered Platform Engineering architecture:
- Monitoring Layer
- Stateful Persistence Layer
- Automation Layer
- Observability Layer
- Incident Response Layer
- Visualization Layer
- Self-Service Platform Layer
- Standardization & Reusability Layer
- Containerized Deployment Layer
👉 This architecture simulates a real-world production monitoring platform.
- Self-service website registration
- Automated background monitoring
- Persistent uptime history storage
- Prometheus metrics integration
- Grafana visualization dashboards
- Real-time Slack alerting
- Dockerized multi-service deployment
- Standardized observability labels
- Production-style monitoring architecture
- Continuous monitoring scheduler
Possible future enhancements include:
- Kubernetes deployment support
- Helm chart packaging
- Multi-user authentication system
- Role-based access control (RBAC)
- Email and Microsoft Teams alert integrations
- HTTPS and TLS support
- Distributed monitoring agents
- Advanced Grafana dashboards
- CI/CD pipeline integration
- Auto-scaling monitoring workers
- Terraform infrastructure provisioning
- Centralized logging with Loki or ELK Stack
Through this project, I gained practical experience in:
- Building observability pipelines
- Prometheus metrics design
- Alert lifecycle management
- Grafana dashboard creation
- Container orchestration with Docker Compose
- Service-to-service networking
- Stateful monitoring architecture
- Platform Engineering concepts
- Monitoring automation strategies
- Production-style infrastructure design
This project demonstrates the design and implementation of a complete production-style monitoring platform using modern observability tools and Platform Engineering principles.
The platform supports:
- Self-service monitoring
- Automated uptime checks
- Persistent historical storage
- Real-time metrics collection
- Alert management and notifications
- Grafana-based visualization
- Standardized observability practices
- Fully containerized deployment
By combining monitoring, alerting, visualization, automation, and container orchestration into a unified architecture, this project reflects real-world DevOps and Platform Engineering workflows used in modern production environments.
Junior DevOps Engineer passionate about:
- Cloud Infrastructure
- Platform Engineering
- Observability & Monitoring
- DevOps Automation
- Infrastructure as Code (IaC)
- LinkedIn: www.linkedin.com/in/philip-oludolamu
- GitHub: github.com/holuphilix
- Portfolio: phillipoludolamu.com






























