Files
ceds-observability-stack/README.md
T
Chase DumphordandGitHub 7d576388b4 Update README.md
Made major changes and added a few sections
2026-04-28 11:50:03 -05:00

354 lines
7.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 🚀 Ceds Observability Stack
![Status](https://img.shields.io/badge/Status-Active%20Development-blue)
![Platform](https://img.shields.io/badge/Platform-Proxmox%20%7C%20K3s-orange)
![Monitoring](https://img.shields.io/badge/Stack-Prometheus%20%7C%20Grafana-red)
![Alerting](https://img.shields.io/badge/Alerting-Alertmanager-yellow)
![License](https://img.shields.io/badge/License-MIT-green)
---
## 🧠 Executive Summary
**Ceds Observability Stack** is a production-style monitoring, metrics, and alerting platform built to provide full visibility into a distributed hybrid infrastructure.
It simulates real-world **SRE / Platform Engineering environments**, delivering:
* 📊 Real-time infrastructure monitoring
* ⚙️ Kubernetes observability (12-node K3s cluster)
* 🖥️ Proxmox HA cluster visibility
* 🌐 Service uptime + network health tracking
* 🚨 Alerting pipelines (Alertmanager)
* 📈 Operational dashboards (Grafana)
> 🎯 **Goal:** Demonstrate enterprise-level observability practices across virtualization, Kubernetes, and self-hosted infrastructure.
---
## 🏗️ Environment Overview
### Core Infrastructure
| System | Purpose |
| ------------------------ | ---------------------------------- |
| 🖥️ Proxmox HA Cluster | Virtualization & high availability |
| ☸️ K3s Cluster (12-node) | Container orchestration |
| 💾 TrueNAS | Storage services |
| 🌐 Nginx Proxy Manager | Reverse proxy & routing |
| ☁️ Cloudflare | DNS, tunnels, external protection |
| 📊 Grafana | Visualization dashboards |
| 📡 Prometheus | Metrics collection |
| 🚨 Alertmanager | Alert routing |
---
## 🚀 Quick Start
### Prerequisites
- Linux server or VM
- Python 3 installed
- Prometheus installed
- Grafana installed
- Network access to homelab systems
---
### Run Service Health Check
bash python3 scripts/service-health-check.py
---
### Run Prometheus
bash prometheus --config.file=prometheus/prometheus.yml
---
### Access Services
- Prometheus: http://localhost:9090
- Grafana: http://localhost:3000
---
## 📡 Monitored Systems
| Target | Example Metrics |
| ----------------------- | ------------------------------------ |
| 🖥️ Proxmox Nodes | CPU, memory, storage, VM + HA status |
| ☸️ K3s Nodes | Node readiness, resource usage |
| 📦 Kubernetes Workloads | Pods, deployments, restarts |
| 🌐 Network Services | Uptime, latency, TCP checks |
| 💾 TrueNAS | Storage + service availability |
| 🔀 Nginx Proxy Manager | Reverse proxy health |
| 📊 Dashy / NOC | Dashboard availability |
| 🎬 Jellyfin | Media service uptime |
---
## 🧩 Architecture
```mermaid
flowchart TD
A[Proxmox HA Cluster] --> P[Prometheus]
B[12-Node K3s Cluster] --> P
C[Node Exporters] --> P
D[Service Health Checks] --> P
E[Proxmox Exporter] --> P
P --> G[Grafana Dashboards]
P --> AM[Alertmanager]
AM --> N[Email / Discord / Slack Alerts]
G --> NOC[Ced's NOC Dashboard]
```
---
## 📸 Dashboards
### Infrastructure Overview
Infrastructure Dashboard
### K3s Cluster Dashboard
K3s Dashboard
### Proxmox HA Dashboard
Proxmox Dashboard
### Service Uptime Dashboard
Services Dashboard
---
## ⚙️ Core Components
### 📡 Prometheus
Collects metrics from:
* Kubernetes endpoints
* Node exporters
* Proxmox exporter
* Custom health scripts
* Static service targets
---
### 📊 Grafana
Provides dashboards for:
* Cluster health
* Resource utilization
* Storage trends
* Service uptime
* Alert visibility
---
### 🚨 Alertmanager
Handles alerting for:
* Node failures
* High CPU / memory
* Service outages
* Pod crash loops
* Proxmox HA issues
---
## 📁 Repo Structure
```
ceds-observability-stack/
├── architecture/
├── prometheus/
├── grafana/
├── exporters/
├── alerting/
├── scripts/
└── docs/
```
---
## 📸 Dashboard Preview (Add Your Screenshots)
> 📌 Replace with real screenshots from your Grafana dashboards
* 🔹 Infrastructure Overview
* 🔹 K3s Cluster Health
* 🔹 Proxmox Cluster Status
* 🔹 Service Uptime Dashboard
---
## 🚀 Deployment (High-Level)
```bash
# Clone repo
git clone https://github.com/ced4568/ceds-observability-stack.git
# Navigate to project
cd ceds-observability-stack
# Deploy Prometheus + exporters
# (Add your actual deployment steps here)
# Access Grafana
http://<your-server-ip>:3000
```
---
## 🎯 Project Roadmap
### Phase 1 — Foundation
* [x] Architecture design
* [x] Repo structure
* [ ] Prometheus base config
* [ ] Grafana datasource
### Phase 2 — Metrics Collection
* [ ] Node exporter
* [ ] K3s metrics
* [ ] Proxmox exporter
* [ ] Uptime checks
### Phase 3 — Dashboards
* [ ] Infrastructure dashboard
* [ ] K3s dashboard
* [ ] Proxmox dashboard
* [ ] Service uptime dashboard
### Phase 4 — Alerting
* [ ] Alertmanager setup
* [ ] Alert rules
* [ ] Notification testing
### Phase 5 — Portfolio Polish
* [ ] Screenshots
* [ ] Architecture diagrams
* [ ] Setup guide
* [ ] Troubleshooting docs
---
## 🧠 Skills Demonstrated
* 📊 Infrastructure Monitoring
* ☸️ Kubernetes Operations
* 📡 Prometheus Configuration
* 📈 Grafana Dashboarding
* 🚨 Alert Engineering
* 🐧 Linux Administration
* 🖥️ Proxmox Virtualization
* ⚙️ SRE Principles
* 🏗️ Platform Engineering
---
## 🔗 Related Projects
| Project | Purpose |
| ----------------- | --------------------------------- |
| Ceds HomeLab | Full infrastructure ecosystem |
| Ceds NOC | Visualization + status dashboards |
| Ceds K3s HomeLab | Kubernetes architecture |
| Ceds APRS iGate | Networking + RF integration |
---
## 🔗 Integration
This observability stack is part of a larger ecosystem:
- Ceds HomeLab → Infrastructure layer
- Ceds Observability Stack → Metrics + monitoring layer
- Ceds NOC → Visualization and operations layer
Data flows from monitored systems into Prometheus, is visualized in Grafana, and feeds Ceds NOC dashboard for real-time system visibility.
---
## 🧪 Verification
To verify the system is working correctly:
### Prometheus Targets
- Navigate to: http://localhost:9090/targets
- Confirm all targets show UP
---
### Node Exporter
bash curl http://<node-ip>:9100/metrics
---
### Service Health Check
bash python3 scripts/service-health-check.py
---
### Grafana
- Confirm dashboards display real-time metrics
- Verify data source connection to Prometheus
- Check for active alerts
---
### Alert Testing
- Stop a service or node temporarily
- Confirm alert triggers in Prometheus
- Confirm alert appears in Grafana
---
## 📌 Status
🟢 **Active Development**
This project is continuously evolving as part of Ceds HomeLab ecosystem and professional portfolio.
---
## 💼 Why This Project Matters
This repository demonstrates the ability to:
* Design and operate **distributed systems**
* Implement **observability at scale**
* Build **production-style monitoring stacks**
* Apply **real-world SRE practices**
---
## 🧠 Future Improvements
- Loki log aggregation
- Tempo tracing
- Cloudflare Access log ingestion
- Automated remediation (self-healing infrastructure)
- Grafana public demo dashboard
- GitOps-based deployment (Argo CD / Flux)
- Multi-cluster Kubernetes monitoring
---
> 🚀 Designed as a **portfolio-grade project** for career growth, promotion, and technical leadership visibility.