Add CedOps V3 SRE dashboard and README updates

This commit is contained in:
2026-05-17 13:44:08 -05:00
parent 4925e52b51
commit 530808c6a2
4 changed files with 155 additions and 100 deletions
+113 -58
View File
@@ -14,38 +14,58 @@
## 📸 NOC Dashboard ## 📸 NOC Dashboard
### V2 — Current CedOps dashboards are intentionally versioned to demonstrate increasing operational maturity — from foundational observability to SRE-style incident response and reliability engineering.
### V3 — Current (SRE / Incident Response)
[![CedOps NOC Dashboard V3](docs/screenshots/dashboard_v3-overview.png)](docs/screenshots/dashboard_v3-overview.png)
#### Incident Response & Dependencies
[![CedOps Incident View](docs/screenshots/dashboard_v3-incidents.png)](docs/screenshots/dashboard_v3-incidents.png)
#### Reliability / SLO Monitoring
[![CedOps SLO View](docs/screenshots/dashboard_v3-slo.png)](docs/screenshots/dashboard_v3-slo.png)
---
### V2 — Operational Reliability
[![CedOps NOC Dashboard V2](docs/screenshots/dashboard_v2.png)](docs/screenshots/dashboard_v2.png) [![CedOps NOC Dashboard V2](docs/screenshots/dashboard_v2.png)](docs/screenshots/dashboard_v2.png)
### V1 ---
[![CedOps NOC Dashboard V1](docs/screenshots/dashboard_V1.png)](docs/screenshots/dashboard_V1.png) ### V1 — Foundational Observability
> **Live:** All 5 public services reporting UP · NOC Health Score 100% · VPS Disk 27% · Prometheus scraping every 30s [![CedOps NOC Dashboard V1](docs/screenshots/dashboard_v1.png)](docs/screenshots/dashboard_v1.png)
> **Live:** Public services monitored every 30 seconds · Prometheus scraping via Tailscale · Grafana visualizing hybrid infrastructure health
--- ---
## 🧭 Overview ## 🧭 Overview
CedOps is a **production-style AI operations platform** built to simulate a real-world Platform Engineering and AI Infrastructure environment. The VPS stack runs on Hetzner and is monitored by Prometheus hosted on a Proxmox homelab node, communicating over a private Tailscale network — with everything visualized in a live Grafana NOC dashboard. CedOps is a **production-style AI operations platform** built to simulate a real-world Platform Engineering and AI Infrastructure environment.
The VPS stack runs on Hetzner and is monitored by Prometheus hosted on a Proxmox homelab node, communicating over a private Tailscale network — with everything visualized in a live Grafana NOC dashboard.
The platform combines: The platform combines:
- **Paperclip** — AI Agent Orchestration Platform managing 8 autonomous agents * **Paperclip** — AI Agent Orchestration Platform managing 8 autonomous agents
- **SugeBot** — Hermes AI agent deployed as Telegram Operations Assistant (CTO/Ops role) * **SugeBot** — Hermes AI agent deployed as Telegram Operations Assistant (CTO/Ops role)
- **n8n** — Automation platform for workflow execution and inter-service routing * **n8n** — Automation platform for workflow execution and inter-service routing
- **OpenRouter** — AI model gateway serving as the LLM backbone for all agents * **OpenRouter** — AI model gateway serving as the LLM backbone for all agents
- **Prometheus** — Hosted on Proxmox homelab, scraping VPS services over Tailscale * **Prometheus** — Hosted on Proxmox homelab, scraping VPS services over Tailscale
- **Grafana** — Full observability NOC dashboard visualizing the entire stack in real-time * **Grafana** — Full observability NOC dashboard visualizing the entire stack in real-time
- **Nginx Proxy Manager** — Reverse proxy + SSL management across all services * **Nginx Proxy Manager** — Reverse proxy + SSL management across all services
- **Docker + Portainer** — Container infrastructure and management layer * **Docker + Portainer** — Container infrastructure and management layer
--- ---
## 🏗️ Architecture Overview ## 🏗️ Architecture Overview
``` ```text
Internet Internet
┌─────────────┴─────────────┐ ┌─────────────┴─────────────┐
@@ -102,7 +122,7 @@ The platform combines:
All agents run inside **Paperclip** — the AI Agent Orchestration Platform at `paperclip.synthossystems.com`. All agents run inside **Paperclip** — the AI Agent Orchestration Platform at `paperclip.synthossystems.com`.
| Agent | Role | Status | | Agent | Role | Status |
|---|---|---| | ----------------------- | ---------------------------------------------------------------------- | ------ |
| **Suge** | CTO / Operations (Hermes Agent) — also deployed as SugeBot on Telegram | ✅ Live | | **Suge** | CTO / Operations (Hermes Agent) — also deployed as SugeBot on Telegram | ✅ Live |
| **CEO** | Executive decision-making and strategic oversight | ✅ Live | | **CEO** | Executive decision-making and strategic oversight | ✅ Live |
| **ClaudeCoder** | Software engineering and code generation | ✅ Live | | **ClaudeCoder** | Software engineering and code generation | ✅ Live |
@@ -121,7 +141,7 @@ All agents run inside **Paperclip** — the AI Agent Orchestration Platform at `
All services run on a **Hetzner VPS** (`synthos-node-1`) behind Cloudflare Tunnels with zero exposed ports. All services run on a **Hetzner VPS** (`synthos-node-1`) behind Cloudflare Tunnels with zero exposed ports.
| Service | Role | URL | | Service | Role | URL |
|---|---|---| | ----------------------- | ------------------------------- | ------------------------ |
| **VPS / Portainer** | Docker container management | Internal | | **VPS / Portainer** | Docker container management | Internal |
| **Nginx Proxy Manager** | Reverse proxy + SSL termination | Internal | | **Nginx Proxy Manager** | Reverse proxy + SSL termination | Internal |
| **Homepage** | CedOps command center dashboard | `ops.synthossystems.com` | | **Homepage** | CedOps command center dashboard | `ops.synthossystems.com` |
@@ -129,7 +149,8 @@ All services run on a **Hetzner VPS** (`synthos-node-1`) behind Cloudflare Tunne
| **Ced's Home Lab** | Homelab / Portfolio platform | `cedshomelab.com` | | **Ced's Home Lab** | Homelab / Portfolio platform | `cedshomelab.com` |
**Traffic Flow:** **Traffic Flow:**
```
```text
Internet → Cloudflare Edge → Tunnel → Nginx Proxy Manager → Docker Services Internet → Cloudflare Edge → Tunnel → Nginx Proxy Manager → Docker Services
``` ```
@@ -137,26 +158,30 @@ Internet → Cloudflare Edge → Tunnel → Nginx Proxy Manager → Docker Servi
## 📊 Observability & NOC ## 📊 Observability & NOC
Prometheus runs on the **Proxmox homelab** and scrapes the Hetzner VPS over a **Tailscale private tunnel** — keeping metrics collection completely off the public internet. Grafana visualizes everything in the live NOC dashboard. Prometheus runs on the **Proxmox homelab** and scrapes the Hetzner VPS over a **Tailscale private tunnel** — keeping metrics collection completely off the public internet.
Grafana visualizes everything in the live NOC dashboard.
### Executive Health Panel ### Executive Health Panel
- **Synthos VPS Status** — UP/DOWN state with instant alerting
- **Public Services Up** — count of healthy endpoints (target: 5/5) * **Synthos VPS Status** — UP/DOWN state with instant alerting
- **NOC Health Score** — composite platform health percentage * **Public Services Up** — count of healthy endpoints (target: 5/5)
- **VPS Uptime** — continuous uptime tracking in days * **NOC Health Score** — composite platform health percentage
* **VPS Uptime** — continuous uptime tracking in days
### VPS Resource Health ### VPS Resource Health
- CPU Usage % — `synthos-node-1` CPU time series
- Memory Utilization % — real-time memory pressure * CPU Usage % — `synthos-node-1` CPU time series
- Disk Used % — current at 27% * Memory Utilization % — real-time memory pressure
- Network Traffic — RX/TX bandwidth graphs (30s refresh) * Disk Used % — infrastructure capacity monitoring
* Network Traffic — RX/TX bandwidth graphs (30s refresh)
### Public Service Availability ### Public Service Availability
Tracked via **Blackbox Exporter** HTTP probes: Tracked via **Blackbox Exporter** HTTP probes.
| Service | Endpoint | Status | | Service | Endpoint | Status |
|---|---|---| | --------------- | ---------------------------- | ------ |
| cedshomelab | cedshomelab.com | ✅ UP | | cedshomelab | cedshomelab.com | ✅ UP |
| n8n | n8n.synthossystems.com | ✅ UP | | n8n | n8n.synthossystems.com | ✅ UP |
| cedops | ops.synthossystems.com | ✅ UP | | cedops | ops.synthossystems.com | ✅ UP |
@@ -164,9 +189,13 @@ Tracked via **Blackbox Exporter** HTTP probes:
| synthos-website | synthossystems.com | ✅ UP | | synthos-website | synthossystems.com | ✅ UP |
### Reliability Metrics ### Reliability Metrics
- HTTP Status Codes per endpoint (all returning `200`)
- Public Endpoint Response Time (ms) — time series graph * HTTP Status Codes per endpoint
- SSL Expiry Days — tracked per certificate with warning thresholds * Public Endpoint Response Time
* SSL Expiry Days tracking
* Service availability monitoring
* SLO / SLA visibility
* Incident timeline monitoring
--- ---
@@ -175,7 +204,7 @@ Tracked via **Blackbox Exporter** HTTP probes:
CedOps is designed with **zero public infrastructure exposure**. CedOps is designed with **zero public infrastructure exposure**.
| Control | Implementation | | Control | Implementation |
|---|---| | ------------------------ | ----------------------------------------------------------- |
| No open ports | Cloudflare Tunnel handles all ingress | | No open ports | Cloudflare Tunnel handles all ingress |
| Private metrics pipeline | Prometheus → Tailscale → VPS; never public | | Private metrics pipeline | Prometheus → Tailscale → VPS; never public |
| SSL everywhere | NPM handles termination for all services | | SSL everywhere | NPM handles termination for all services |
@@ -186,19 +215,23 @@ CedOps is designed with **zero public infrastructure exposure**.
## 📂 Repository Structure ## 📂 Repository Structure
``` ```text
cedops-ai-infrastructure-noc/ cedops-ai-infrastructure-noc/
├── docs/ ├── docs/
│ ├── architecture/ │ ├── architecture/
│ ├── screenshots/ │ ├── screenshots/
│ │ ├── dashboard_v1.png │ │ ├── dashboard_v1.png
│ │ ── dashboard_v2.png │ │ ── dashboard_v2.png
│ │ ├── dashboard_v3-overview.png
│ │ ├── dashboard_v3-incidents.png
│ │ └── dashboard_v3-slo.png
│ └── sops/ │ └── sops/
├── grafana/ ├── grafana/
│ ├── cedops-ai-infrastructure-noc-v1.json │ ├── cedops-ai-infrastructure-noc-v1.json
── cedops-ai-infrastructure-noc-v2.json ── cedops-ai-infrastructure-noc-v2.json
│ └── cedops-sre-incident-response-v3.json
├── Prometheus/ ├── Prometheus/
│ ├── prometheus-example.yml │ ├── prometheus-example.yml
@@ -210,14 +243,14 @@ cedops-ai-infrastructure-noc/
└── README.md └── README.md
``` ```
> **Note:** Prometheus and Blackbox configs are sanitized for public sharing. Production configs are managed privately via environment variables and are never committed to this repo. > **Note:** Prometheus and Blackbox configs are sanitized for public sharing. Production configs are managed privately and never committed to this repo.
--- ---
## 🛠️ Stack ## 🛠️ Stack
| Layer | Technology | | Layer | Technology |
|---|---| | -------------------- | ---------------------------------------------- |
| **Compute** | Hetzner VPS (synthos-node-1) | | **Compute** | Hetzner VPS (synthos-node-1) |
| **Containers** | Docker + Portainer | | **Containers** | Docker + Portainer |
| **Reverse Proxy** | Nginx Proxy Manager | | **Reverse Proxy** | Nginx Proxy Manager |
@@ -227,38 +260,59 @@ cedops-ai-infrastructure-noc/
| **AI Gateway** | OpenRouter | | **AI Gateway** | OpenRouter |
| **Automation** | n8n | | **Automation** | n8n |
| **Ops Agent** | SugeBot (Hermes / Telegram) | | **Ops Agent** | SugeBot (Hermes / Telegram) |
| **Metrics** | Prometheus (Proxmox) + Node Exporter + Blackbox Exporter | | **Metrics** | Prometheus + Node Exporter + Blackbox Exporter |
| **Visualization** | Grafana | | **Visualization** | Grafana |
| **Virtualization** | Proxmox VE (homelab, hosts Prometheus + Grafana) | | **Virtualization** | Proxmox VE |
--- ---
## ✅ What's Built ## ✅ What's Built
### V1 Dashboard ### V1 Dashboard — Foundational Observability
- [x] VPS UP/DOWN status panel
- [x] Public service availability (5 endpoints via Blackbox Exporter)
- [x] VPS resource health (CPU, Memory, Disk, Network)
- [x] HTTP status codes per endpoint
- [x] Public endpoint response times
- [x] SSL certificate expiry tracking
- [x] Cross-environment scraping via Tailscale (Proxmox → Hetzner VPS)
### V2 Dashboard * [x] VPS UP/DOWN status panel
- [x] Agent activity metrics (Paperclip runs, success rate, token spend) * [x] Public service availability monitoring
- [x] SugeBot command log panel * [x] VPS resource health (CPU, Memory, Disk, Network)
- [x] Per-agent health monitoring * [x] HTTP status code monitoring
- [x] n8n workflow execution tracking * [x] Endpoint response time tracking
- [x] Telegram alert integration via SugeBot * [x] SSL certificate monitoring
* [x] Cross-environment scraping via Tailscale
### V2 Dashboard — Operational Reliability
* [x] SSL expiration tracking
* [x] Service reliability scoring
* [x] Public service availability metrics
* [x] Latency monitoring
* [x] Capacity awareness
* [x] Prometheus scrape visibility
* [x] Infrastructure health scoring
### V3 Dashboard — SRE / Incident Response
* [x] Incident command overview
* [x] Service dependency monitoring
* [x] Alert timeline visibility
* [x] SLO/SLA monitoring
* [x] Error budget awareness
* [x] Incident response workflow
* [x] Infrastructure pressure detection
* [x] Reliability engineering metrics
--- ---
## 🔮 Roadmap ## 🔮 Roadmap
- [ ] AI Incident Detection — anomaly-triggered agent escalation ### V4 — AI-Assisted Operations
- [ ] Automated Remediation — n8n workflows triggered by Prometheus alerts
- [ ] Multi-environment monitoring (homelab + VPS unified NOC) * [ ] AI Incident Detection
- [ ] Discord alert channel * [ ] Automated Remediation (n8n workflows triggered by alerts)
* [ ] Telegram alert integration via SugeBot
* [ ] Discord notifications
* [ ] Paperclip agent telemetry dashboard
* [ ] Agent execution analytics
* [ ] Token spend / cost observability
* [ ] Multi-environment monitoring (Homelab + VPS unified NOC)
--- ---
@@ -277,7 +331,8 @@ Digital Systems Engineer | Platform Engineering | AI Infrastructure | Observabil
## 🔗 Related Repos ## 🔗 Related Repos
| Repo | Description | | Repo | Description |
|---|---| | ------------------------------------------------------------- | ------------------------------------------------------------------ |
| [ceds-homelab](https://github.com/ced4568/ceds-homelab) | 6-node Proxmox cluster + 12-node K3s + full homelab infrastructure | | [ceds-homelab](https://github.com/ced4568/ceds-homelab) | 6-node Proxmox cluster + 12-node K3s + full homelab infrastructure |
| [ced-k3s-homelab](https://github.com/ced4568/ced-k3s-homelab) | 12-node Raspberry Pi K3s cluster detail | | [ced-k3s-homelab](https://github.com/ced4568/ced-k3s-homelab) | 12-node Raspberry Pi K3s cluster detail |
| [ceds-observability-stack](https://github.com/ced4568) | Observability stack configs and dashboards |
| [ceds-noc](https://github.com/ced4568/ceds-noc) | Custom-built public NOC status page | | [ceds-noc](https://github.com/ced4568/ceds-noc) | Custom-built public NOC status page |
Binary file not shown.

After

Width:  |  Height:  |  Size: 214 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 313 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 231 KiB