GreenKube

GreenKube Architecture

This document describes the technical architecture of GreenKube. The goal is to create a lightweight, modular, and extensible platform to measure, report, and optimize the carbon footprint and cost of Kubernetes workloads.

Architecture Diagram

flowchart TB
    subgraph K8s["Kubernetes Cluster"]
        Prom["Prometheus"]
        OC["OpenCost"]
        K8sAPI["K8s API"]
        EM["Electricity Maps API"]
        WN["Wattnet API"]
        Boavizta["Boavizta API"]
    end

    subgraph GreenKube["GreenKube"]
        direction TB
        subgraph Collectors["Collectors (Input Adapters)"]
            PC["PrometheusCollector"]
            OCC["OpenCostCollector"]
            NC["NodeCollector"]
            PodC["PodCollector"]
            EMC["ElectricityMapsCollector"]
            WNC["WattnetCollector"]
            BC["BoaviztaCollector"]
        end

        subgraph Core["Core (Business Logic)"]
            DP["DataProcessor<br/>(Facade)"]
            CO["CollectionOrchestrator<br/>(Parallel Fetch)"]
            Est["BasicEstimator<br/>(CPU → Joules)"]
            Calc["CarbonCalculator<br/>(Joules → CO₂e)"]
            MA["MetricAssembler<br/>(Build CombinedMetric)"]
            NZM["NodeZoneMapper<br/>(Cloud → eMaps Zone)"]
            PRM["PrometheusResourceMapper<br/>(Per-Pod Resources)"]
            CN["CostNormalizer<br/>(Per-Step Cost)"]
            ESvc["EmbodiedEmissionsService<br/>(Boavizta Cache + Fallback)"]
            SR["SummaryRefresher<br/>(Pre-computed Cache)"]
            HRP["HistoricalRangeProcessor<br/>(Chunked Range)"]
            Rec["Recommender<br/>(9 Types)"]
        end

        subgraph Storage["Storage (Output Adapters)"]
            PG["PostgreSQL"]
            SQLite["SQLite"]
        end

        subgraph Presentation["Presentation"]
            API["FastAPI<br/>REST API"]
            CLI["Typer CLI"]
            Dash["SvelteKit<br/>Dashboard"]
            Grafana["Grafana<br/>(via Prometheus)"]
        end
    end

    Prom --> PC
    OC --> OCC
    K8sAPI --> NC
    K8sAPI --> PodC
    EM --> EMC
    WN --> WNC
    Boavizta --> BC

    PC --> CO
    OCC --> CO
    PodC --> CO
    NC --> DP

    CO --> DP
    EMC --> DP
    WNC --> DP
    BC --> ESvc

    DP --> Est
    DP --> NZM
    DP --> PRM
    DP --> CN
    DP --> MA
    DP --> HRP
    DP --> SR

    MA --> Calc
    MA --> ESvc
    MA --> Storage
    SR --> Storage

    Storage --> API
    Storage --> CLI
    API --> Dash
    API --> Grafana
    Storage --> Rec
    Rec --> API
    Rec --> CLI

Overview

GreenKube operates as an asynchronous agent that collects, processes, analyzes, and reports data. It runs as both a scheduled service (continuous monitoring) and an on-demand CLI tool (ad-hoc reporting).

The system is designed around the principles of Clean Architecture and Hexagonal Architecture:

All I/O operations use Python’s asyncio for high-performance, non-blocking concurrent execution.

Architectural Principles

1. Database Agnosticism

The core business logic (src/greenkube/core) NEVER depends on a specific storage implementation. This is enforced through:

Supported backends:

2. Cloud Provider Agnosticism

Cloud-specific details (AWS, GCP, Azure, OVH, Scaleway) are isolated in:

3. Asynchronous & Non-Blocking

Core Components

Collectors (Input Ports)

All collectors are fully asynchronous and implement a common pattern.

PrometheusCollector

NodeCollector

PodCollector

OpenCostCollector (optional)

ElectricityMapsCollector

WattnetCollector

BaseElectricityProvider

BoaviztaCollector

Estimator (Business Logic)

BasicEstimator

Processor (Use Case Orchestrator)

DataProcessor

The main orchestrator that coordinates the data pipeline from collection to metric assembly.

Architecture: The processor acts as a facade, delegating specialized work to focused collaborators while managing the overall pipeline flow.

Key Responsibilities:

  1. Data Collection: Coordinates parallel collection from Prometheus, Kubernetes, OpenCost, and external APIs
  2. Energy Estimation: Converts resource usage into energy consumption (Joules)
  3. Zone Mapping: Resolves cloud regions to carbon intensity zones
  4. Carbon Calculation: Computes CO2e emissions from energy and grid intensity
  5. Metric Assembly: Combines energy, cost, resources, and metadata into unified metrics
  6. Embodied Emissions: Integrates hardware manufacturing emissions

Pipeline Stages:

Implementation Pattern:

Calculator (Business Logic)

CarbonCalculator

SummaryRefresher (Business Logic)

SummaryRefresher

Recommender (Business Logic)

Recommender / RecommenderV2

Analyzes CombinedMetric data to identify optimization opportunities.

Recommendation Types:

  1. Zombie Pods: Workloads consuming resources but producing minimal value
    • Criteria: Low CPU/energy, high cost, extended idle time
    • Savings: Potential cost and emission reduction
  2. Rightsizing:
    • Over-provisioned CPU: Request » actual usage
    • Over-provisioned memory: Request » actual usage
    • Headroom calculation for safe downsizing
    • Savings estimate based on cloud provider pricing
  3. Autoscaling Candidates:
    • High coefficient of variation (CV) in usage
    • Spike detection (max/avg ratio)
    • HPA/VPA recommendations
  4. Carbon-Aware Scheduling:
    • Identifies high-carbon-intensity periods
    • Suggests workload time-shifting for batch jobs
  5. Idle Namespace Cleanup:
    • Namespaces with minimal activity
    • Low-value resource consumption

Configuration: All thresholds configurable via config.py and Helm values

Repositories (Output Ports)

Repositories use asynchronous drivers for high-performance database interactions. All implement abstract base classes to ensure database agnosticism.

CarbonIntensityRepository

Abstract base class for carbon intensity storage.

Implementations:

NodeRepository

Abstract base class for node state snapshots.

Implementations:

CombinedMetricRepository

Stores final aggregated metrics (energy + carbon + cost + resources).

Schema (33 columns):

Migrations: All backends support automatic schema evolution (ADD COLUMN IF NOT EXISTS)

EmbodiedRepository

Caches Boavizta API responses for hardware embodied emissions.

Schema: provider, instance_type, gwp_manufacture, lifespan_hours, last_updated

SummaryRepository

Stores pre-computed KPI scalar totals per window.

Schema: window_slug, namespace, total_co2e_grams, total_embodied_co2e_grams, total_cost, total_energy_joules, pod_count, namespace_count, updated_at

TimeseriesCacheRepository

Stores pre-computed time-series buckets per window and granularity.

Schema: window_slug, namespace, bucket_ts, co2e_grams, embodied_co2e_grams, total_cost, joules

API & Presentation Layer

FastAPI Server

Endpoints:

Report export query parameters:

SvelteKit Dashboard

Pages:

Features:

CLI

Grafana Integration

Data Flow

Instant Collection (run())

Used for real-time monitoring and scheduled collection.

  1. Phase 1 — Node Discovery (sequential, single K8s API call):
    NodeCollector.collect()
    └─ Node metadata (instance type, zone, capacity, provider)
    └─ Build node_instance_map directly from Phase-1 data
    
  2. Phase 2 — Zone Resolution (depends on Phase-1 data):
    NodeZoneMapper.map_nodes(nodes_info)
    └─ Map cloud region → Electricity Maps zone per node
    
  3. Phase 3 — Parallel Collection + Boavizta (concurrent, Phase-1/2 data used for enrichment):
    asyncio.gather(
      ┌─ CollectionOrchestrator.collect_all(nodes_info):
      │   ├─ PrometheusCollector.collect()
      │   │   ├─ CPU usage (8 concurrent queries)
      │   │   ├─ Memory usage
      │   │   ├─ Network I/O (rx + tx)
      │   │   ├─ Disk I/O (read + write)
      │   │   ├─ Restart counts
      │   │   └─ Node labels (enriched with Phase-1 data)
      │   ├─ OpenCostCollector.collect()
      │   │   └─ Cost allocation data
      │   └─ PodCollector.collect()
      │       └─ Resource requests (CPU, memory, storage)
      └─ EmbodiedEmissionsService.prepare_embodied_data(nodes_info):
          ├─ Check EmbodiedRepository cache
          ├─ Fetch missing profiles from Boavizta API
          └─ Inject fallback profile (DEFAULT_EMBODIED_EMISSIONS_KG) for unknowns
    )
    
  4. Phase 4 — Assembly:
    ├─ BasicEstimator.estimate() → EnergyMetric per pod (Joules)
    ├─ MetricAssembler.prefetch_intensities()
    │   ├─ Group pods by zone
    │   └─ Prefetch intensity for (zone, timestamp) pairs
    └─ MetricAssembler.assemble()
        ├─ CarbonCalculator.calculate_emissions()
        ├─ EmbodiedEmissionsService.calculate_pod_embodied()
        │   └─ Mark metric is_estimated=True if fallback used
        └─ Build CombinedMetric (energy + carbon + cost + resources + metadata)
    
  5. Persistence:
    Repository.write_combined_metrics()
    └─ Batch insert to database (Postgres/SQLite)
    
  6. Background (hourly scheduler):
    SummaryRefresher.run()
    ├─ For each window (24h, 7d, 30d, 1y, ytd):
    │   ├─ aggregate_summary() → MetricsSummaryRow → upsert metrics_summary
    │   └─ aggregate_timeseries() → TimeseriesCachePoint[] → upsert metrics_timeseries_cache
    └─ Repeated for each namespace
    

Historical Analysis (run_range())

Used for reporting over time ranges with historical accuracy.

  1. Historical Node State Loading:
    NodeRepository.get_latest_snapshots_before(start)
    NodeRepository.get_snapshots(start, end)
    └─ Reconstruct node timeline
    
  2. Chunked Processing (day-sized chunks to prevent OOM):
    For each chunk (1 day):
      ├─ Prometheus range queries (concurrent):
      │   ├─ CPU usage over time
      │   ├─ Network I/O over time
      │   ├─ Disk I/O over time
      │   ├─ Restart counts over time
      │   └─ (5 queries via asyncio.gather)
      │
      ├─ Parse time-series data:
      │   └─ Build per-pod resource maps from range results
      │
      ├─ Energy estimation per timestamp:
      │   └─ Use historical node profile at each timestamp
      │
      ├─ Carbon calculation:
      │   └─ Historical intensity lookup
      │
      └─ Generate CombinedMetrics for chunk
    
  3. Aggregation:
    Collect all chunks → Filter by namespace → Return
    

Recommendation Flow

  1. Data Collection:
    Read CombinedMetrics from repository (last N days)
    
  2. Analysis:
    RecommenderV2.generate_recommendations()
    ├─ Aggregate pod metrics by stable workload owner when available
    ├─ Calculate statistics (weighted mean, observed max, percentile, CV)
    ├─ Apply thresholds:
    │   ├─ Zombie detection
    │   ├─ Rightsizing analysis using average and retained maximum usage
    │   ├─ Autoscaling candidates
    │   └─ Carbon-aware opportunities
    ├─ Upsert active recommendations by full target identity
    ├─ Mark previously active recommendations as stale when absent from the latest generation
    └─ Calculate savings (cost + CO2e)
    
  3. Output:
    Return List[Recommendation] with:
    ├─ Type and severity
    ├─ Affected resources
    ├─ Current vs. recommended
    ├─ Estimated savings
    └─ Actionable commands
    

Configuration & Deployment

Environment Variables (12-Factor App)

All configuration flows through src/greenkube/core/config.py:

Database:

External Services:

Collection:

Carbon:

Recommendations:

Helm Chart Structure

helm-chart/
├── Chart.yaml              # Chart metadata
├── values.yaml             # Default configuration
└── templates/
    ├── deployment.yaml     # GreenKube deployment
    ├── service.yaml        # API service (LoadBalancer/ClusterIP)
    ├── configmap.yaml      # Non-secret config
    ├── secret.yaml         # Tokens and credentials
    ├── postgres-*.yaml     # PostgreSQL StatefulSet (optional)
    ├── serviceaccount.yaml # K8s RBAC
    ├── clusterrole.yaml    # Read permissions for nodes/pods
    └── post-install-hook.yaml  # Schema initialization

Container Structure (Multi-Stage Build)

Dockerfile (3 stages):
1. Frontend Builder (Node 20):
   ├─ npm ci (install deps)
   ├─ npm run build (SvelteKit SSG)
   └─ Output: static files in build/

2. Python Builder (Python 3.14):
   ├─ pip install build
   ├─ Build wheel from pyproject.toml
   └─ pip install to /install prefix

3. Final Image (Python 3.14-slim):
   ├─ Copy Python packages from builder
   ├─ Copy frontend build from frontend-builder
   ├─ Run as non-root user 'greenkube'
   └─ ENTRYPOINT: greenkube CLI

Normalization and Caching

Timestamp Normalization

Controlled by NORMALIZATION_GRANULARITY:

Caching Strategy

CarbonCalculator Cache:

Repository Cache:

Boavizta Cache:

Dashboard Summary Cache:

Performance Considerations

Concurrency

Memory Management

Scalability

Testing Strategy

Unit Tests

Integration Tests

Test Patterns

Design Principles & Guidelines

For Contributors

Database Agnosticism (Critical)

Error Handling

Async Best Practices

Code Quality

Logging Levels

Commit Messages

Follow conventional commits:

Future Architecture Evolution

Planned Enhancements

Multi-Cluster Support:

Advanced Analytics:

Real-Time Streaming:

Extended Metrics:

Enhanced Reporting:

Extension Points

New Collectors:

  1. Implement collector interface
  2. Add to core/factory.py
  3. Update DataProcessor to use it
  4. Add tests

New Storage Backend:

  1. Implement CarbonIntensityRepository abstract
  2. Add to core/factory.py mapping
  3. Add schema migration logic
  4. Add tests

New Cloud Provider:

  1. Add region mappings to utils/region_mapping.py
  2. Add instance profiles to energy/instance_profiles.py
  3. Add PUE values to core/config.py
  4. Update tests

Resources

License

Apache 2.0 - See LICENSE file for details.