Documentation

Observability

A built-in Prometheus endpoint for control plane health, ready-made Grafana dashboards, logs and profiling.

Built-in metrics endpoint

KubeSolo has its own Prometheus endpoint that reports the health of the control plane it runs: whether each component is up, when it became ready, certificate expiry and the size of the database. It is off by default and listens on localhost when enabled:

yaml
1metrics:
2 enabled: true
3 bindAddress: 127.0.0.1:9105 # 0.0.0.0:9105 to scrape from another machine
bash
1$ sudo kubesoloctl config set metrics.enabled true
2$ sudo systemctl restart kubesolo
3$ curl -s http://127.0.0.1:9105/metrics | grep kubesolo_component_up

The endpoint serves /metrics and /healthz over plain HTTP with no authentication, so bind it to a non-local address only on a network you trust. Probes run every 15 seconds.

Metric reference

MetricTypeLabelsMeaning
kubesolo_component_upgaugecomponent1 if the component is reachable and healthy, 0 otherwise.
kubesolo_component_ready_timestamp_secondsgaugecomponentWhen the component first signalled readiness. 0 until ready.
kubesolo_component_last_probe_timestamp_secondsgaugecomponentWhen the component was last probed.
kubesolo_certificate_expiry_timestamp_secondsgaugenameWhen the certificate expires. Absent if it cannot be read.
kubesolo_certificate_validgaugename1 if the certificate is readable and inside its validity window.
kubesolo_kine_db_size_bytesgaugeSize of the SQLite database file.
kubesolo_build_infogaugeversion, commit, build_date, go_version, archAlways 1; build metadata in the labels.
kubesolo_start_time_secondsgaugeWhen the metrics endpoint started.
kubesolo_uptime_secondsgaugeSeconds since the metrics endpoint started.

The standard Go runtime (go_*) and process (process_*) collectors are included too.

LabelValues
componentruntime, kine, apiserver, controller, kubelet, kubeproxy, coredns, webhook
name (certificates)ca, apiserver, controller-manager, kubelet, admin, webhook, request-header-ca, request-header-client, plus d2k-server and d2k-client when d2k is enabled

Scraping and alerting

prometheus.yml
1scrape_configs:
2 - job_name: kubesolo
3 static_configs:
4 - targets: ["192.168.1.20:9105"]
PromQL
1# a component is down
2kubesolo_component_up == 0
 
4# a certificate expires within 30 days
5kubesolo_certificate_expiry_timestamp_seconds - time() < 30 * 24 * 3600

Grafana dashboards

The repository ships a KubeSolo Cluster Overview dashboard in examples/grafana. It covers the cluster and node overview, API server request rate, errors and latency, admission webhook latency, kubelet pods and containers, per-pod CPU, memory, network and throttling from cAdvisor, controller workqueues, and Go runtime health.

The dashboard reads the Kubernetes components' own metrics, not the KubeSolo endpoint above. Its data comes from two scrape targets: the API server's /metrics, and kubelet plus cAdvisor metrics through the API server's node proxy.

FileUse
kubesolo-grafana-dashboard-classic.jsonClassic dashboard JSON. Recommended for Grafana OSS, Enterprise and Grafana Cloud.
kubesolo-grafana-dashboard-apiv2.jsonThe same dashboard in the dashboard.grafana.app/v2 format, for Grafana instances that support Dashboard API v2.
alloy-config.yamlGrafana Alloy deployment that scrapes both targets and remote-writes to a Prometheus-compatible endpoint, with the RBAC it needs.
otel-collector-config.yamlThe same pipeline with the OpenTelemetry Collector. Needs the contrib distribution.

To import, in Grafana go to Dashboards → New → Import, upload the JSON and select your Prometheus data source for ${datasource}. If Grafana reports unsupported apiVersion/kind, use the classic file.

bash
1$ kubectl apply -f https://raw.githubusercontent.com/portainer/kubesolo/v1.2.1/examples/grafana/alloy-config.yaml

Then set GRAFANA_CLOUD_URL, GRAFANA_CLOUD_USERNAME and GRAFANA_CLOUD_API_KEY on the Alloy deployment, ideally from a Kubernetes Secret, or change its prometheus.remote_write block to point at your own endpoint.

Logs

KubeSolo logs every component to one stream. With systemd:

bash
1$ journalctl -u kubesolo -f
2$ journalctl -u kubesolo | grep configapi # configuration API audit lines

On OpenRC the service logs to /var/log/messages, on SysV init to /var/log/syslog, and in daemon mode to /var/log/kubesolo.log. Set logging.debug: true for debug output.

Profiling with pprof

logging.pprof: true (or --pprof-server=true at install) starts Go's pprof server on port 6060 on all interfaces, without authentication. Enable it only while profiling, and firewall the port.

bash
1$ curl -s http://localhost:6060/debug/pprof/ | head
2$ go tool pprof http://localhost:6060/debug/pprof/heap