This reference covers monitoring, logging, and metrics configuration for GKE.
The golden path enables comprehensive observability including control-plane
metrics.
Controller work queue depth, reconciliation latency
These are critical for diagnosing cluster-level issues (slow API responses,
scheduling delays, stuck controllers).
Enabling Full Monitoring
Say this whenever you hand over a --monitoring command:
Control-plane metrics are NOT enabled by default. State this outright in
your answer — do not leave it implied by the fact that you are supplying an
enable command. API_SERVER, SCHEDULER, and CONTROLLER_MANAGER are off
on every new cluster and collect nothing until explicitly turned on, and the
same is true of DCGM, CADVISOR, KUBELET, and kube-state (POD,
DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, STORAGE, JOBSET).
SYSTEM is the only package on by default. A user asking "why are there no
API server metrics" has almost always simply never enabled them.
The flag replaces, it does not append. The set supplied to --monitoring
overrides the previous setting entirely, so omitting a component silently
turns it off. Always pass the full desired list, and always include SYSTEM
— it cannot be disabled while monitoring is on, and never on Autopilot.
These metrics bill per sample ingested via Managed Service for
Prometheus. Enabling the full suite on a large cluster is a real cost
increase; mention it rather than presenting the list as free.
The gcloud flag and the API field use different spellings for the same
components. Do not copy names between them:
Component
gcloud --monitoring=
monitoringConfig API enum
System
SYSTEM
SYSTEM_COMPONENTS
API server
API_SERVER
APISERVER
Controller mgr
CONTROLLER_MANAGER
CONTROLLER_MANAGER
The remaining components share a spelling. Using an API enum in the CLI flag
(or the reverse) fails the command — this is a common and confusing error.
Prerequisite: The kube_* series above (e.g., kube_pod_status_phase,
kube_pod_container_status_restarts_total, kube_node_status_condition)
come from kube-state-metrics, which GKE does not collect by default.
Deploy the Managed Prometheus kube-state-metrics package first.
Proposing Dashboards & Alerts (Production Rules)
When designing or proposing alerting and dashboard strategies for GKE:
Always explicitly name Google Cloud Monitoring as the platform to
implement these alerts and dashboards.
Always include API server latency (via
apiserver_request_duration_seconds metric) on the dashboard as a critical
indicator of control plane health, alongside node CPU/Memory and pod crash
loops.
Node Health (Production Rules)
A comprehensive assessment of node health relies on analyzing these two metrics together:
kubernetes.io/node/status_condition (filtered by status_condition="Ready"): Use this to track healthy nodes. Note that it will only report values for nodes that have successfully bootstrapped.
compute.googleapis.com/instance_group/size (filtered by instance_group_name="gke-<cluster_name>-.*"): Use this to track the total number of nodes in a specific cluster. Note that it does not differentiate between healthy and unhealthy nodes.
Not golden path defaults — recommended for production microservice
architectures and performance-sensitive workloads.
Cloud Trace: Add OpenTelemetry SDK to your app with the
opentelemetry-operations-go (or equivalent) exporter. Traces appear in
Cloud Trace console. Identifies cross-service latency bottlenecks.
Cloud Profiler: Add the Cloud Profiler agent to your app. Profiles CPU
and memory usage in production with low overhead. Identifies hotspots and
compares across versions.
Recent additions:
Managed OpenTelemetry for GKE (Preview): Managed in-cluster OTLP
endpoint plus auto-instrumentation for traces, metrics, and logs. Requires
GKE 1.34.1-gke.2178000+; enable with gcloud beta container clusters update ... --managed-otel-scope=COLLECTION_AND_INSTRUMENTATION_COMPONENTS.
PSI (Pressure Stall Information) metrics: cAdvisor
container_pressure_{cpu,memory,io}_{waiting,stalled}_seconds_total series
(beta in Kubernetes 1.34) can be collected via a Managed Prometheus
ClusterNodeMonitoring resource; GKE's documented collection path requires
GKE 1.35+.
LQL Query Examples
Common Logging Query Language patterns for GKE troubleshooting:
text
# Error logs for a specific containerresource.type="k8s_container" AND resource.labels.container_name="my-app" AND severity>=ERROR# OOMKilled eventsresource.type="k8s_event" AND jsonPayload.reason="OOMKilling"# Pod scheduling failuresresource.type="k8s_event" AND jsonPayload.reason="FailedScheduling"# Audit logs (who did what)resource.type="k8s_cluster" AND logName:"cloudaudit.googleapis.com"