Troubleshooting error message reference

This reference lists all documented error messages and error codes across troubleshooting topics for IBM Cloud Kubernetes Service. Entries are grouped by component and link to the full troubleshooting topic.

Clusters and masters

Cluster and master error messages
Error message Troubleshooting topic
Cannot complete cluster master operations because the cluster has a broken webhook application. Why do cluster master operations fail due to a broken webhook?
The master is approaching its allotted memory resource limit (93%). Why does my cluster master status say it is approaching its resource limit?
etcd database size is approaching the maximum Why do I see an etcd database size is approaching the maximum error?
The 'configuration' field is not a valid Kubernetes PodSecurityConfiguration setting. Why do I get an error that my PodSecurityConfiguration is not valid?
No VPC is available. Create a VPC. VPC: Why is no VPC available when I create a cluster in the console?
Your cluster can't pull images from the 'icr.io' domains because an IAM access policy could not be created. Why can't the cluster pull images from IBM Cloud Container Registry during creation?
Image security enforcement update canceled. CAE008: can't enable Portieris image security enforcement because the cluster already has a conflicting image admission controller installed. Why is my Portieris cluster image security enforcement installation canceled?
incorrect account for worker, Worker deploy failed due to network communications failing, Unable to connect to the IBM Cloud account. Why can't I create or delete clusters or worker nodes?
Unable to create cluster. The 'vpc-gen2' infrastructure operation failed with the message: the provided token is not authorized to view the specified subnet Why do I get an infrastructure operation failed error when creating a VPC cluster?
No resources found., connection timed out, dial tcp: connect: connection timed out Debugging common CLI issues with clusters
Encrypted storage cannot be configured. Review the customer root key configuration for the worker pool. Why can't I create a VPC cluster with encrypted worker nodes?
Pending security group creation When I create a VPC cluster, my worker nodes are stuck in Pending security group creation
Infrastructure instance status is 'failed': Can't start instance because provisioning failed. Why do I see DNS failures after adding a custom DNS resolver?
Version update canceled. CAE009: Cannot complete cluster master operations because the cluster does not pass Pod Security upgrade prerequisites. Why does my cluster upgrade fail due to Pod Security upgrade prerequisites?
Cannot complete cluster master operations because there is a migration in progress Resolving cluster master upgrade issues: Migration in progress error

Worker nodes

Worker node error messages
Error message Troubleshooting topic
The worker node instance ID changed. Reload the worker node if bare metal hardware was serviced. Classic: Why is the bare metal instance ID inconsistent with worker records?
The dedicated hosts for the zone 'eu-de-2' are not ready. VPC: Why can't I create worker nodes on dedicated hosts?
SoftLayerAPIError(SoftLayer_Exception_Public): Could not obtain network VLAN with id #123456. Classic: Why can't I add worker nodes with an invalid VLAN ID?
Registration failed – The plan containers.kubernetes.vpc.gen2.roks is not available in <region>. Why do I see a Registration failed error when I try to provision or reload worker nodes?
A VSI with this profile will put user over quota. VPC worker nodes fail to provision due to quota limits
warning: Container container-00 is unable to start due to an error: Back-off pulling image "registry.redhat.io/rhel8/support-tools" After creating a version 1.30 cluster, my app no longer works

Network health check (NHC) errors

The following error codes appear in the output of the ibmcloud ks cluster health issues command.

Network health check (NHC) error codes
Error code Severity Description Troubleshooting topic
NHC001 Warning Tigera operator has been reporting that Calico is in 'progressing' state for over an hour. Why does the Network status show an NHC001 error?
NHC003 Warning Some worker nodes in the cluster can not reach container image registries to pull images. Why does the Network status show an NHC003 error?
NHC004 Warning Some worker nodes in the cluster can not resolve VPE gateway hostnames. Why does the Network status show an NHC004 error?
NHC005 Warning Tigera operator is reporting that Calico is in 'degraded' state. Why does the Network status show an NHC005 error?
NHC006 Warning One or more DNS resolvers are not reachable from certain worker nodes. Why does the Network status show an NHC006 error?
NHC007 Warning One or more DNS resolvers are not reachable from certain worker nodes. Why does the Network status show an NHC007 error?
NHC009 Error The IAM token exchange request failed. Why does the Network status show an NHC009 error?
NHC010 Error Exceeded security group rules related quota. Why does the Network status show an NHC010 error?
NHC011 Error Exceeded security group related quota. Why does the Network status show an NHC011 error?

Ingress status errors (ERR and ESS codes)

The following error codes appear in the output of the ibmcloud ks ingress status-report get command. Shared codes appear in both IBM Cloud Kubernetes Service and Red Hat OpenShift on IBM Cloud.

Ingress status error codes (ERR and ESS)
Error code Error message Troubleshooting topic
ERRDSIA The subdomain has incorrect addresses registered. Ingress error: ERRDSIA
ERRDRISS The subdomain has DNS resolution issues. Ingress error: ERRDRISS
ERRDSAISS The external provider for the given subdomain has authorization issues. Ingress error: ERRDSAISS
ERRDSISS The subdomain has TLS secret issues. Ingress error: ERRDSISS
ERRSAM The load balancer service address is missing. Ingress error: ERRSAM
ESSDNE The secret is not present on the cluster or is in the wrong namespace. Ingress error: ESSDNE
ESSEC The certificate for TLS secret expired or will expire soon. Ingress error: ESSEC
ESSEF The Opaque secret field expired or will expire soon. Ingress error: ESSEF
ESSSMG Could not find the secret group. Ingress error: ESSSMG
ESSSMI Could not access Secrets Manager instance. Ingress error: ESSSMI
ESSSMINF The Secrets Manager instance is not found. Ingress error: ESSSMINF
ESSVC The CRN does not match the default secret with the same domain. Ingress error: ESSVC
ESSWS The secret status shows a warning. Ingress error: ESSWS
ERRADNF The ALB deployment is not found on the cluster. Ingress error: ERRADNF
ERRADRUH One or more ALB pods are not in the running state. Ingress error: ERRADRUH
ERRAHCF The ALB is unable to respond to health requests. Ingress error: ERRAHCF
ERRAHINF One or more ALB health Ingress resource is not found on the cluster. Ingress error: ERRAHINF
ERRAHSNF One or more ALB health service is not found on the cluster. Ingress error: ERRAHSNF
ERRAVUS The ALB version is no longer supported. Ingress error: ERRAVUS
ERRHPAETPI Autoscaling is ineffective. Ingress error: ERRHPAETPI
ERRHPAIWC The cluster does not have enough worker nodes to satisfy the autoscaling configuration. Ingress error: ERRHPAIWC
ERRHPANA Autoscaling is failing. Ingress error: ERRHPANA
ERRHPANF The autoscaler resource is missing. Ingress error: ERRHPANF
ERRICCNF The Ingress controller ConfigMap is not found on the cluster. Ingress error: ERRICCNF
ERRSNF The load balancer service is missing. Ingress error: ERRSNF
0/3 nodes are available: 1 node(s) didn't match pod affinity/anti-affinity Why do ALB pods not deploy to worker nodes?
No valid subnets found for the specified zone. Classic clusters: Why does enabling Ingress ALBs result in subnet errors?
admission webhook "validate.nginx.ingress.kubernetes.io" denied the request: nginx.ingress.kubernetes.io/configuration-snippet annotation cannot be used. Ingress resource operations refused by validating webhook

Load balancers

Load balancer error messages
Error message Troubleshooting topic
The VPC load balancer that routes requests to this Kubernetes LoadBalancer service is offline. VPC clusters: Why can't my app connect via load balancer?
The subnet with ID(s) '<subnet_id>' has insufficient available ipv4 addresses. VPC clusters: Why does a Kubernetes LoadBalancer service fail with no IPs?
The load balancer was created in zone <zone>. This setting cannot be changed. VPC Clusters: My VPC NLB has a zone error and does not update
Warning CreatingCloudLoadBalancerFailed ... Failed ensuring LoadBalancer: FindLoadBalancer failed ... 401 Unauthorized ... BXNIM0430E Why do I see SyncLoadBalancerFailed errors when creating a VPC cluster?
Security group protocol mismatch events on load balancer creation or update VPC clusters: Security group protocol error creating or updating a LoadBalancer

Apps and services

App and service error messages
Error message Troubleshooting topic
Failed to create pod sandbox: rpc error: ... failed to request 1 IPv4 addresses. IPAM allocated only 0 Why don't my containers start?
ImagePullBackOff or image pull authorization errors Why do images fail to pull from registry with ImagePullBackOff or authorization errors?
pull QPS exceeded errors during image pulls Why do pods show pull QPS exceeded errors during image pulls?
Error: failed to download "<helm_repo>/<chart_name>" Troubleshooting helm chart installation updated configuration values
This service doesn't support creation of keys Resolving service binding errors in IBM Cloud clusters
Pod remains in Pending state Why do pods remain in pending state?
Pod repeatedly fails to restart or is unexpectedly removed Why do pods repeatedly fail to restart or are unexpectedly removed?
unable to validate against any pod security policy Why do my pods fail to deploy after applying a pod security policy?
Cluster or service instance already exists with the same name Why does binding a service to a cluster result in a same name error?
The server is currently unable to handle the request (get pods.metrics.k8s.io) Troubleshooting metrics server issues in Kubernetes clusters

Permissions and credentials

Permission and credential error messages
Error message Troubleshooting topic
User doesn't have permissions to create or manage Storage What permissions do I need to manage storage and create PVCs?

Secure by default (SBD)

Secure by default (SBD) error messages
Error message Troubleshooting topic
warning: Container container-00 is unable to start due to an error: Back-off pulling image "registry.redhat.io/rhel8/support-tools" After creating a version 1.30 cluster, my app no longer works
Pending security group creation When I create a VPC cluster, my worker nodes are stuck in Pending security group creation
Infrastructure instance status is 'failed': Can't start instance because provisioning failed. Why do I see DNS failures after adding a custom DNS resolver?
Nodeport apps not working after updating to version 1.30 Fixing nodeport apps after updating cluster version 1.30 or later
Other clusters in the VPC failing after creating a version 1.30 cluster After creating a version 1.30 cluster, applications running in other clusters in my VPC are failing
VSIs cannot access VPE gateway Why can't my VSIs access VPE gateway?

File Storage

Error message Troubleshooting topic
MountVolume.SetUp failed for volume ... mount.nfs: access denied by server while mounting Classic: Why am I denied server access when mounting a volume to a worker node?
write-permission or non-root user ownership errors on NFS mount path Why does my app fail when a non-root user owns the NFS file storage mount path?
Group ID error applying NFS file storage permissions Why does my app fail with a group ID error for NFS file storage permissions?
Non-root user cannot add access to persistent storage Why can't I add non-root user access to persistent storage?
File systems for worker nodes changed to read-only Why are the file systems for worker nodes changed to read-only?
PVC remains in pending state (file storage) Why does my file storage PVC stay in a pending state?
MetadataServiceNotEnabled Why do I see a MetadataServiceNotEnabled error for File Storage for VPC?
MountingTargetFailed or rpc error: code = DeadlineExceeded desc = context deadline exceeded Why do I see a MountingTargetFailed error for File Storage for VPC?
SubnetFindFailed or rpc error: code = FailedPrecondition on PVC creation Why does PVC creation fail for File Storage for VPC?
UnresponsiveMountHelperContainerUtility Why do I see an UnresponsiveMountHelperContainerUtility error for File Storage for VPC?
shares_snapshot_operation_not_allowed Why can't I create File Storage for VPC snapshots?
shares_snapshot_not_found on PVC restore Why can't I restore my File Storage for VPC snapshot to a PVC?
VPC File Storage snapshot cannot be deleted Why can't I delete my File Storage for VPC snapshot?
'rfs' profile is not accessible or stunnel manager is not initialized Troubleshooting Regional File Storage encryption in transit

| VPC File Storage deployment permissions error | Why does my File Storage for VPC deployment fail due to a permissions error? | | App pod stuck in Container creating when mounting VPC File Storage | Why is my app pod stuck in Container creating when trying to mount File Storage for VPC? | | File Storage add-on in Critical state | Why is the File Storage for VPC add-on in Critical state? |

Block Storage

Block Storage error messages
Error message Troubleshooting topic
failed to mount the volume as "ext4", it already contains xfs. Mount error: mount failed: exit status 32 Why does mounting existing block storage to a pod fail with the wrong file system?
Volume not attached Why do I get a Volume not attached error when trying to expand a Block Storage for VPC volume?
Block storage changes to read-only Why does block storage change to read-only?
Message: 50% throttling of CPU in namespace kube-system for container ibmcloud-block-storage-driver-container Why does the Block storage plug-in Helm chart give CPU throttling warnings?
Block storage PVC remains in pending state Why does my block storage PVC stay in a pending state?
App cannot access or write to PVC (block) Why can't my app access or write to a PVC?
Labels: ibm.io/pv-connectivity-status: limited Why does my Block Storage persistent volume show a limited connectivity status?
Block Storage API key reset causes provisioning failure Block Storage for VPC PVC creation fails after API key reset
UNEXPECTED INCONSISTENCY; RUN fsck MANUALLY. Why does mounting Block Storage for Classic fail with a file system check error?
Block storage volume snapshot cannot be deleted Why can't I delete my Block Storage for VPC volume snapshot resources?
Block storage snapshot creation fails Why can't I create Block Storage for VPC snapshots?
Charges still appear for block storage devices after deleting the cluster Why am I still seeing charges for block storage devices after deleting my cluster?

Object Storage

Object Storage error messages
Error message Troubleshooting topic
pvc:...:can't access bucket <bucket_name>: NotFound: Not Found Why can't my PVC access an existing bucket?
Error: symlink ... helm-ibmc: file exists Why does installing the Object storage Helm plug-in fail?
d--------- 1 root root 0 Jan 1 1970 <file_name> (non-root user cannot access files) Resolving non-root user access issues to files in IBM Cloud
EPERM: operation not permitted Why does my app pod fail with an Operation not permitted error?
chown: changing ownership of '<volume_mount_path>': Input/output error Why can't the ownership of the mount path be changed?
Error: rendered manifest contains a resource that already exists. ... existing_kind: storageClass Why does installing the IBM Cloud Object Storage plug-in fail?
Bad value for ibm.io/object-store-endpoint ... scheme is missing. Why do I see wrong s3fs or IAM API endpoints when I create a PVC?
SignatureDoesNotMatch: The request signature we calculated does not match the signature you provided. Why do I see wrong credentials or access denied messages when I create a PVC?
Object Storage PVC remains in pending state Why does my PVC remain in a pending state?
can't get credentials: can't get secret tsecret-key: secrets "secret-key" not found Why does PVC or pod creation fail due to not finding the Kubernetes secret?
Transport endpoint is not connected. Why is the transport endpoint not connected?
Error mounting volume: s3fs mount failed: s3fs: error while loading shared libraries: libfuse.so.2 Why do I see a volume mounting error when using the IBM Cloud Object Storage plug-in?
Transport endpoint not connected errors when using the COS cluster add-on Why do I see transport endpoint not connected errors when using the IBM Cloud Object Storage cluster add-on?

Portworx Storage

Portworx Storage error messages
Error message Troubleshooting topic
kp.Error: ... msg='Unauthorized: The user does not have access to the specified resource' Why does encryption fail with an invalid KMS endpoint?