Docs

Troubleshooting

Work from the environment outward: first the install command, then the pod or task, then the connection to Alien, then the operation.

The install command stops

The generated commands and the chart check the target before they change anything. When a check fails, the command prints one of these messages and stops.

Kubernetes

MessageCauseWhat to do
The existing setup-owned credentials Secret differs from operator-credentials.yaml. Refusing overwrite.A Secret with the same name exists and has other values.Use the operator-credentials.yaml that belongs to this installation. Do not overwrite the existing Secret.
The setup-owned credentials Secret <namespace>/<name> is missing. Create it before installing or upgrading this release.Helm ran without the Secret.Run the generated install or enable command, which creates the Secret first.
The setup-owned credentials Secret encryption key does not match remoteOperator.existingSecret.encryptionKeySha256.The Secret and operator-values.yaml come from different sets of values.Use both files from the same Generate test values result.
This exact Remote Operator installation completed previously. Register a new installation instead.The dedicated installation was installed and uninstalled before.Register a new installation.
The shared access-request CRD differs from this reviewed chart.The cluster has another version of the shared CRD.Stop. A cluster administrator must review compatibility for every installed operator before a separate CRD migration. Helm never changes the CRD.
The identity volume is missing. Restore it; an upgrade must not bootstrap a replacement identity even when every other managed resource was deleted.The identity PersistentVolumeClaim was deleted.Restore the volume. An upgrade never creates a new identity.
Remote Operator cannot be enabled on the initial Helm install.A product chart was installed with the operator enabled.Install the product release with the operator disabled first, then run the enable command.
remoteOperator.bootstrapIdentity has already been consumed by this release. Set it to false; ...An upgrade set bootstrapIdentity=true again.Upgrade with remoteOperator.bootstrapIdentity=false.
Disabling Remote Operator on an existing release deletes its workload and permissions. ...A product release upgrade set remoteOperator.enabled=false.To remove the operator, confirm it with remoteOperator.confirmRemoval. See Remove only the operator.
Recover the failed or pending Helm revision first using the guarded recovery command. (dedicated) or Recover the failed or pending product release first using the guarded recovery command. (product)The release is not deployed.Use the guarded recovery.
Use Helm 3 or 4 for this reviewed chart.Another Helm major version is installed.Install Helm 3.13 or later, or Helm 4.

Amazon ECS

MessageCauseWhat to do
AWS account mismatch: expected <id>, got <id>Your AWS credentials belong to another account.Switch to credentials for the account you entered in setup.
Choose at most one subnet per Availability Zone, or supply an existing EFS filesystem and access pointEFS allows one mount target per Availability Zone.Remove the extra subnets, or use existing EFS storage.
This installation's EFS mount targets use subnets ...; keep the same subnets in the same orderAn update changed the subnets of a stack that created its own EFS storage.Enter the original subnets in the original order.
Existing registration secret uses customer-managed KMS key ...The secret <stack>-registration uses a customer-managed KMS key.Use the default aws/secretsmanager key.
Registration token is requiredThe bootstrap command's token prompt got an empty answer.Paste the unexpired registration token from setup. Every bootstrap attempt before the operator registers prompts, including retries after a failed first deployment. An update or rollback of a registered installation asks for no token. See Apply an update.
curl is required to verify the registration tokenThe bootstrap command must check a new token with Alien, and curl is not installed.Install curl and run the command again.
Could not confirm this setup token belongs to an unregistered stack; no secret was changedAlien rejected the token check. The token is expired, revoked, or mistyped; the account, Region, cluster, or stack differs from the one recorded at setup; or an operator has already registered from this stack.Before registration, generate a replacement registration token in setup and run the command again. After registration, select Regenerate reviewed template and run the update command, which asks for no token. See Apply an update.
Unexpected bootstrap check response <code>; no secret was changedThe token check returned a status other than 204, such as a proxy redirect.Make sure the machine reaches Alien's API directly, then run the command again.
Registration secret <stack>-registration could not be readAn update or rollback command could not read the secret the installation registered with.Check that your AWS credentials can read the secret, or restore it. The command never creates a replacement secret for a registered installation.

The pod or task does not start

On Kubernetes:

  1. Run Inspect this release from the setup page.
  2. Check the operator Deployment, ServiceAccount, Secret, and image.
  3. Read the pod events before you change permissions.

On ECS, read the CloudWatch Logs group in the stack output LogGroupName, then use Repair task. See Repair a connection.

The operator runs but Alien shows no connection

Alien counts an installation as connected only while it receives heartbeats less than five minutes old. Check outbound DNS and HTTPS to the management endpoint in the template.

First find out whether the Secret is missing, the key is rejected, or the network path is blocked. Do not regenerate values for a registered installation. See Repair a connection.

An operation fails

Trace its dependencies in order:

operation selected in Alien
          ↓
present in the generated operator image
          ↓
allowed by Kubernetes RBAC or the cloud identity
          ↓
target reachable over the network
          ↓
target accepts the credential and request

Fix the first step that fails. Do not grant broad cluster access as a shortcut.

If the setup page shows Permissions need an update, the installed permissions no longer match the enabled operations. See Upgrade and roll back.

A Kubernetes Helm upgrade fails

Do not retry right away, and do not delete Helm release Secrets, the identity volume, or the shared CRD to clear a lock. Check helm history, then follow Recover a failed or stopped Helm operation.

An Amazon ECS update fails

Check the CloudFormation stack status and events for the failing resource. If an update or rollback is still running, wait for it to finish before issuing another command. After CloudFormation restores the previous stack state, correct the failure and rerun Apply reviewed update or Rollback; neither asks for a token. If CloudFormation cannot finish its rollback, restore the stack to a stable state before retrying. If the stack updated but its task does not start, see Repair a connection.

On this page