Skip to main content
Version: 3.7

NodeDrain CRD Reference

NodeDrain (storagenode.io/v1, kind NodeDrain) is a namespaced custom resource that the Portworx Operator creates and reconciles to track the drain lifecycle of a single Kubernetes node during migration of its datastore from PX-StoreV1 to PX-StoreV2. The operator creates one NodeDrain custom resource per node (or per batch of storageless nodes) being drained and updates its status as the node progresses through cordoning, data evacuation, and node cleanup. For the full workflow, see Migrate Portworx Datastore from PX-StoreV1 to PX-StoreV2.

To inspect a NodeDrain custom resource:

kubectl get nodedrain -n <portworx>
kubectl describe nodedrain <node-drain-name> -n <portworx>

NodeDrain​

FieldDescriptionType
apiVersionAPIVersion defines the versioned schema of this representation of an object.
Servers should convert recognized schemas to the latest internal value and
may reject unrecognized values.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#resources
string
kindKind is a string value representing the REST resource that this object represents.
Servers may infer this value from the endpoint to which the client submits requests.
Cannot be updated.
In CamelCase.
More info: https://git.k8s.io/community/contributors/devel/sig-architecture/api-conventions.md#types-kinds
string
specNodeDrainSpec defines the desired state of the drain operation.object
statusObserved state of the drain operation.object

spec fields​

FieldDescriptionType
spec.typeIdentifies the drain workflow this custom resource drives. For the PX-StoreV1-to-V2 migration workflow, this is always PXStoreMigration.string
spec.sourceNodesPortworx node IDs of the source nodes being drained.array
spec.targetNodesPortworx node IDs of the target nodes to which the pool drain should evacuate data. May be empty, in which case the pool drain selects target nodes itself.array
spec.isStoragelessNodeSpecifies whether the source node is a storageless node rather than a storage node.boolean
spec.restartSet by the StorageCluster controller (or an operator) to request that the NodeDrain controller replay unfinished steps from the first condition that hasn't succeeded. The controller acts only when spec.restart.attempt is greater than status.lastRestart.attempt.object
spec.restart.attemptA monotonically increasing counter that identifies this restart request.integer
spec.restart.reasonExplanation of why the restart was requested.string
spec.restart.timestampWhen the restart request was written.string

status fields​

FieldDescriptionType
status.phaseThe current phase of the drain operation. See status.phase values.string
status.jobIDThe ID of the underlying Portworx job for the current phase, for example, the pool drain job. Empty when the current phase has no associated job.string
status.nextStateThe phase to which the controller plans to transition next. Empty for terminal phases (Completed, Cancelled, Failed).string
status.repl3VolumeIdsIDs of volumes that were running with replication factor 3 and were temporarily reduced to replication factor 2 during the drain so that they can be restored to replication factor 3 afterward. Discovered once, early in the drain; volumes created after discovery aren't added to this list.array
status.conditionsTracks the observed state of each major workflow stage. Each entry corresponds to one stage, for example, VolumeAttachmentsDrained or PoolDrainComplete, and is updated as the controller progresses.array
status.conditions.typeThe workflow stage that this condition reports. See status.conditions types.string
status.conditions.reasonShort CamelCase token for the condition's current status: InProgress, Succeeded, Failed, or Paused.string
status.conditions.messageDetails about the current state or failure.string
status.conditions.lastTransitionTimeWhen this condition last changed status.string
status.retriesPer-step retry counts, keyed by condition type and persisted so that retry counts survive operator restarts.object
status.stepTimersThe time when each long-running step first started, keyed by condition type and persisted so that step timeouts survive operator restarts.object
status.lastErrorThe most recent failure that moved the custom resource into the Failed phase. Cleared when a spec.restart request is accepted.object
status.lastError.stepThe condition type of the stage that failed.string
status.lastError.reasonShort CamelCase token that describes why the step failed.string
status.lastError.messageComplete error message.string
status.lastError.retryableSpecifies whether the failure is transient and can be retried automatically or resolved with an operator restart.boolean
status.lastError.retriesThe retry count at the time of the final failure.integer
status.lastError.lastObservedWhen this error was last recorded.string
status.lastPhaseChangeAtWhen status.phase was last updated.string
status.lastRestartThe most recently accepted spec.restart request.object
status.lastRestart.attemptMirrors the accepted spec.restart.attempt.integer
status.lastRestart.reasonMirrors the accepted spec.restart.reason.string
status.lastRestart.acceptedAtWhen the controller recorded acceptance of the restart request.string
status.preflightResults of the per-node preflight checks that run before cordoning.object
status.preflight.verdictOverall preflight result: Passed or Failed.string
status.preflight.checksIndividual check results.array
status.preflight.checks.nameThe check's identifier, for example, KVDBQuorumHealthy or SufficientFreeSpaceForDrain.string
status.preflight.checks.passedSpecifies whether this check succeeded.boolean
status.preflight.checks.messageDetails returned by the check, especially on failure.string
status.preflight.checks.lastRunWhen this check last ran.string
status.preflight.lastEvaluatedWhen preflight was last run.string

status.phase values​

ValueDescription
NotStartedInitial state. Preflight checks run first; if they pass, the node is cordoned and volume attachments are evacuated. Advances to ReducingVolumeReplicas (three-storage-node clusters) or Pending once the pool drain job is submitted.
ReducingVolumeReplicasThree-storage-node clusters only. Replication-factor-3 volumes on the source node are being reduced to replication factor 2 before the pool drain.
PendingThe pool drain job has been submitted and is queued.
RunningThe pool drain job is actively evacuating data.
PausedThe underlying pool drain job is paused.
CancelledThe pool drain job was canceled. The migration is marked Failed; restarting resets this custom resource to NotStarted.
DrainedThe pool drain job finished; all data has been evacuated from the source node's pools.
RestoringVolumeReplicasThree-storage-node clusters only. Volumes temporarily reduced to replication factor 2 are being restored to replication factor 3.
NodeWipeRunningThe node-wiper job is running on the source node.
PxRestartingPortworx is restarting on the source node with PX-StoreV2.
DeletingOfflineNodesThe source node is being removed from Portworx cluster membership.
CompletedThe drain operation finished successfully.
FailedThe drain operation failed. See status.lastError for details.
UnknownThe phase couldn't be determined.

status.conditions types​

TypeDescription
PreflightPassedPer-node preflight checks passed before the drain workflow begins.
K8sNodeCordonedThe source Kubernetes node was marked unschedulable.
VolumeAttachmentsDrainedAll volume attachments were moved off the source node.
VolumeRepl3ReducedThree-storage-node clusters only. Replication-factor-3 volumes with a replica on the source node were temporarily reduced to replication factor 2.
PoolDrainCompleteThe pool drain job finished; all data was evacuated from the source node.
PXStoppedPortworx was stopped on the source node.
NodesRemovedFromClusterThe source node was removed from Portworx cluster membership.
NodeWipeCompleteThe storage wipe on the source node finished.
CloudDriveConfigMapClearedCloud storage only. The source node's entry was removed from the Portworx cloud-drive ConfigMap.
StorageProvisionedNew storage was provisioned on the source node after restart.
PXRestartedPortworx restarted and is healthy on the source node.
KVDBHealthyAll KVDB members are healthy after the node rejoined the cluster.
VolumeRepl3RestoredThree-storage-node clusters only. Replication factor 3 was restored on volumes after the drain phase completed.
NodeUncordonedThe source node was made schedulable again after migration.
DrainCompletedThe full drain and migration workflow completed successfully for this node.

The reason field for each condition has one of the following values: InProgress, Succeeded, Failed, or Paused.