# Process group stuck in ResourcesTerminating state

**URL:** <https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232>\
**Category:** Kubernetes Operator\
**Created:** [March 28, 2022, 1:28pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232 "2022-03-28T13:28:10Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![larshagen](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/larshagen/32/776_2.png) [@larshagen](https://forums.foundationdb.org/u/larshagen)\
**Post date:** [March 28, 2022, 1:28pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/1 "2022-03-28T13:28:10Z")

</div>

This happens when running the v1.0.0 operator. A pod has been deleted from the cluster, the process group looks like this:

```auto
{
                        "addresses": [
                            "10.83.79.152"
                        ],
                        "exclusionTimestamp": "2022-03-28T11:29:34Z",
                        "processClass": "log",
                        "processGroupConditions": [
                            {
                                "timestamp": 1648466973,
                                "type": "ResourcesTerminating"
                            }
                        ],
                        "processGroupID": "log-1",
                        "removalTimestamp": "2022-03-28T11:20:55Z"
                    },

```

The timestamp of the process group condition is the same as the exclusion timestamp. Neither the PVC nor the service is deleted, and the cluster has been stuck in this state for a few hours.

The operator keeps logging `Waiting for volume claim to get torn down`, so it seems it is trying to confirm that the process group is deleted, but it does not actually remove the PVC and service, so it doesn’t move forward.

---

<div class="post-metadata">

**Author:** ![johscheuer](https://avatars.discourse-cdn.com/v4/letter/j/f475e1/32.png) [@johscheuer](https://forums.foundationdb.org/u/johscheuer)\
**Post date:** [March 28, 2022, 3:04pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/2 "2022-03-28T15:04:24Z")

</div>

What is the state of the PVC and the underlying PV? The operator will delete all related resources (if they exist) and then wait/check until they have a deletionTimestamp [fdb-kubernetes-operator/remove\_process\_groups.go at main · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/main/controllers/remove_process_groups.go#L337-L363).

Just to clarify the Pod was deleted but the PVC and the Service is not deleted? If the label config ([fdb-kubernetes-operator/cluster\_spec.md at main · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/main/docs/cluster_spec.md#labelconfig)) wouldn’t match I would suspect that the operator is not able to delete any resource.

Could you validate if the service and the PVC have those labels: [fdb-kubernetes-operator/foundationdb\_labels.go at main · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/main/api/v1beta2/foundationdb_labels.go#L48-L55)?

---

<div class="post-metadata">

**Author:** ![larshagen](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/larshagen/32/776_2.png) [@larshagen](https://forums.foundationdb.org/u/larshagen)\
**Post date:** [March 28, 2022, 4:13pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/3 "2022-03-28T16:13:14Z")

</div>

Yes, pod has been deleted, but not PVC nor service.  
PVC labels:

```auto
labels:
    foundationdb.org/fdb-cluster-name: foundationdb-cluster
    foundationdb.org/fdb-process-class: log
    foundationdb.org/fdb-process-group-id: log-1

```

Service labels

```auto
labels:
    foundationdb.org/fdb-cluster-name: foundationdb-cluster
    foundationdb.org/fdb-process-class: log
    foundationdb.org/fdb-process-group-id: log-1

```

---

<div class="post-metadata">

**Author:** ![larshagen](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/larshagen/32/776_2.png) [@larshagen](https://forums.foundationdb.org/u/larshagen)\
**Post date:** [March 28, 2022, 6:53pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/4 "2022-03-28T18:53:26Z")

</div>

If a process group is markedForRemoval, and the pod is deleted externally, then the pod will be given the status terminating, right? [fdb-kubernetes-operator/update\_status.go at 5023f43b5ede3b72183a50033d261090379d1d5b · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/5023f43b5ede3b72183a50033d261090379d1d5b/controllers/update_status.go#L413)

In this case we will remove all other conditions [fdb-kubernetes-operator/update\_status.go at 5023f43b5ede3b72183a50033d261090379d1d5b · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/5023f43b5ede3b72183a50033d261090379d1d5b/controllers/update_status.go#L594)

When we later do zoned removal, the process group is placed in the terminating zone, which is skipped during zone deletion [fdb-kubernetes-operator/remove.go at 5023f43b5ede3b72183a50033d261090379d1d5b · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/5023f43b5ede3b72183a50033d261090379d1d5b/internal/removals/remove.go#L66)

Am I on the right track here?

---

<div class="post-metadata">

**Author:** ![johscheuer](https://avatars.discourse-cdn.com/v4/letter/j/f475e1/32.png) [@johscheuer](https://forums.foundationdb.org/u/johscheuer)\
**Post date:** [March 28, 2022, 6:57pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/5 "2022-03-28T18:57:57Z")

</div>

Okay, that’s an interesting bug. You have to delete the PVC and the service manually. The issue is that the process group is in the `ResourcesTerminating` state and those process groups are skipped in the removal step and they will only be validated if all resources are deleted. The simplest solution is to change [fdb-kubernetes-operator/remove\_process\_groups.go at main · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/main/controllers/remove_process_groups.go#L337) to `append(processGroupsToRemove, terminatingProcessGroups...)` in order to not issue multiple deletion we can check in the `removeProcessGroup` method if the resource still exists and has no `deletionTimestamp` to trigger the deletion.

---

<div class="post-metadata">

**Author:** ![johscheuer](https://avatars.discourse-cdn.com/v4/letter/j/f475e1/32.png) [@johscheuer](https://forums.foundationdb.org/u/johscheuer)\
**Post date:** [March 28, 2022, 6:58pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/6 "2022-03-28T18:58:57Z")

</div>

You’re on the right track. Feel free to provide a PR with the fix otherwise I try to fix this issue tomorrow.

Thanks for reporting this issue!

---

<div class="post-metadata">

**Author:** ![larshagen](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/larshagen/32/776_2.png) [@larshagen](https://forums.foundationdb.org/u/larshagen)\
**Post date:** [March 28, 2022, 7:34pm UTC](https://forums.foundationdb.org/t/process-group-stuck-in-resourcesterminating-state/3232/7 "2022-03-28T19:34:52Z")

</div>

BTW, is the way we set terminating state in [fdb-kubernetes-operator/update\_status.go at 5023f43b5ede3b72183a50033d261090379d1d5b · FoundationDB/fdb-kubernetes-operator · GitHub](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/5023f43b5ede3b72183a50033d261090379d1d5b/controllers/update_status.go#L391) something we should change as part of issue #970? It sets the state to terminating if the process group is marked for removal and the pod is missing or terminating, but we should probably count it as MissingPod if the process is not fully excluded in FDB.
