# Cluster stuck with status "Locked coordination state" even when all coordination servers available

**URL:** <https://forums.foundationdb.org/t/cluster-stuck-with-status-locked-coordination-state-even-when-all-coordination-servers-available/2537>\
**Category:** Using FoundationDB\
**Created:** [February 1, 2021, 9:27am UTC](https://forums.foundationdb.org/t/cluster-stuck-with-status-locked-coordination-state-even-when-all-coordination-servers-available/2537 "2021-02-01T09:27:01Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![ashishgupta](https://avatars.discourse-cdn.com/v4/letter/a/c67d28/32.png) [@ashishgupta](https://forums.foundationdb.org/u/ashishgupta)\
**Post date:** [February 1, 2021, 9:27am UTC](https://forums.foundationdb.org/t/cluster-stuck-with-status-locked-coordination-state-even-when-all-coordination-servers-available/2537/1 "2021-02-01T09:27:01Z")

</div>

We run a cluster population automation that does the following -

1. Starts a new, triple-replicated cluster on 5 machines (1 coordinator per machine).
2. Transacts a custom metadata key into the database.
3. Sets the replication factor to single.
4. Populates the cluster using a number of clients transacting objects simultaneously in a loop.
5. Once populated, we move it back to triple replication.

We’ve been facing issues where clients timeout in a read operation right at the beginning of step 4. The read operation is on the same key that was transacted in step 2. We dump the cluster status at such a time, and get the following -

```
fdbcli --exec "status details"

Locking coordination state. Verify that a majority of coordination server
processes are active.

10.134.188.8:4271 (reachable)
10.134.188.9:4271 (reachable)
10.134.188.109:4271 (reachable)
10.134.188.116:4271 (reachable)
10.134.188.119:4271 (reachable)

```

Any insights into why the cluster might be stuck? This happens only sometimes, so hard to repro the exact scenario.

---

<div class="post-metadata">

**Author:** ![andrew.noyes](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/andrew.noyes/32/443_2.png) [@andrew.noyes](https://forums.foundationdb.org/u/andrew.noyes)\
**Post date:** [February 1, 2021, 5:28pm UTC](https://forums.foundationdb.org/t/cluster-stuck-with-status-locked-coordination-state-even-when-all-coordination-servers-available/2537/2 "2021-02-01T17:28:58Z")

</div>

Can you look in the trace logs and see if there are any “severity 40” events?

---

<div class="post-metadata">

**Author:** ![osamarin](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/osamarin/32/905_2.png) [@osamarin](https://forums.foundationdb.org/u/osamarin)\
**Post date:** [February 4, 2021, 12:29pm UTC](https://forums.foundationdb.org/t/cluster-stuck-with-status-locked-coordination-state-even-when-all-coordination-servers-available/2537/3 "2021-02-04T12:29:32Z")

</div>

`Locking coordination state` message means that FDB recovery is in progress. See

> <https://github.com/apple/foundationdb/blob/master/design/recovery-internals.md>

for details.

Unfortunally, there is no any troubleshooting guide on the recovery stuck.

Some similar issue that causes the same result is [The database gets unavailable after changing usable\_regions · Issue #3925 · apple/foundationdb · GitHub](https://github.com/apple/foundationdb/issues/3925)
