# Upgrading fdb without downtime

**URL:** <https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208>\
**Category:** Running FoundationDB\
**Created:** [November 2, 2023, 6:03pm UTC](https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208 "2023-11-02T18:03:17Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![pdeva](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/pdeva/32/1655_2.png) [@pdeva](https://forums.foundationdb.org/u/pdeva)\
**Post date:** [November 2, 2023, 6:03pm UTC](https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208/1 "2023-11-02T18:03:17Z")

</div>

we are trying to deploy fdb on ec2 hosts.

the [documentation on upgrading](https://github.com/apple/foundationdb/wiki/Upgrading-FoundationDB) says:

> you need to upgrade all of the processes at once, because the old and new processes will be unable to communicate with each other.

however, that is impossible to accomplish across multiple hosts at once. even the documentation only talks about upgrade process for a single host. Does that mean that upgrading a production fdb instance, which will definitely have multiple hosts, it is not possible to avoid downtime?

---

<div class="post-metadata">

**Author:** ![rajivr](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/rajivr/32/1100_2.png) [@rajivr](https://forums.foundationdb.org/u/rajivr)\
**Post date:** [November 3, 2023, 1:59am UTC](https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208/2 "2023-11-03T01:59:44Z")

</div>

Here is a [gist](https://gist.github.com/rajivr/f8fb0369c5caba39178e8c60427c3478) from my notes. Hope it helps.

I would recommend that you please consider using [`fdb-kubernetes-operator`](https://github.com/FoundationDB/fdb-kubernetes-operator/blob/main/docs/manual/upgrades.md) to manage the lifecycle of the cluster.

---

<div class="post-metadata">

**Author:** ![pdeva](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/pdeva/32/1655_2.png) [@pdeva](https://forums.foundationdb.org/u/pdeva)\
**Post date:** [November 3, 2023, 5:57pm UTC](https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208/3 "2023-11-03T17:57:10Z")

</div>

we are not using k8s so the k8s operator cannot be used.

regarding your gist:

i am not sure i understand it fully. does it suggest that all processes should be killed via `fdbcli --exec kill` and then restored? wont that result in downtime?

---

<div class="post-metadata">

**Author:** ![rajivr](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/rajivr/32/1100_2.png) [@rajivr](https://forums.foundationdb.org/u/rajivr)\
**Post date:** [November 4, 2023, 1:43am UTC](https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208/4 "2023-11-04T01:43:56Z")

</div>

> [@pdeva](#):
>
> am not sure i understand it fully. does it suggest that all processes should be killed via `fdbcli --exec kill` and then restored? wont that result in downtime?

Maybe FDB core developers can provide a better answer.

But from what I understand, there will be a small impact to client latency (FDB uses a fat client), and inflight transactions will be retried. It is still your responsibility to ensure that your transactions are _idempotent_.

At a high level from what I understand, FDB has one recovery path (when individual components fail) and one upgrade path that is extensively tested using simulation testing.

---

<div class="post-metadata">

**Author:** ![markus.pilman](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/markus.pilman/32/379_2.png) [@markus.pilman](https://forums.foundationdb.org/u/markus.pilman)\
**Post date:** [November 4, 2023, 1:41pm UTC](https://forums.foundationdb.org/t/upgrading-fdb-without-downtime/4208/5 "2023-11-04T13:41:45Z")

</div>

> [@pdeva](#):
>
> i am not sure i understand it fully. does it suggest that all processes should be killed via `fdbcli --exec kill` and then restored?

Yes, this is correct.

> [@pdeva](#):
>
> wont that result in downtime?

Yes, but if done well, downtime should be done short (1-5 seconds).

Not supporting rolling upgrades has been a deliberate design choice. Basically, it reduces the testing surface drastically (there’s a few other benefits).

If you write your own operations, the burden of making sure this will be fast is on you. This is roughly how we do it:

1. Make sure the new client version is installed on all clients (in _addition_ to the old one, you want the application to be able to talk to both versions of FDB so it can fail over – FDBs multi-version client will take care of this for you).
2. Install the new version of FDB on all machines (we use the versioned RPM, but you could also just copy the `fdbserver` binaries over)
3. Either change the `foundationdb.conf` file to point to the new binary or change the symlink to the binary on all machines (depending on how it is installed). If you change the `foundationdb.conf` you also need to set an option that `fdbmonitor` doesn’t automatically trigger restarts when the file changes.
4. Run `kill; kill all` on `fdbcli`. The `fdbserver` processes on all machines will kill themselves and `fdbmonitor` will immediately restart them.

If you do the above correctly, the downtime should be very short. Obviously, the processes won’t restart all at exactly the same time (as you pointed out, this is impossible to achieve). But that is fine: as soon as a majority of the processes run the new version, the old processes won’t be able to rejoin the cluster until they also get bounced.

You might want to include some automation here to also check the cluster health (check whether `kill` in `fdbcli` returns all processes that you expect to see, check whether the cluster reports healthy etc).
