# FoundationDB Summit 2019: Managing FoundationDB at Scale

**URL:** <https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763>\
**Category:** Community\
**Created:** [November 18, 2019, 10:29pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763 "2019-11-18T22:29:54Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![alexmiller](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/alexmiller/32/326_2.png) [@alexmiller](https://forums.foundationdb.org/u/alexmiller)\
**Post date:** [November 18, 2019, 10:29pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/1 "2019-11-18T22:29:54Z")

</div>

Speaker: @john_brownlee  
Slides: [Managing FoundationDB](https://static.sched.com/hosted_files/foundationdbsummit2019/c4/FDB%20Summit%202019%20Presentation.key)  
Recording: [https://www.youtube.com/watch?v=A3U8M8pt3Ks](https://www.youtube.com/watch?v=A3U8M8pt3Ks)

---

<div class="post-metadata">

**Author:** ![alexmiller](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/alexmiller/32/326_2.png) [@alexmiller](https://forums.foundationdb.org/u/alexmiller)\
**Post date:** [November 18, 2019, 11:00pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/2 "2019-11-18T23:00:03Z")

</div>

The Kubernetes Operator is now public at [FoundationDB/fdb-kubernetes-operator](https://github.com/FoundationDB/fdb-kubernetes-operator).

---

<div class="post-metadata">

**Author:** ![KrzysFR](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/krzysfr/32/43_2.png) [@KrzysFR](https://forums.foundationdb.org/u/KrzysFR)\
**Post date:** [November 22, 2019, 2:18pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/3 "2019-11-22T14:18:58Z")

</div>

Is there a PDF version available of the slides? I’m not sure how to open a `.key` file.

---

<div class="post-metadata">

**Author:** ![janl](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/janl/32/493_2.png) [@janl](https://forums.foundationdb.org/u/janl)\
**Post date:** [November 22, 2019, 2:38pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/4 "2019-11-22T14:38:33Z")

</div>

PDF version here: [http://jan.prima.de/u/FDB\_Summit\_2019\_Presentation.pdf](http://jan.prima.de/u/FDB_Summit_2019_Presentation.pdf)

---

<div class="post-metadata">

**Author:** ![baiwfg2](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/baiwfg2/32/1173_2.png) [@baiwfg2](https://forums.foundationdb.org/u/baiwfg2)\
**Post date:** [October 19, 2021, 8:54am UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/5 "2021-10-19T08:54:13Z")

</div>

@john_brownlee Hi, John. I’m not a native speaker and don’t know what `bounce` means.

For the following slide, my understanding is: As long as one of proxies or logServers dies, we restart all fdb processes (including proxies, LogServers, StorageServer, fdbmonitor, coordinators ? who won’t be restarted ?). That’s called `bounce everything at once`. Right ?

 ![image](https://global.discourse-cdn.com/foundationdb/original/2X/9/98b7b7a74ce67468c2628fba3d6d0f6232cdd007.jpeg)

---

<div class="post-metadata">

**Author:** ![john\_brownlee](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/john_brownlee/32/22_2.png) [@john\_brownlee](https://forums.foundationdb.org/u/john_brownlee)\
**Post date:** [October 19, 2021, 3:12pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/6 "2021-10-19T15:12:48Z")

</div>

Yes, that’s correct. “Bounce” in this context means “restart”, and every process in the cluster gets restarted.

---

<div class="post-metadata">

**Author:** ![baiwfg2](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/baiwfg2/32/1173_2.png) [@baiwfg2](https://forums.foundationdb.org/u/baiwfg2)\
**Post date:** [October 20, 2021, 3:26am UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/7 "2021-10-20T03:26:22Z")

</div>

You mean all processes ? Failure of a single process of any role will trigger all processes of all role bounce ? Is the strategy overly performed ? Sounds controversial indeed.

Then how long does this strategy cause unavailability ? And the upper Layer (client) can tolerate that (maybe by retrying) ?

---

<div class="post-metadata">

**Author:** ![john\_brownlee](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/john_brownlee/32/22_2.png) [@john\_brownlee](https://forums.foundationdb.org/u/john_brownlee)\
**Post date:** [October 20, 2021, 2:35pm UTC](https://forums.foundationdb.org/t/foundationdb-summit-2019-managing-foundationdb-at-scale/1763/8 "2021-10-20T14:35:55Z")

</div>

This slide is describing deliberate restarts, such as when changing the command-line parameters for the processes, rather than failures. Failures don’t trigger restarts of the other processes, but the failure of a process in the transaction subsystem will cause a recovery, where the database recruits a new transaction subsystem. These recoveries take a single-digit number of seconds, which the clients will experience as operations blocking until the database is available again. For the case of process restarts, the unavailability period will be slightly longer, but when running through fdbmonitor that additional period can be kept small.
