# FoundationDB 7.1.24 - the memory usage after clean startup of fdbserver process is too high

**URL:** <https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863>\
**Category:** Using FoundationDB\
**Created:** [March 20, 2023, 7:37am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863 "2023-03-20T07:37:15Z")\
**Posts on this page:** 11\
**Page:** 1

<div class="post-metadata">

**Author:** ![deng](https://avatars.discourse-cdn.com/v4/letter/d/9fc29f/32.png) [@deng](https://forums.foundationdb.org/u/deng)\
**Post date:** [March 20, 2023, 7:37am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/1 "2023-03-20T07:37:15Z")

</div>

We have a cluster of 7 nodes, with a total of 56 fdbserver processes, storing over 50 billion entries. After a clean reboot of an fdbserver process, the memory usage exceeds 17GB. If we set the memory limit in the `foundationdb.conf` file to below 17GB, the fdbserver processes cannot successfully boot.

 ![image](https://global.discourse-cdn.com/foundationdb/original/2X/9/989d483c8bba0ea56b2f525f81df6a2b15676aa4.jpeg)

We are unsure why so much memory is needed, and if we were to deploy the cluster to store 100 billion data entries, how much memory would be required?

Any help would be much appreciated.

---

<div class="post-metadata">

**Author:** ![losfair](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/losfair/32/1149_2.png) [@losfair](https://forums.foundationdb.org/u/losfair)\
**Post date:** [March 21, 2023, 4:31pm UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/2 "2023-03-21T16:31:23Z")

</div>

What storage engine are you using? If it’s `memory` it will take up a lot memory.

---

<div class="post-metadata">

**Author:** ![deng](https://avatars.discourse-cdn.com/v4/letter/d/9fc29f/32.png) [@deng](https://forums.foundationdb.org/u/deng)\
**Post date:** [March 22, 2023, 2:23am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/3 "2023-03-22T02:23:29Z")

</div>

Thanks for your reply. The storage engine is `ssd-2`, `fdbcli --exec "status;"` gives us:

```auto
$ fdbcli
Using cluster file `fdb.cluster'.

The database is available.

Welcome to the fdbcli. For help, type `help'.
fdb> status

Using cluster file `fdb.cluster'.

Configuration:
  Redundancy mode - triple
  Storage engine - ssd-2
  Coordinators - 7
  Usable Regions - 1

Cluster:
  FoundationDB processes - 56
  Zones - 7
  Machines - 7
  Memory availability - 32.0 GB per process on machine with least available
  Retransmissions rate - 1 Hz
  Fault Tolerance - 2 machines
  Server time - 03/22/23 10:17:44

Data:
  Replication health - Healthy
  Moving data - 0.000 GB
  Sum of key-value sizes - 49.688 TB
  Disk space used - 197.873 TB

Operating space:
  Storage server - 3321.1 GB free on most full server
  Log server - 3321.9 GB free on most full server

Workload:
  Read rate - 345451 Hz
  Write rate - 0 Hz
  Transactions started - 337321 Hz
  Transactions committed - 0 Hz
  Conflict rate - 0 Hz

Backup and DR:
  Running backups - 0
  Running DRs - 0

Client time: 03/22/23 10:17:44

```

---

<div class="post-metadata">

**Author:** ![harikb](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/harikb/32/1033_2.png) [@harikb](https://forums.foundationdb.org/u/harikb)\
**Post date:** [March 22, 2023, 9:49pm UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/4 "2023-03-22T21:49:33Z")

</div>

Not sure if this helps, but …

You have 7 machines, holding 56 processes. I am assuming this is 8 per machine. What is the configuration for [fdbserver] section?  
[https://apple.github.io/foundationdb/configuration.html#fdbserver-section](https://apple.github.io/foundationdb/configuration.html#fdbserver-section)  
Particularly `memory=`. If it is defaulting to 8 GB, as far as I know, you need 64 GB max per machine.

In addition, your earlier message sounded as is the memory usage goes back up immediately after reboot (even if there is no activity on it). Your status output indicates a fairly good amount of traffic. Is it really going that high as soon as db is up or after getting enough traffic?

Clarifying these might help someone answer the question.

---

<div class="post-metadata">

**Author:** ![SteavedHams](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/steavedhams/32/18_2.png) [@SteavedHams](https://forums.foundationdb.org/u/SteavedHams)\
**Post date:** [March 23, 2023, 6:23am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/5 "2023-03-23T06:23:21Z")

</div>

I’m not familiar with the tool you are using but “Memory” of 17gb and “Memory Used” of 52% would seem to indicate a 17GB virtual memory footprint and about half of that is actually Resident. Is that correct?

Virtual Memory will often be a factor of 2x or 3x larger than Resident memory, but Resident is what matters. In FDB 7.1, the malloc implementation was changed from glibc to jemalloc, which has a much larger virtual : resident ratio.

There are two options to `fdbserver` which limit memory, `--memory` sets a resident limit and `--memory-vsize` sets the virtual memory limit.

---

<div class="post-metadata">

**Author:** ![deng](https://avatars.discourse-cdn.com/v4/letter/d/9fc29f/32.png) [@deng](https://forums.foundationdb.org/u/deng)\
**Post date:** [March 24, 2023, 3:21am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/6 "2023-03-24T03:21:30Z")

</div>

Thanks for your questions!

> [@harikb](#):
>
> You have 7 machines, holding 56 processes. I am assuming this is 8 per machine.

Yes, we have 8 fdbserver processes on each machine, and each process manages 1 NVMe disk.

> [@harikb](#):
>
> What is the configuration for [fdbserver] section?

The [fdbserver] section is:

```auto
## Default parameters for individual fdbserver processes
[fdbserver]
command = /usr/sbin/fdbserver
public-address = auto:$ID
listen-address = public
datadir = /foundationdb/data/$ID
logdir = /var/log/foundationdb
logsize = 200MiB
maxlogssize = 8MiB
# machine-id =
# datacenter-id =
# class =
memory = 32GiB
# storage-memory = 1GiB
cache-memory = 8GiB
# metrics-cluster =
# metrics-prefix =

```

> [@harikb](#):
>
> Particularly `memory=` . If it is defaulting to 8 GB, as far as I know, you need 64 GB max per machine.

Each machine in our cluster is equipped with 512GB memory.

> [@harikb](#):
>
> In addition, your earlier message sounded as is the memory usage goes back up immediately after reboot (even if there is no activity on it). Your status output indicates a fairly good amount of traffic. Is it really going that high as soon as db is up or after getting enough traffic?

The memory usage increases significantly as soon as the database (db) is up, and during the booting time, the CPU usage of each fdbserver process reaches approximately 100%. Once the cluster is back to a healthy state, there is no traffic from outside clients or internal traffic between fdbserver processes.

---

<div class="post-metadata">

**Author:** ![deng](https://avatars.discourse-cdn.com/v4/letter/d/9fc29f/32.png) [@deng](https://forums.foundationdb.org/u/deng)\
**Post date:** [March 24, 2023, 4:02am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/7 "2023-03-24T04:02:34Z")

</div>

Thanks for your reply.

> [@SteavedHams](#):
>
> I’m not familiar with the tool you are using but “Memory” of 17gb and “Memory Used” of 52% would seem to indicate a 17GB virtual memory footprint and about half of that is actually Resident. Is that correct?

The tool reports resident memory usage, and `/proc` also gives that the resident memory usage of fdbserver is above 17GB. In one node, execute the following command, we get resident and virtual memory usage in KiB bytes unit (the first column is resident memory and the second column is virtual memory):

```auto
$ pidof fdbserver | sed -e 's: : -p :g' | xargs ps -o rss= -o vsz= -p
16868952 17029776
16898596 17085836
16866468 17006220
16779736 17007628
16819000 17029008
16829896 16999052
16802448 16991244
16746752 16969484

```

and, the current fdb details status is:

```auto
>>> status details
Using cluster file `/etc/foundationdb/fdb.cluster'.

Configuration:
  Redundancy mode - triple
  Storage engine - ssd-2
  Coordinators - 7
  Usable Regions - 1

Cluster:
  FoundationDB processes - 56
  Zones - 7
  Machines - 7
  Memory availability - 32.0 GB per process on machine with least available
  Retransmissions rate - 1 Hz
  Fault Tolerance - 2 machines
  Server time - 03/24/23 11:55:00

Data:
  Replication health - Healthy
  Moving data - 0.000 GB
  Sum of key-value sizes - 49.702 TB
  Disk space used - 196.752 TB

Operating space:
  Storage server - 3352.1 GB free on most full server
  Log server - 3343.0 GB free on most full server

Workload:
  Read rate - 29 Hz
  Write rate - 1 Hz
  Transactions started - 5 Hz
  Transactions committed - 0 Hz
  Conflict rate - 0 Hz

Backup and DR:
  Running backups - 0
  Running DRs - 0

Process performance details:
  10.142.7.74:4500 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.74:4501 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.74:4502 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.74:4503 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.6 GB / 32.0 GB RAM )
  10.142.7.74:4504 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.74:4505 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.74:4506 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.74:4507 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.7 GB / 32.0 GB RAM )
  10.142.7.85:4500 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.5 GB / 32.0 GB RAM )
  10.142.7.85:4501 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.6 GB / 32.0 GB RAM )
  10.142.7.85:4502 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.5 GB / 32.0 GB RAM )
  10.142.7.85:4503 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.5 GB / 32.0 GB RAM )
  10.142.7.85:4504 ( 1% cpu; 0% machine; 0.002 Gbps; 2% disk IO;17.5 GB / 32.0 GB RAM )
  10.142.7.85:4505 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.5 GB / 32.0 GB RAM )
  10.142.7.85:4506 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.5 GB / 32.0 GB RAM )
  10.142.7.85:4507 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;17.4 GB / 32.0 GB RAM )
  10.142.7.86:4500 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.2 GB / 32.0 GB RAM )
  10.142.7.86:4501 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.2 GB / 32.0 GB RAM )
  10.142.7.86:4502 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.3 GB / 32.0 GB RAM )
  10.142.7.86:4503 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.86:4504 ( 1% cpu; 0% machine; 0.003 Gbps; 2% disk IO;16.2 GB / 32.0 GB RAM )
  10.142.7.86:4505 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.86:4506 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.86:4507 ( 1% cpu; 0% machine; 0.003 Gbps; 1% disk IO;16.2 GB / 32.0 GB RAM )
  10.142.7.88:4500 ( 1% cpu; 0% machine; 0.004 Gbps; 2% disk IO;16.2 GB / 32.0 GB RAM )
  10.142.7.88:4501 ( 1% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.88:4502 ( 1% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.88:4503 ( 1% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.7 GB / 32.0 GB RAM )
  10.142.7.88:4504 ( 1% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.88:4505 ( 1% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.2 GB / 32.0 GB RAM )
  10.142.7.88:4506 ( 1% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.88:4507 ( 3% cpu; 0% machine; 0.004 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4500 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4501 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4502 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4503 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4504 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4505 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4506 ( 1% cpu; 0% machine; 0.002 Gbps; 2% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.89:4507 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4500 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4501 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4502 ( 1% cpu; 0% machine; 0.002 Gbps; 2% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4503 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4504 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4505 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4506 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.90:4507 ( 1% cpu; 0% machine; 0.002 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4500 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4501 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4502 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4503 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4504 ( 0% cpu; 0% machine; 0.001 Gbps; 0% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4505 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4506 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )
  10.142.7.91:4507 ( 1% cpu; 0% machine; 0.001 Gbps; 1% disk IO;16.1 GB / 32.0 GB RAM )

Coordination servers:
  10.142.7.74:4500 (reachable)
  10.142.7.85:4500 (reachable)
  10.142.7.86:4500 (reachable)
  10.142.7.88:4500 (reachable)
  10.142.7.89:4500 (reachable)
  10.142.7.90:4500 (reachable)
  10.142.7.91:4500 (reachable)

Client time: 03/24/23 11:55:00

```

> [@SteavedHams](#):
>
> There are two options to `fdbserver` which limit memory, `--memory` sets a resident limit and `--memory-vsize` sets the virtual memory limit.

We used `--memory` option of `foundationdb.conf` to limit resident memory.

---

<div class="post-metadata">

**Author:** ![SteavedHams](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/steavedhams/32/18_2.png) [@SteavedHams](https://forums.foundationdb.org/u/SteavedHams)\
**Post date:** [March 24, 2023, 6:50am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/8 "2023-03-24T06:50:17Z")

</div>

Thanks for confirming that the 17 GB is RSS.

I see the problem. You have 49 TB of KV data and 56 processes. After triple replication, this is 49 \* 3 / 56 = 2.6 TB of KV data that each process is responsible for.

There is a data structure that storage servers have called the Byte Sample which stores a deterministic random sample of keys. This data is persisted on disk in the storage engine and is loaded immediately upon storage server startup. Unfortunately, its size is not tracked or reported, but grows linearly with KV size and I suspect yours is somewhere around 4GB-6GB based on the memory usage I’ve seen for smaller storage KV sizes.

The byte sample size is technically configurable to a smaller sample rate, however changing its knob once a cluster is created is undefined behavior. FDB relies on the byte sample’s determinism to know how much logical data is in each shard and each storage server. **Weird things will happen if you change this knob on a existing cluster and the cluster may become unavailable**. I don’t think any data loss would occur but you could easily get into a situation that is hard to get out of.

If you need to reduce memory usage for you disk sizes you could reduce the cache memory setting.

If you want to reduce the size of the byte sample, you would have to create a new cluster and migrate your data to it. You would also have to make sure that **no storage servers on the new cluster ever start up without the knob override**. The option is `knob_byte_sampling_factor` and the default is 250. Multiply this by N to reduce the byte sample size by a factor of N.

---

<div class="post-metadata">

**Author:** ![deng](https://avatars.discourse-cdn.com/v4/letter/d/9fc29f/32.png) [@deng](https://forums.foundationdb.org/u/deng)\
**Post date:** [March 30, 2023, 2:40am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/9 "2023-03-30T02:40:11Z")

</div>

Thank you for providing that valuable information. We’ll try it in our future benchmarking.

And, we are curious whether the data structure `Byte Sample` is a global data structure that each `fdbserver` process keeps one replica of, or if it is a local data structure that only contains random sample keys from its managed key-value pairs.

---

<div class="post-metadata">

**Author:** ![jzhou](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/jzhou/32/445_2.png) [@jzhou](https://forums.foundationdb.org/u/jzhou)\
**Post date:** [March 30, 2023, 3:16am UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/10 "2023-03-30T03:16:52Z")

</div>

`Byte Sample` is a local data structure that each storage server maintains, used for estimating shard sizes on the storage server.

---

<div class="post-metadata">

**Author:** ![PierreZ](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/pierrez/32/866_2.png) [@PierreZ](https://forums.foundationdb.org/u/PierreZ)\
**Post date:** [April 22, 2024, 5:37pm UTC](https://forums.foundationdb.org/t/foundationdb-7-1-24-the-memory-usage-after-clean-startup-of-fdbserver-process-is-too-high/3863/11 "2024-04-22T17:37:41Z")

</div>

I took the liberty of [blogging](https://pierrezemb.fr/posts/redwood-memory-tuning/) about some Redwood memory tuning, including Byte Sample and page-cache 😉
