# We have some machines and cpu and a little ssd

**URL:** <https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532>\
**Category:** Using FoundationDB\
**Created:** [July 22, 2019, 1:53pm UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532 "2019-07-22T13:53:58Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Vasilii](https://avatars.discourse-cdn.com/v4/letter/v/258eb7/32.png) [@Vasilii](https://forums.foundationdb.org/u/Vasilii)\
**Post date:** [July 22, 2019, 1:53pm UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532/1 "2019-07-22T13:53:58Z")

</div>

We have 10 machines with 40cpu and 6 ssd 2t per machine.

We read documentation and found this

[https://apple.github.io/foundationdb/configuration.html#guidelines-for-setting-process-class](https://apple.github.io/foundationdb/configuration.html#guidelines-for-setting-process-class)

> `class` : Process class specifying the roles that will be taken in the cluster. Recommended options are `storage` , `transaction` , `stateless`

but here we found

> [@Why doesn't my cluster performance scale when I double the number of machines?](https://forums.foundationdb.org/t/why-doesnt-my-cluster-performance-scale-when-i-double-the-number-of-machines/640/10):
>
> Okay Alex, thank you! This is what I have in my cluster of 10 machines at the moment: ip port cpu% mem% iops net class roles --------------- ------ ---- ---- ---- --- ----------- -------------------- 172.31.32.74 4500 1 6 6 0 log log 4501 0 4 - 0 stateless 4502 0 3 - 0 stateless 4503 5 3 - 2 stateless cluster\_controlle…

and class settings is proxy, log and we found class tests

where we can found full description for this opion `class` ???

and how calculate amount of process of each class type for machine?

---

<div class="post-metadata">

**Author:** ![markus.pilman](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/markus.pilman/32/379_2.png) [@markus.pilman](https://forums.foundationdb.org/u/markus.pilman)\
**Post date:** [July 22, 2019, 11:13pm UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532/2 "2019-07-22T23:13:02Z")

</div>

Documentation is definitely a thing that we need to improve on…

As of FDB 6.2 (which will be released soon), this will be the possible classes (from the source code):

```auto
	enum ClassType { UnsetClass,
                                      StorageClass,
                                      TransactionClass,
                                      ResolutionClass,
                                      TesterClass,
                                      ProxyClass,
                                      MasterClass,
                                      StatelessClass,
                                      LogClass,
                                      ClusterControllerClass,
                                      LogRouterClass,
                                      DataDistributorClass,
                                      CoordinatorClass,
                                      RatekeeperClass, 
                                      InvalidClass = -1 };

```

As you can see, there are many more than mentioned in the documentation (although to be fair, some of those don’t exist in FDB 6.1). However, for most workloads you probably won’t need to explicitly set these classes.

The strategy I would recommend for you:

1. Start with something simple. I think a good strategy would be to run per machine:

- two TLogs (these can share one disk)
- 10 storage processes (two per disk)
- 1-2 Stateless processes

1. This probably will run reasonably well.
2. If you run into any issues, you should reiterate and maybe refine the configuration.

You probably should make sure that you have 8GB memory per process. I would expect that a configuration like this will be good enough for a long time before you’ll run into any issues. Afterwards you will need to identify bottlenecks and refine accordingly.

---

<div class="post-metadata">

**Author:** ![alexmiller](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/alexmiller/32/326_2.png) [@alexmiller](https://forums.foundationdb.org/u/alexmiller)\
**Post date:** [July 23, 2019, 12:49am UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532/3 "2019-07-23T00:49:42Z")

</div>

> [@markus.pilman](#):
>
> However, for most workloads you probably won’t need to explicitly set these classes.

The warning from fdbcli that he’s getting is from not explicitly setting these classes. By default, there’s nothing that stops FDB from recruiting storage servers on TLogs, which is a sufficiently frequent source of poor performance that a warning was finally added for it.

> [@markus.pilman](#):
>
> - two TLogs (these can share one disk)

I don’t think I’ve ever benchmarked more than one TLog per disk. I would expect that this would increase latency due splitting the required work of fsync()ing data into two calls?

---

<div class="post-metadata">

**Author:** ![Vasilii](https://avatars.discourse-cdn.com/v4/letter/v/258eb7/32.png) [@Vasilii](https://forums.foundationdb.org/u/Vasilii)\
**Post date:** [July 23, 2019, 7:29am UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532/4 "2019-07-23T07:29:18Z")

</div>

Ok  
for stateless class wich ssd i must use?  
Tlog ssd or one state less process for 1 disk  
6 ssd = 1 ssd tlog + 5\*2 storage and each 6 ssd stateless???

or

6 ssd = 1 ssd tlog + 5\*2 storage and 2 stateless on tlog ssd ???

and we have 400g ram per machine

---

<div class="post-metadata">

**Author:** ![alexmiller](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/alexmiller/32/326_2.png) [@alexmiller](https://forums.foundationdb.org/u/alexmiller)\
**Post date:** [July 23, 2019, 8:00am UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532/5 "2019-07-23T08:00:00Z")

</div>

Stateless processes won’t actually use the disk, so you can put them wherever without impacting much. Putting them all on one SSD is fine, spreading it out is fine. They only store a 4KB file in the data directory, and emit a normal amount of logs.

6 stateless per host seems a bit much, so 2\*stateless per TLog sounds fine.

---

<div class="post-metadata">

**Author:** ![markus.pilman](https://sea1.discourse-cdn.com/foundationdb/user_avatar/forums.foundationdb.org/markus.pilman/32/379_2.png) [@markus.pilman](https://forums.foundationdb.org/u/markus.pilman)\
**Post date:** [July 23, 2019, 6:05pm UTC](https://forums.foundationdb.org/t/we-have-some-machines-and-cpu-and-a-little-ssd/1532/6 "2019-07-23T18:05:38Z")

</div>

> [@Vasilii](#):
>
> and we have 400g ram per machine

Sadly FDB is not great when it comes to utilizing CPU cores (if you have many more processes, your disks might be over-utilized). This makes sense to a large degree because FDB is mostly IO-bound.

However, you can make better use of your memory. To be on the safe side, I would recommend to give each process at least 8GB of memory (which is the default). On your machines this will give you ~100 GB of unused memory (104 - but you probably want some memory reserved for the OS and operational tools you’re going to run on these machines). We were very successful by using our memory for caching. So you could give each storage process an additional 10GB of memory for caching which will, if the workload is skewed to some degree a great performance improvement.

To do so you can change the `--memory` option for all storage servers. To use the memory for the cache you will also need to set two knobs. Sadly you will need to set these options for each storage process (which might not be too bad if you auto-generate your foundationdb.conf file). But if your storages run on ports 4503-4513 (as an example) your conf might look a bit like this:

```auto
[fdbserver.4503]
memory = 18GiB
knob_page_cache_4k = 12884901888 # 12 GiB

```

(the way this math works is: 2GiB page cache is default, we added 10GiB memory, so we can add these to the 2GiB we already have. 12\*2^30 = 12884901888 if I did the math right).

For our workloads with a similar config we have a cache-hit rate of ~99.99% which makes a big difference for us…

If your cache hit rate is similar, you might even be able to double the number of storage nodes per disk (making the average cache size 5GiB which should give you the same cache hit rate as the amount of data per process will be lower). This way you could improve on CPU utilization and write throughput (we found that fdb is kind of limited when it comes to utilizing disks for writes - a problem that redwood is hopefully going to solve).
