Storage cache sizing calculator
You have plenty of slow storage and want a fast tier in front of it: SSDs in front of spinning disks, say, or in front of object storage. Tell us how much data you move each day and the average read speed you want. You'll get how big the cache has to be and how long it has to hold on to data.
This is the number that matters most. It sits between the second tier's speed and the first's: the closer you get to the first, the bigger the cache has to be.
[[ t.error_title ]]
[[ errorText ]]
[[ t.out_size ]]
[[ gb(result.size) ]]
[[ tb(result.size) ]] · [[ devicesText ]]
[[ fmt(t.out_split, { w: gb(result.fromWrites), r: gb(result.fromReads) }) ]]
[[ t.no_cache_title ]]
[[ fmt(t.no_cache_text, { speed: speed(values.speed2) }) ]]
- [[ t['stat_' + s.k] ]]
- [[ s.v ]]
- [[ t['stat_' + s.k + '_help'] ]]
[[ hoverPoint.speed ]]
[[ t.tooltip_size ]]: [[ hoverPoint.size ]]
[[ t.tooltip_hit ]]: [[ hoverPoint.hit ]]
[[ fmt(t.chart_note, { speed: speed(values.speed1) }) ]]
[[ t.table_toggle ]]
[[ t.table_help ]]
| [[ t.table_hit ]] | [[ t.table_speed ]] | [[ t.table_days ]] | [[ t.table_size ]] |
|---|---|---|---|
| [[ speed(row.target) ]] | [[ smart(row.days) ]] | [[ gb(row.size) ]] |
How it works
The calculator looks at two storage tiers: a small, fast one (the cache) and a large, slow one where the data actually lives. A read that finds its data in the cache runs at first-tier speed, any other read at second-tier speed. So the average speed comes down to the share of reads served from the cache: the hit ratio.
The hit ratio you need
The right way to average two speeds is the harmonic mean, because it's the read times that add up, not the speeds. With v1 the first tier's speed, v2 the second's and T your target, the hit ratio p you need is:
1 / T = p / v1 + (1 − p) / v2
p = v1 · (T − v2) / (T · (v1 − v2))
How long data stays cached
The model assumes data is used heavily right after it's written and less and less afterwards, with exponential decay: after one half-life h, the chance of reading it again has halved. If the cache keeps each piece of data for t days, the share of reads that come later, and so go to the second tier, is 2^(−t/h). To keep that at or below 1 − p:
t = h · log₂(1 / (1 − p))
How big
Over those t days the cache has to hold the data written (W per day) plus the data read (R per day). How much the reads weigh depends on what happens when you go back to older data. If it's a spot access, like resending an old invoice to a customer, you read it and that's it: it counts once. If it's rework instead, say you dig out old files to rerun the numbers or train a model, the old data becomes current again and drags the data around it along, which from then on gets read as if it were new. Each read pulls in another one with probability p, so the chain averages 1 + p + p² + … = 1 / (1 − p) reads. With a high p this is by far the biggest share.
S = W · t + R · t / (1 − p) redo the whole job
S = W · t + R · t spot access
Miss latency is the second tier's latency plus the time to copy a whole chunk into the cache. Drives needed is the cache size over the capacity per drive, rounded up, with no redundancy. Drive writes per day (DWPD) count everything that enters the cache: in steady state, with time-based eviction, the whole cache is replaced every t days.
Tips
- Use the time in cache as the eviction timeout. If the cache drops data sooner, the hit ratio falls and the average speed with it.
- Don't aim for the first tier's full speed. Look at the chart: the last few MB/s cost more than all the others put together.
- On top of the drives computed here, add the ones for redundancy (RAID, replicas) and keep some free space: many SSDs slow down when they're nearly full.
- Here 1 GB = 1000 MB and 1 TB = 1000 GB, the same units drive makers use for capacity.