Skip to main content

Tutorial 05: latency tuning

Mirrored from the FEMU repository

This page is hw/femu/docs/tutorials/05-latency-tuning.md at FEMU 39a55eeb6 (2026-10-02), licensed GPL-2.0-or-later. Send corrections to the FEMU repository.

You change the parts of FEMU's timing model one at a time: the flat NAND times, the cell type, the channel bus, the host link and the controller firmware cost. For each one you predict the effect, measure it with fio, and compare. It takes about twenty minutes.

You need: tutorial 01, the variables from Before you start, and about 5 GiB of free host memory. Read performance tuning first if you want numbers that repeat to the microsecond: a poller or FTL thread that has to share a core adds delay the model did not ask for.

Background​

FEMU copies the data of a command at once, computes when the command would have finished on the emulated device, and posts the completion no earlier than that (the timing model). A queue depth 1 latency is therefore the model's time plus a fixed overhead of the guest and of FEMU itself. Each step below changes one term of the model; NAND media and timing has the formulas.

1. The device and the measurement​

Every run uses a small BlackBox SSD, 512 MiB of NAND in 2 channels of 4 LUNs, and adds one option to it:

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25

Start QEMU with the command line of tutorial 01, step 2 and this -device line, restarting QEMU for each configuration. In the guest, fill the first 256 MiB (a read of a page never written costs no NAND time), then run four measurements:

sudo fio --name=fill --filename=/dev/nvme0n1 --direct=1 --ioengine=libaio \
--rw=write --bs=128k --iodepth=16 --size=256M
F="--filename=/dev/nvme0n1 --direct=1 --ioengine=libaio --size=256M --runtime=8 --time_based"
sudo fio --name=rd4k $F --rw=randread --bs=4k --iodepth=1
sudo fio --name=rd128k $F --rw=randread --bs=128k --iodepth=1
sudo fio --name=qd32 $F --rw=randread --bs=4k --iodepth=32
sudo fio --name=wr4k $F --rw=randwrite --bs=4k --iodepth=1

Read the median (clat 50.00th percentile) of the QD1 tests and the IOPS of the QD32 test. Keep the write test last. Random overwrites move pages to new places, so a 128 KiB read after them finds its 32 pages spread unevenly over the LUNs and takes longer than the model of a freshly filled range (in a first run with the write test second, it took 289 us instead of 165 us). On this small drive, background garbage collection may also start in the last second of the write test; that moves its tail, not its median.

2. The baseline​

With the default timing (40 us read, 200 us program, no channel bus):

TestResult
4 KiB read, QD144.3 us
4 KiB write, QD1203.8 us
128 KiB read, QD1164.9 us
4 KiB read, QD32176,880 IOPS

A 4 KiB read is pg_rd_lat plus about 4 us. A 128 KiB read is 32 pages. The FTL spreads consecutive pages over the channels and then the LUNs, so the 32 pages sit on all 8 LUNs, 4 each. The LUNs work in parallel and each reads its 4 pages one after another: 4 x 40 us = 160 us. Copying the 128 KiB takes about 30 us (measured with NAND times set to 0), but FEMU copies at once and holds only the completion, so the copy overlaps the modelled time instead of adding to it. At queue depth 32, 8 LUNs reading in parallel allow at most 8 / 40 us = 200,000 reads per second; FEMU delivers 88% of that.

3. Flat NAND times​

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25,pg_rd_lat=80000,pg_wr_lat=400000

Prediction: reads and writes take 40 us and 200 us longer, and QD32 throughput halves.

TestBaseline80 us / 400 us
4 KiB read, QD144.3 us84.5 us
4 KiB write, QD1203.8 us423.9 us
4 KiB read, QD32176,880 IOPS89,317 IOPS

4. Cell type​

nand_cell_type replaces the flat times with built-in tables that give each page type in a wordline its own time (cell types). TLC (3) has lower, center and upper pages read in 56.5, 77.5 and 106 us:

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25,nand_cell_type=3
TestBaselineTLC
4 KiB read, QD144.3 us81.4 us
4 KiB read, QD1, 99th percentile52.0 us114.2 us
4 KiB write, QD1203.8 us2310.1 us
4 KiB read, QD32176,880 IOPS89,449 IOPS

The pages of a block are split about equally over the lower, center and upper types, so the read median is the center page's 77.5 us plus overhead, and the 99th percentile is near the upper page. TLC programs take milliseconds. pg_rd_lat, pg_wr_lat and blk_er_lat have no effect while nand_cell_type is set.

5. The channel bus​

By default, moving a page between controller and NAND takes no time. Any non-zero bus phase turns the channel bus on; pg_xfer_lat is the data transfer per page, and transfers on one channel run one at a time:

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25,pg_xfer_lat=20000
TestBaseline20 us transfer
4 KiB read, QD144.3 us63.7 us
128 KiB read, QD1164.9 us383.0 us
4 KiB read, QD32176,880 IOPS98,056 IOPS

A single page read gains exactly one transfer. A 128 KiB read moves 16 pages over each channel, one at a time: 320 us of transfers, which now dominate the read (with NAND times set to 0 it still takes 342 us). Under load the two channels become the limit: each moves one page per 20 us, 100,000 pages per second for both, which is where QD32 lands.

pcie_bandwidth_mbps charges each Read and Write its size divided by the link bandwidth, on one queue per direction (host link):

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25,pcie_bandwidth_mbps=1000
TestBaseline1000 MB/s link
4 KiB read, QD144.3 us47.9 us
128 KiB read, QD1164.9 us305.2 us

128 KiB at 1000 MB/s is 131 us, and the 128 KiB read grew by 140 us. A 4 KiB read pays 4 us. pcie_prop_delay_ns adds a fixed delay on top.

7. Controller firmware​

fw_cpu_ns charges a fixed cost per Read, Write and Zone Append on one modelled controller core, so commands queue behind each other there:

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25,fw_cpu_ns=20000
TestBaseline20 us firmware
4 KiB read, QD144.3 us63.7 us
4 KiB write, QD1203.8 us230.4 us
4 KiB read, QD32176,880 IOPS49,534 IOPS

At queue depth 1 the cost simply adds. Under load one core finishes one command per 20 us, 50,000 per second, and that becomes the limit no matter how many LUNs the NAND has.

What you learned​

  • Queue depth 1 latency is the modelled time plus a few microseconds; each timing option moves it by what you configured.
  • Throughput is set by whichever resource saturates first: LUNs, channels with the bus on, or the firmware core.
  • nand_cell_type gives a spread of read times rather than one value.

The numbers above come from one run per configuration on a 20-core host where other work was running, so expect yours to differ by a few microseconds. The relations between them should hold.

Next​