Skip to main content

Parameter manual

Mirrored from the FEMU repository

This page is hw/femu/docs/reference/parameter-manual.md at FEMU 39a55eeb6 (2026-10-02), licensed GPL-2.0-or-later. Send corrections to the FEMU repository.

This manual explains how to configure an emulated device: the ways a parameter reaches FEMU, then every group of parameters by component, with what each one means, its unit, the values it accepts, and how it interacts with the others. It ends with worked configurations.

The property reference is generated from the binary and is the authority on names, types and defaults. This manual never states a default; each group links to its table there. Run-time properties and counters are in runtime-properties.md.

Contents​

  1. How parameters reach FEMU
  2. Devices and how they combine
  3. Controller
  4. Namespaces and subsystems
  5. NAND geometry
  6. NAND timing
  7. FTL: mapping, caches and write buffer
  8. Garbage collection
  9. Reliability, wear and fault insertion
  10. Host link and controller firmware
  11. ZNS
  12. FDP
  13. OCSSD, CSD and KV
  14. CXL SSD
  15. Accepted for compatibility, no effect
  16. Worked configurations

1. How parameters reach FEMU​

1.1 Device options on the QEMU command line​

Every FEMU parameter is a property of one of three QEMU devices, given on the command line when QEMU starts:

-device femu-subsys,id=SUB,prop=value,... NVMe subsystem (FDP, shared namespaces)
-device femu-cxl-ssd,id=CXL,prop=value,... CXL Type-3 SSD
-device femu,id=CTRL,prop=value,... NVMe controller and its namespaces

The rules of QEMU's option syntax apply:

  • Options are separated by commas. A comma inside a value is written twice: namespace_modes=bbssd,,znssd is the list bbssd,znssd.
  • Booleans take on or off (QEMU also accepts yes/no and true/false).
  • Properties of type size take a number of bytes or a QEMU size with a suffix (64M, 4G). Integer properties take plain numbers, decimal or 0x hexadecimal; their unit is in the property name or in this manual (_mb MiB, _ns and _lat nanoseconds, _pcent percent).
  • A property that links to another device (subsys, cxl_ssd) names the other device's id, and that device must come earlier on the command line.
  • A misspelled name stops QEMU with Property 'femu.NAME' not found.

Static properties are read once, when QEMU creates (realizes) the device. Realize checks the values and their combinations and stops QEMU with a message naming the rule that failed; the "Limits and refusals" section of each mode page lists them. ./qemu-system-x86_64 -device femu,help prints every property with its type, default and description; femu-subsys,help and femu-cxl-ssd,help do the same for the other two devices.

1.2 Configuration files​

hw/femu/scripts/ssd-config.sh expands a file of key = value lines into the same -device options, so a configuration can be read, commented and kept in version control. Keys are property names, mode = bbssd stands for femu_mode, and a [subsys] section becomes a femu-subsys device that the controller joins. The script checks names against -device femu,help (and femu-subsys,help for [subsys] keys) of the binary in FEMU_BIN, or of one it finds in the build directories; without a binary it only warns. It also rejects an unknown mode and malformed lines. QEMU still checks the values. Tutorial 09 walks through it, and scripts and tools describes the shipped files.

1.3 Run-time properties over QMP​

A few properties can be read or changed while the guest runs, with QMP qom-get and qom-set on /machine/peripheral/<id>:

  • femu: simulate-power-loss, a write-only trigger for the power-cut model (power loss).
  • femu-cxl-ssd: the cache tunables cache-ways, prefetch-degree and prefetch-stride (also accepted on -device), the actions flush-cache, stats-reset, der-ratio and control-command, and every counter.

The full list is runtime-properties.md. femu and femu-subsys have no counter properties; their counters are NVMe log pages (log pages and counters).

1.4 Run-time switches from the guest​

These are not parameters, but they change the model while the guest runs:

  • Vendor admin command 0xEF on a BlackBox controller turns GC time and the flat NAND times off and on, and counts completions posted after their due time. Code 3 sets the built-in flat times, not the ones on your command line; code 4 sets them to 0 (timing model).
  • Set Features 06h (Volatile Write Cache) turns the write buffer off and on when vwc=1; Set Features 20h (Key Value Configuration) sets EDNEK on a KV namespace; Set Features 04h and 0Bh set the temperature threshold and the asynchronous events the host wants.

1.5 Environment variables​

A few debugging and host-placement aids are environment variables of the QEMU process, not properties: environment variables.

2. Devices and how they combine​

femu-subsys (optional) femu-cxl-ssd (optional)
FDP, shared namespaces CXL memory over NAND
^ subsys=SUB ^ cxl_ssd=CXL
| |
femu (NVMe controller) ---------------------+
controller parameters section 3
femu_mode, namespaces section 4
per-mode parameters:
bbssd, CSD geometry 5, timing 6, FTL 7, GC 8, reliability 9
KV geometry 5, timing 6, part of 8
ZNS zns_* 11
OCSSD, CSD section 13
host link, firmware section 10

A femu controller runs one mode per namespace. femu_mode picks it for every namespace, and namespace_modes per namespace. Only the parameters of the modes in use have an effect; the description of each property in properties.md names the modes it applies to. A property of a mode that no namespace runs has no effect. The properties of section 15 have no effect in any mode and warn when set.

3. Controller​

These shape the NVMe controller the guest driver sees, whatever the mode.

3.1 Queues and pollers​

Reference: queues, pollers and interrupts.

ParameterUnitValid valuesMeaning
queuescount1 to 2047I/O submission and completion queue pairs the controller offers; MSI-X vectors are queues + 1
entriescount, 0's based1 to 65534CAP.MQES; queues may have up to entries + 1 entries
multipoller_enabled0 or 10, 10: one poller thread for every I/O queue; 1: several, see poller_ratio
poller_ratioqueues per poller0 or more (0 counts as 1)with multipoller_enabled=1, ceil(queues / poller_ratio) pollers, each serving a round-robin share of the queues
stridepower of two0 to 12doorbell stride CAP.DSTRD: doorbells are 4 << stride bytes apart
max_sqes, max_cqespower of two6 and 4 onlyIdentify SQES and CQES; the only accepted values are the 64-byte and 16-byte entries
aerlcount, 0's based0 to 255outstanding Asynchronous Event Requests the controller holds
elpecount, 0's based0 to 255Error Information log entries kept
hiops_inlineboolon, offNoSSD only: performance option, on by default; set off only when debugging
intc, intc_thresh, intc_timeas namedintc 0 or 1initial values of the Interrupt Coalescing and Interrupt Vector Configuration features; reported only, interrupts are not coalesced

Interactions:

  • Every poller spins on a host core while the controller is enabled. The number of pollers follows queues, not the number of queues the guest creates, so set queues to the guest's vCPU count when you turn on more pollers (performance tuning).

3.2 Identity, capabilities and transfers​

Reference: controller identity and capabilities.

ParameterUnitValid valuesMeaning
vid, didPCI IDs16-bitPCI vendor and device ID; vid is also Identify VID
mdtspower of two0 to 255; 0 = no limitlargest transfer, 2^(12 + mpsmin + mdts) bytes
mpsmin, mpsmaxpower of twompsmin <= mpsmax <= 15host memory page sizes the controller accepts, 2^(12 + n) bytes
cqr0 or 10, 11 requires physically contiguous queues
vwc0 or 10, 1advertise a volatile write cache and make feature 06h usable; Flush drains the write buffer either way
oacsbit maskonly bit 1 (0x2)Format NVM support; clearing it refuses Format NVM
oncsbit mask0x1 Compare, 0x2 Write Uncorrectable, 0x4 Dataset Management, 0x8 Write Zeroes, 0x10 Save/Select, 0x80 Verify, 0x100 Copyoptional NVM commands; Timestamp is always added
sglboolon, offaccept scatter gather lists; OCSSD ignores it
cmbsz, cmblocregisterscmbsz size x unit a power of two; cmbloc BAR field 2Controller Memory Buffer; 0 means none
temperaturekelvin16-bitcomposite temperature in the SMART log, compared with the threshold feature
pel_filehost patha writable pathkeeps the Persistent Event log and power cycle count across QEMU runs (retention)

Interactions:

  • Write Zeroes, Compare, Copy and Verify are unreachable until you set their oncs bit. A test of those commands on a default controller tests nothing. Flush is always accepted and drains the bbssd write buffer even with vwc=0, though Linux sends no Flush to a controller that advertises no cache.
  • With ZNS, mdts also caps Zone Append when zns_zasl_bs=0.

3.3 LBA formats, metadata and protection​

Reference: LBA formats, metadata and protection.

ParameterUnitValid valuesMeaning
nlbafcount1 to 16, at most 8 with metaLBA formats offered: 512 bytes, doubling with each
lba_indexindexbelow nlbafformat the namespaces boot with; ZNS needs 4 KiB or less
metabytes per blockNoSSD and bbssd onlymetadata per logical block; each format is then also offered with metadata
mcbit maskbit 0 interleaved, bit 1 separateMetadata Capabilities; required with meta
extended0 or 1needs mc bit 0boot with metadata interleaved (extended LBAs)
pibooltakes effect only with meta >= 8offer protection information types 1 to 3 to Format and Create
dpc, dpsbit masksdpc 0 with meta; leave dps 0Identify fields; use pi instead

Interactions: meta is refused with FDP, with dpc or dps, and with any namespace that is not NoSSD or bbssd; pi is refused with power_loss and cxl_ssd (namespace management, metadata and PI).

4. Namespaces and subsystems​

Reference: mode, capacity and namespaces, namespace management, streams and power loss, shared namespaces. Design: subsystems, controllers and namespaces.

ParameterUnitValid valuesMeaning
femu_modemode number0 OCSSD, 1 bbssd, 2 NoSSD, 3 ZNS, 4 CSD, 5 KVmode of every namespace unless namespace_modes is set
devsz_mbMiB32-bithost memory backend, split across the namespaces; ignored with op_pcent
namespacescount1 to 256; OCSSD and FDP 1namespaces present from boot
namespace_sizesbytes, QEMU sizesone entry per namespace, each >= 512 bytes, sum <= backendsize of each namespace; unset splits the backend evenly
namespace_modeslistnossd, bbssd, znssd, ocssd, csd, kvssd, one per namespacemode of each namespace
op_pcentpercentbbssd; not with cxl_ssdback the device with the full NAND and expose NAND / (1 + op_pcent/100)
ns_mgmt (femu)boolstandalone NoSSD or bbssd controllerNamespace Management and Attachment from the guest
bbssd_ns_limitcount1 to 256, >= namespacesmost bbssd namespaces ns_mgmt may allocate, each with its own FTL
streams, streams.maxbool, countstreams.max 1 to 32the Streams directive; bbssd separates streams per FTL page
power_lossboolsee interactionsroll back writes still in the write buffer on a simulated power cut
subsysdevice ida femu-subsys listed earlierjoin a subsystem
ns_mgmt (femu-subsys)boolNoSSD and bbssd controllers, not with fdpone namespace table and backend shared by every controller of the subsystem
nqn (femu-subsys)stringanysubsystem NQN suffix

Interactions:

  • Namespaces are packed one after another in the backend. Each bbssd or CSD namespace has its own FTL built from the whole NAND geometry, so each must fit that geometry on its own with room for GC (the reserve). Each ZNS namespace builds its zones from its own size.
  • op_pcent overrides devsz_mb: the backend becomes the raw NAND, and each namespace gets its share of NAND x 100 / (100 + op_pcent).
  • A controller may hold at most one CSD namespace; OCSSD takes the whole controller.
  • ns_mgmt on femu stays off without an error unless every namespace runs femu_mode with dps 0; with a shared subsystem use the subsystem's ns_mgmt.
  • streams needs bbssd with mapping page or dftl (or NoSSD, where it has no placement effect). It is refused with FDP, with shared namespaces (ns_mgmt on the subsystem), and when a second controller joins the subsystem. bbssd reserves streams.max + 1 lines for it.
  • power_loss needs bbssd (femu_mode=1), buffer_size > 0 and page-aligned namespaces. Without vwc=1 it is accepted but the write buffer is off, so there is nothing to roll back. It is refused with meta, pi, ns_mgmt, subsys, namespace_modes and cxl_ssd.

5. NAND geometry​

Applies to bbssd, CSD and KV. Reference: NAND geometry. Design: geometry.

ParameterUnitValid valuesMeaning
secszbytes> 0sector size
secs_per_pgsectors1 to 256sectors per NAND page; page size = secs_per_pg x secsz
pgs_per_blkpages1 to 65536; <= 512 with nand_cell_typepages per block
blks_per_plblocks1 to 65536blocks per plane; also the number of lines
pls_per_lunplanes1 to 16planes per LUN
luns_per_chLUNs1 to 128LUNs (dies) per channel
nchschannels1 to 4096channels

Derived quantities:

page = secs_per_pg x secsz bytes
NAND capacity = nchs x luns_per_ch x pls_per_lun x blks_per_pl x pgs_per_blk x page
lines = blks_per_pl (a line is one block on every plane of every LUN)
line size = nchs x luns_per_ch x pls_per_lun x pgs_per_blk x page
parallel units = nchs x luns_per_ch (LUNs work in parallel; planes of a LUN together)

Interactions:

  • The total sector count must fit in a signed 32-bit integer.
  • Without op_pcent, the namespace is devsz_mb and must leave GC its reserve of lines; the reserve grows with gc_thres_pcent_high, hot_cold_sep, log-block mapping and Streams (the reserve).
  • Fewer, larger lines make GC coarser; more LUNs and channels raise throughput, not single-request latency.

6. NAND timing​

Applies to bbssd, CSD and KV; ZNS has its own in section 11. Reference: NAND timing. Design: NAND media and timing.

ParameterUnitValid valuesMeaning
pg_rd_lat, pg_wr_lat, blk_er_latns32-bitflat page read, page program and block erase times, used when nand_cell_type is 0
nand_cell_typetype0 flat, 1 SLC, 2 MLC, 3 TLC, 4 QLC; others fall back to 0built-in per-page-type tables instead of the flat times
pgtype_lat, cell_pagesflag, bits per cellcell_pages 0 to 5with nand_cell_type 0, scale the program time by the page's position in its wordline
cmd_addr_latns>= 0channel bus command and address phase
pg_xfer_latns>= 0channel bus data phase per page; 0 uses ch_xfer_lat
ch_xfer_latns>= 0data phase when pg_xfer_lat is 0; also the OCSSD 1.2 transfer time
status_latns>= 0channel bus status phase
tplebsyns>= 0busy time between the planes of a multi-plane erase (GC with pls_per_lun > 1)
pe_suspend, tsusp_nsflag, nstsusp_ns >= 0let a read suspend a program or erase on its LUN, at a cost of tsusp_ns
trim_lat_nsns per range>= 0; bbssd and CSD; refused with FDPtime charged per Dataset Management deallocate range

Interactions:

  • The channel bus is modelled only when cmd_addr_lat, pg_xfer_lat (or ch_xfer_lat) or status_lat is non-zero; phases on one channel then run one at a time.
  • Vendor command 0xEF codes 3 and 4 change only the flat times: code 3 sets the built-in values, not the configured ones, and code 4 sets them to 0. Neither has an effect while nand_cell_type is set.
  • Read time also grows with ecc_step_ns (section 9).
  • Tutorial 05 measures each of these.

7. FTL: mapping, caches and write buffer​

Applies to bbssd and CSD. Reference: garbage collection, mapping and caches. Design: the BlackBox FTL.

ParameterUnitValid valuesMeaning
mappingnamepage, dftl, hybrid, fastlogical-to-physical scheme: a full page table, a cached page table, BAST or FAST log-block mapping
mapping_cache_mbMiB32-bit; 0 picks the built-in sizewith mapping=dftl, the cached part of the table; a miss costs a NAND read
read_cache_mbMiB0 disablesDRAM read cache; a hit costs DRAM time instead of a NAND read
cache_evictnameclock, random, lru, arcread cache eviction policy
hot_cold_sepboolacts with page or dftl; ignored under hybrid and fast, which still reserve its line; refused with FDPwrite overwrites of mapped pages to separate hot lines
buffer_sizeNAND pages, not bytes>= 0; refused with FDPDRAM write buffer; 0 programs every write directly
buffer_thres_pcentpercent1 to 100 when buffer_size > 0fill level at which buffered pages are written to NAND
debug_ftlboolprint FTL invariant violations and log-block merge counts

Interactions:

  • The write buffer models timing; the data lives in the backend. With vwc=1 the guest sees a volatile write cache and feature 06h turns it off. Flush drains the buffer and FUA writes skip it with either setting. With power_loss=on the buffer also holds data, which simulate-power-loss drops.
  • hybrid and fast reserve one more line and count merges in log page C0h; hot_cold_sep reserves one more line.
  • mapping_cache_mb has no effect except under dftl.
  • Unknown names for mapping, gc_policy and cache_evict are refused.

8. Garbage collection​

Applies to bbssd and CSD; KV uses gc_thres_pcent only (but see KV). Reference: garbage collection, mapping and caches. Design: garbage collection.

ParameterUnitValid valuesMeaning
gc_thres_pcentpercent of lines in use1 to 100background GC starts; KV: share of the NAND usable for values
gc_thres_pcent_highpercent of lines in usegc_thres_pcent to 100forced GC inside writes; also sizes the reserve
gc_policynamegreedy, random, cost-benefit, fifo, d-choice; without FDPvictim line policy
gc_strategynumber0 greedy, 1 cost-benefit, 2 random, 4 per-handle; FDP onlyvictim reclaim unit policy
fdp_trim_erase_allflagFDP onlya deallocate resets every reclaim unit instead of the given ranges

Interactions:

  • Background GC takes only a line with at least 1/8 of its pages invalid; forced GC takes the best line whatever it holds.
  • Set gc_thres_pcent above the share of the NAND the namespace fills. Below it, background GC never stops and the WAF settles near 8 (tutorial 02).
  • The reserve is (1 - gc_thres_pcent_high / 100) x blks_per_pl lines, rounded down, plus the write pointers; the namespace must fit in the rest.

9. Reliability, wear and fault insertion​

Applies to bbssd, CSD and KV unless noted. Reference: reliability and wear. Design: wear, read reclaim and retention refresh.

ParameterUnitValid valuesMeaning
ecc_step_nsns per tier0 disablesextra read time per ECC tier: one per 750 erases of the block, one per ecc_retention_sec of data age, at most 4
ecc_retention_secseconds0 counts wear only; refused with FDPdata age that adds one tier
pe_cycles_ratedP/E cycles0 takes the rating of nand_cell_typedenominator of SMART Percentage Used
nand_bad_blocksblockscapped at the block countblocks bad from the start; lowers SMART Available Spare
err_read_unc_ppmper million0 disables; bbssd and CSDreads that fail as Unrecovered Read Error, at a fixed period
err_write_fail_ppmper million0 disables; bbssd, CSD and ZNSwrites that fail; a ZNS zone then becomes read only
read_reclaim_limitreads0 disables; bbssd and CSD; refused with FDPa block read this often since its erase gets its line rewritten on a following write
retention_limit_secseconds0 disables; bbssd and CSD; refused with FDPa read that hits a line filled at least this long ago queues the line, which is rewritten on a following write

Interactions: faults come at a fixed period, so a run repeats exactly. Read reclaim and retention refresh act only when the host reads and then writes; their cost appears as write amplification.

Applies to every NVMe mode. Reference: host link and controller firmware. Design: host link and firmware CPU.

ParameterUnitValid valuesMeaning
pcie_bandwidth_mbpsMB/s (10^6 bytes)0 disableseach Read and Write is charged its size at this rate, on one queue per direction
pcie_prop_delay_nsns0 disablesfixed delay added to each Read and Write after its transfer
fw_cpu_nsns per command0 disablesfirmware time per Read, Write and Zone Append on one modelled core; caps the rate at about one command per fw_cpu_ns

11. ZNS​

Reference: ZNS. Design: ZNS.

ParameterUnitValid valuesMeaning
zns_num_chchannels1 to 128channels
zns_num_lunLUNs>= 1LUNs per channel
zns_num_planeplanes1 to 8planes per LUN; the program unit grows with it
zns_num_blkblocks>= 1blocks per plane
zns_chnls_per_zonechannelsdivides zns_num_ch; 0 = allzone width
zns_zone_capbytesone block to the zone size; 0 = zone sizewritable part of each zone
zns_flash_typetype1 SLC, 2 MLC, 3 TLC, 4 QLC, 5 PLCcell type; MLC and PLC need the three times below
zns_pg_rd_lat, zns_pg_wr_lat, zns_blk_er_latns>= 0; 0 = built-inoverride the cell type's times
zns_cmd_addr_lat, zns_pg_xfer_lat, zns_status_latns>= 0channel bus phases
zns_pe_suspend, zns_tsusp_nsflag, nsreads suspend a program or erase on their plane
zns_max_open, zns_max_activezones<= zone count, open <= active; 0 = no limitMaximum Open and Active Resources
zns_num_wcwrite caches<= zone count; 0 picks from zns_max_openzone write caches
zns_zasl_bsbytespower-of-two multiple of 4 KiB; 0 follows mdtsZone Append size limit
zns_zd_ext_sizebytesmultiple of 64, up to 16320zone descriptor extension size
zns_num_conv_zoneszonescapped at the zone countleading conventional zones; Linux rejects them
zns_zrwa_size, zns_zrwafg_size, zns_zrwa_numblocks, blocks, zonesall three set or all 0Zone Random Write Area window, flush granularity and resources
zns_cross_zone_readboolallow reads across zone boundaries

Derived quantities (when the divisions are exact):

pages per block = namespace size / 16 KiB / (zns_num_ch x zns_num_lun x zns_num_blk)
zone width = zns_chnls_per_zone, or zns_num_ch when 0
zone size = zone width x zns_num_lun x zns_num_plane x pages per block x 16 KiB
zone count = zns_num_ch x zns_num_blk / (zone width x zns_num_plane)

Interactions:

  • The zone count does not depend on the namespace size; the zone size grows with it. Linux uses a zoned namespace only when the zone size is a power of two, so keep the size and the geometry powers of two.
  • lba_index must select a block of 4 KiB or less.
  • zns_zrwa_size must be a multiple of zns_zrwafg_size, and the zone capacity a multiple of zns_zrwafg_size.
  • The BlackBox geometry and timing properties do not apply to ZNS.

12. FDP​

Set on femu-subsys, which a bbssd controller joins with subsys=. Reference: Flexible Data Placement. Design: FDP.

ParameterUnitValid valuesMeaning
fdpboolFDP in endurance group 1 for the joining controller
fdp.nruhhandles1 to fdp.nru; required with fdp=onreclaim unit handles, the placement identifiers the host can name
fdp.nrureclaim unitsfdp.nruh to 65536reclaim units per group; bbssd uses at most one per line and needs 2 x fdp.nruh + 1
fdp.nrggroups1 onlyreclaim groups
fdp.runsbytes0, or one line's size with bbssdreclaim unit size; 0 lets FEMU choose
fdp.isolation_modenumber0, or any other value0: every handle Persistently Isolated; otherwise the last one Initially Isolated

Interactions:

  • The subsystem takes one controller with one namespace, and must come first on the command line.
  • The controller's gc_strategy and fdp_trim_erase_all apply; gc_policy other than greedy, mapping other than page, buffer_size, hot_cold_sep, read_reclaim_limit, retention_limit_sec, ecc_retention_sec, trim_lat_ns, meta, streams and ns_mgmt are refused.
  • KV is refused under FDP; NoSSD, ZNS and OCSSD report FDP but ignore it for placement.

13. OCSSD, CSD and KV​

OCSSD​

Reference: OCSSD.

ParameterUnitValid valuesMeaning
lverversion1 (1.2) or 2 (2.0)Open-Channel version
flash_typetype1 SLC, 2 MLC, 3 TLC, 4 QLCcell type for the built-in timing tables
lnum_ch, lnum_luncountlnum_ch 1 to 32, lnum_ch x lnum_lun <= 128channels (2.0 groups) and LUNs (parallel units)
lnum_plnplanes> 0; 1.2: 1, 2 or 4planes per LUN
lpgs_per_blk, lsecs_per_pgcount> 0; 1.2: pages <= 512pages per block, sectors per page
lsec_size, lmetasize, lmax_sec_per_rqbytes, bytes, sectors1.2 onlysector size, out-of-band bytes, most sectors per vector command
oc12_channel_timingbool1.2 onlycharge channel transfer per page, ch_xfer_lat or the flash_type value
learly_resetflag2.0 onlyreport the early reset capability

OCSSD takes the whole controller: one namespace, no namespace_modes neighbours.

CSD​

Reference: CSD. CSD builds the bbssd FTL, so sections 5 to 9 apply.

ParameterUnitValid valuesMeaning
fdm_sizeMiB> 0, requiredfunctional data memory
nr_cucompute units1 to 64programs run on the first free unit
csf_runtime_scalemultiplier> 0host run time multiplier for programs that declare no run time
csd_program_dirhost directorywhere shared-object and uBPF programs load from; unset allows only the built-in program type

KV​

KV uses the geometry (section 5), the NAND timing (section 6), and gc_thres_pcent as the share of the NAND usable for values. The FTL, mapping, cache, write buffer and other GC properties have no effect on KV, but KV runs the same validation as bbssd, so an out-of-range or unknown value is still refused. The GC capacity check of bbssd is not made: a namespace larger than the NAND starts, with its value space clamped. Design: KV.

14. CXL SSD​

femu-cxl-ssd is a separate device below a CXL root port. Reference: femu-cxl-ssd and its runtime properties. Design: CXL SSD.

ParameterUnitValid valuesMeaning
volatile-memdevmemory backend idrequired; a non-zero multiple of 256 MiB, at most 120 GiBthe backend that holds the data; its size is the device capacity
cache-pages4 KiB pages0, or up to the media page count and divisible by cache-waysDRAM page cache; 0 sends every access to the media
cache-wayswaysnon-zero, divides cache-pagesset associativity; 1 is direct mapped; changeable at run time
cache-policynamefifo, lifo, clock, s3-fiforeplacement within a set
prefetch-degree, prefetch-stridepagesup to the media page countpages inserted after a miss, and their distance; changeable at run time
ftlbooloff charges no media time and cannot be linked to an NVMe controller
channels, luns-per-channelcount1 to 4096, 1 to 128NAND channels and LUNs (one plane each)
pages-per-block, blocks-per-planecount1 to 65536; blocks 2 to 65536, or 0 to size itNAND blocks; 0 leaves spare room for GC
read-ns, program-ns, erase-ns, channel-nsnsat most one secondNAND times
gc-threshold, gc-threshold-highpercent1 to 100, high >= lowGC watermarks
dernameoff, memslot, cylondirect mapping of cached pages into the guest
der-replace-rateper second0 disablesmemslot alias replacements a hot page may cause
cylon-kernel-ackboolmust be on with der=cylonstates that the host runs a fixed Cylon kernel
concurrent-misseson, off, automisses to different pages wait for the media together
cylon-first-touch-program, cylon-free-writebackboolmedia rules of the published Cylon experiments
ccaboolthe caching API on BAR 5
lsa-controlboolnot with an lsa backendexperiment commands through Get LSA; trusted guests only
log-dir, tracefs-dir, log-limitpath, path, byteslog-limit 0 opens no I/O logwhere the control channel writes, and how much

Interactions:

  • The machine needs cxl=on, a pxb-cxl host bridge, a cxl-rp root port and a cxl-fmw window at least as large as the backend; run-cxlssd.sh builds them.
  • der=memslot needs KVM. der=cylon needs a Cylon host kernel and a shared, preallocated hugetlbfs backend; without them, or when the memory cannot be locked, it falls back to MMIO with a warning. It also needs -machine smm=off, which the device does not check (der=cylon).
  • A femu controller with cxl_ssd=<id> serves the same medium as an NVMe namespace. It needs femu_mode=1, one namespace, ftl=on on the CXL device, devsz_mb unset or equal to the medium's size, and none of ns_mgmt, subsys, streams, power_loss, buffer_size, op_pcent, meta, pi or dps. The medium's geometry and timing apply, not the controller's (CXL NVMe link).

15. Accepted for compatibility, no effect​

These properties exist so that old command lines still start. Setting one to anything but its default prints a warning at realize.

DeviceProperties
femuserial, ms, ms_max, dlfeat, tplpbsy, tplrbsy, trcbsy, nr_thread, time_slice, context_switch_time

nr_thread is still refused at 0 by CSD. Identify Controller reports a serial number FEMU generates, whatever serial says.

16. Worked configurations​

Each configuration below is started by the documentation checks, which also write and read one block through it.

A TLC drive with a channel bus and read suspend​

A 4 GiB namespace over the default geometry, TLC page-type timing, a 10 us page transfer and 2 us command and status phases on each channel, and reads that suspend a program or erase for 15 us of overhead:

-device femu,femu_mode=1,devsz_mb=4096,nand_cell_type=3,cmd_addr_lat=2000,pg_xfer_lat=10000,status_lat=2000,pe_suspend=1,tsusp_ns=15000

DFTL with caches and a write buffer​

A DFTL table with 8 MiB cached, a 64 MiB read cache with LRU eviction, and a write buffer of 4096 NAND pages that the guest sees as a volatile write cache. oncs=0x1c enables Write Zeroes (0x8) together with Dataset Management (0x4) and Save/Select (0x10):

-device femu,femu_mode=1,devsz_mb=2048,mapping=dftl,mapping_cache_mb=8,read_cache_mb=64,cache_evict=lru,buffer_size=4096,buffer_thres_pcent=75,vwc=1,oncs=0x1c

A GC study drive​

512 MiB of NAND exposed as NAND / 1.25, background and forced GC both starting at 95% of lines in use, cost-benefit victims and hot/cold separation, as in tutorial 02:

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25,gc_thres_pcent=95,gc_thres_pcent_high=95,gc_policy=cost-benefit,hot_cold_sep=on

A ZNS drive with 4 KiB blocks, limits and ZRWA​

64 zones of 64 MiB, 4 KiB logical blocks, 8 open and 16 active zones, and a ZRWA of 64 blocks flushed in units of 8 on up to 4 zones at once:

-device femu,femu_mode=3,devsz_mb=4096,zns_num_ch=8,zns_num_lun=4,zns_num_plane=2,zns_num_blk=128,lba_index=3,zns_max_open=8,zns_max_active=16,zns_zrwa_size=64,zns_zrwafg_size=8,zns_zrwa_num=4

FDP with eight handles​

Eight placement handles over 256 lines (blks_per_pl=256), with cost-benefit reclaim unit selection:

-device femu-subsys,id=fdp0,fdp=on,fdp.nruh=8 -device femu,femu_mode=1,devsz_mb=2048,blks_per_pl=256,subsys=fdp0,gc_strategy=1

A 3 GiB BlackBox namespace and a 1 GiB ZNS namespace on one controller, with a 3500 MB/s host link, 1 us of propagation delay and 2 us of firmware time per command:

-device femu,femu_mode=1,devsz_mb=4096,namespaces=2,namespace_sizes=3G,,1G,namespace_modes=bbssd,,znssd,pcie_bandwidth_mbps=3500,pcie_prop_delay_ns=1000,fw_cpu_ns=2000

Wear and faults​

A drive rated for 3000 program/erase cycles, with 100 bad blocks, an ECC step of 5 us, one read in 100,000 uncorrectable and one write in 1,000,000 failed:

-device femu,femu_mode=1,devsz_mb=1024,pe_cycles_rated=3000,nand_bad_blocks=100,ecc_step_ns=5000,err_read_unc_ppm=10,err_write_fail_ppm=1

Namespace management​

A controller that starts with one bbssd namespace and lets the guest allocate up to four bbssd namespaces in all, the boot one included, so three more:

-device femu,femu_mode=1,devsz_mb=2048,ns_mgmt=on,bbssd_ns_limit=4

A CXL SSD and an NVMe view of it​

A 256 MiB femu-cxl-ssd with a 1024-page, 4-way s3-fifo cache below a CXL host bridge, and a femu controller that serves the same medium as an NVMe namespace:

./qemu-system-x86_64 -machine q35,cxl=on \
-object memory-backend-ram,id=cxlmem,size=256M \
-device pxb-cxl,id=cxl.0,bus=pcie.0,bus_nr=52 \
-device cxl-rp,id=cxl-rp0,bus=cxl.0,chassis=0,slot=0 \
-device femu-cxl-ssd,id=cxlssd,bus=cxl-rp0,volatile-memdev=cxlmem,cache-pages=1024,cache-ways=4,cache-policy=s3-fifo \
-device femu,id=nvme0,bus=pcie.0,femu_mode=1,cxl_ssd=cxlssd \
-M cxl-fmw.0.targets.0=cxl.0,cxl-fmw.0.size=256M