Skip to main content

FEMU architecture

Mirrored from the FEMU repository

This page is hw/femu/docs/concepts/architecture.md at FEMU 39a55eeb6 (2026-10-02), licensed GPL-2.0-or-later. Send corrections to the FEMU repository.

FEMU is QEMU with a set of emulated storage devices under hw/femu/. This page describes those devices as a stack of layers, from what the guest sees down to the host memory that holds the data. Each section says what the layer does, where its code lives, which threads run it, and where it charges time.

Two ideas explain most of the design:

  • Data and time are separate. The guest's data is copied into a host memory buffer as soon as a command is parsed. The FTL and NAND model never touch the data; they only compute how long the command would take on a real device, and the completion is held back until that time has passed. (The power-loss model is the exception: with power_loss=on the FTL thread applies the data.)
  • NVMe I/O commands run on FEMU's own threads, not on the vCPU. A doorbell write only records the new queue tail. Polling threads fetch and execute the commands, and an FTL thread computes their timing.

The CXL SSD (femu-cxl-ssd) is the exception to the second rule: its loads and stores are served on the vCPU thread that issued them. The CXL SSD design note covers it in depth.

The layers​

FEMU’s six layers, the code that implements each one, and the thread that runs it; data is copied from the frontend straight to the memory backend, past the FTL and NAND model.

Figure: FEMU's six layers, the code that implements each one, and the thread that runs it; data is copied from the frontend straight to the memory backend, past the FTL and NAND model.

Solid arrows carry commands and timing. Dotted arrows carry the data, which goes straight to the memory backend.

LayerMain source filesThreads
1. Guest-visible interfacehw/femu/femu.c, hw/femu/femu-props.c, hw/femu/intr.c, hw/femu/cxlssd/qemu-adapter.cvCPU threads, QEMU main loop
2. Frontendhw/femu/nvme-io.c, hw/femu/nvme-admin.c, hw/femu/dma.c, hw/femu/lib/, hw/femu/cxlssd/cxlssd.c, hw/femu/cxlssd/cache.cfemu-poller, vCPU threads
3. Mode backendshw/femu/nossd/, hw/femu/bbssd/bb.c, hw/femu/zns/zns.c, hw/femu/ocssd/, hw/femu/kvssd/, hw/femu/csd/femu-poller
4. FTLhw/femu/bbssd/, hw/femu/zns/zftl.c, hw/femu/kvssd/kvssd-ftl.cFEMU-FTL-Thread, femu-poller (KV), femu-cxl-ftl
5. NAND media timinghw/femu/nand/, hw/femu/timing-model/the caller's thread
6. Memory backendhw/femu/backend/dram.c; for CXL, a QEMU memory backend objectthe thread that copies the data

1. Guest-visible interface​

FEMU registers three device types.

-device femu is a PCIe NVMe controller. It reports NVMe version 1.4 and PCI ID 1d1d:1f1f by default (vid, did). BAR0 holds the controller registers and the doorbells, MSI-X has one vector per queue plus the admin queue, and an optional Controller Memory Buffer (cmbsz) must be placed on BAR2 with cmbloc. femu_mode selects which kind of SSD the controller emulates (see layer 3 and choosing a mode).

Namespaces are not separate devices. There is no FEMU namespace device to put on the command line; the controller builds its namespaces at realize from namespaces, namespace_sizes and namespace_modes, and Namespace Management can add more at run time when ns_mgmt=on. Migration and snapshots are blocked for this device.

-device femu-subsys is an NVMe subsystem. It is not on any bus. A controller joins it with subsys=<id>, and the subsystem must appear first on the command line. It is needed for Flexible Data Placement (fdp=on) and for namespaces shared between controllers (ns_mgmt=on on the subsystem).

-device femu-cxl-ssd is a CXL Type-3 volatile memory device, a subclass of QEMU's cxl-type3. Its data lives in the memory backend named by volatile-memdev. FEMU installs an I/O region named femu-cxl-media over every CXL fixed memory window that can reach the device, so guest loads and stores to the device's address range reach FEMU. Optional extras are a cache control interface on BAR5 (cca=on) and an experiment control channel through the label storage area (lsa-control=on). A femu controller can serve the same medium as an NVMe namespace with cxl_ssd=<id>.

Where it lives:

  • hw/femu/femu.c: QOM types femu and femu-subsys, femu_realize(), nvme_check_constraints(), PCI and BAR setup, MMIO and doorbell handlers (nvme_mmio_write()).
  • hw/femu/femu-props.c: the help text of the femu and femu-subsys properties. The properties and their defaults are defined in femu.c (femu_props[]). The generated property reference is built from the binary.
  • hw/femu/intr.c: MSI-X, MSI and pin interrupts (nvme_isr_notify_io()).
  • hw/femu/cxlssd/qemu-adapter.c: QOM type femu-cxl-ssd, realize checks, window overlay and address decoding, and its properties (cxl_props[]). hw/femu/cxlssd/props.c: their help text.

Threads: register and doorbell writes are MMIO exits handled on the vCPU thread that made them, under QEMU's big lock (BQL).

2. Frontend​

NVMe queues and the admin path​

The admin queue is handled synchronously. A write to the admin submission queue doorbell calls nvme_process_sq_admin() on the vCPU thread, which runs nvme_admin_cmd() in hw/femu/nvme-admin.c for each new entry. Commands the generic code does not know go to the mode's admin hook, for example the BBSSD vendor command 0xEF.

I/O queues are handled asynchronously. A write to an I/O submission queue doorbell only stores the new tail (nvme_process_db_io()). The host can also set up shadow doorbells with the standard Doorbell Buffer Config command (nvme_set_db_memory()); a queue with a shadow doorbell is then read from that buffer, and FEMU publishes EventIdx values so the host can skip doorbell writes it does not need.

Pollers​

When the host enables the controller, nvme_start_dataplane() creates the poller threads (femu-poller) once. Each poller loops over the I/O queues it owns (nvme_poller() in hw/femu/nvme-io.c). How many pollers there are is set by two properties:

multipoller_enabledPollersQueues per poller
0 (default)1all I/O queues
1ceil(queues / poller_ratio)queues i, i + N, i + 2N, ... for poller i of N

Any other value fails realize. A poller fetches commands from a queue without a lock, so every queue must have exactly one owner.

Each poller owns three structures, all indexed by poller number:

  • to_ftl[i]: a ring of requests waiting for timing.
  • to_poller[i]: a ring of timed requests coming back from the FTL thread.
  • pq[i]: a priority queue of timed requests, ordered by completion time.

The rings are in hw/femu/lib/rte_ring.c and the priority queue in hw/femu/lib/pqueue.c.

I/O dispatch​

For each new submission queue entry, nvme_process_sq_io():

  1. copies the command and stamps the request with the current host time (QEMU_CLOCK_REALTIME) as both its start time and its completion time;
  2. calls nvme_io_cmd(), which handles the commands every block mode shares (Flush, Dataset Management, Compare, Write Zeroes, Copy, Verify, Write Uncorrectable, I/O Management) and passes the rest to the namespace's mode (ns->ext_ops.io_cmd);
  3. puts the request on to_ftl[i].

For Read and Write on NoSSD, BBSSD and CSD, the mode handler calls nvme_rw(), which checks the command, maps the guest's PRP or SGL list (hw/femu/dma.c) and copies the data between guest memory and the memory backend with backend_rw(). The data transfer is finished at this point, on the poller thread, before any time has been charged. ZNS and OCSSD have their own read and write paths that copy to the same backend, and KV copies values into a store of its own.

On a NoSSD controller, NoSSD namespaces skip the rings: the poller posts the completion inside the same sweep, unless the host-link or firmware-CPU model is enabled.

Completion​

nvme_process_cq_cpl() runs at the end of every sweep. It drains the timed requests (from to_poller[i] when an FTL thread exists, otherwise straight from to_ftl[i]), adds the optional host-link and firmware-CPU time, and inserts them into pq[i]. Then it posts every request whose completion time has passed and raises one interrupt per completion queue it posted to. A request that is not due yet stays in the queue for the next sweep. This is where FEMU enforces latency: nothing sleeps, the completion is simply not posted before its time.

CXL accesses​

The CXL device has no queues. A guest load or store to its range is an MMIO access on the vCPU thread, routed through the host bridge and endpoint decoders to femu_cxl_access() in hw/femu/cxlssd/cxlssd.c. A page-granular cache (hw/femu/cxlssd/cache.c, policies fifo, lifo, clock, s3-fifo) decides whether the access needs the FTL. The cache keeps only page numbers and dirty bits; the data is always in the memory backend. With der=memslot or der=cylon, cached pages can be mapped into the guest so that later accesses do not reach QEMU at all (see the CXL load walkthrough).

3. Mode backends​

A mode is a table of handlers (FemuExtCtrlOps in hw/femu/nvme.h). nvme_register_extensions() in hw/femu/femu.c installs the table for the controller's femu_mode, and nvme_register_extensions_ns() gives each namespace the table for its own mode, so one controller can mix modes with namespace_modes. I/O commands are dispatched on the namespace's mode, not the controller's.

Modefemu_modeCodeWhat the I/O handler doesWhere time is computed
OCSSD0hw/femu/ocssd/oc12.c (lver=1), hw/femu/ocssd/oc20.c (lver=2)Open-Channel vector commands; the host runs the FTLin the poller, from chip and channel timestamps (hw/femu/timing-model/timing.c)
BBSSD1hw/femu/bbssd/bb.cRead and Write through nvme_rw()FTL thread, bb_ftl_process_req()
NoSSD2hw/femu/nossd/nop.cRead and Write through nvme_rw()none
ZNS3hw/femu/zns/zns.czoned command set, zone state machineFTL thread, zns_ftl_process_req()
CSD4hw/femu/csd/csd.ccomputational storage commands plus block I/OFTL thread for NAND; programs and compute-unit time on the femu-csd-cu threads
KV5hw/femu/kvssd/kvssd.ckey-value command set: Store, Retrieve, Delete, Exist, Listin the poller, through the BBSSD NAND model

Features that are not modes:

  • Flexible Data Placement: enabled on femu-subsys; only BBSSD places data by reclaim unit (hw/femu/bbssd/ftl-fdp.c).
  • Several namespaces and per-namespace modes: namespaces, namespace_sizes, namespace_modes (nvme_init_namespaces() in hw/femu/femu.c).
  • Namespace Management: ns_mgmt=on on a NoSSD or BBSSD controller.
  • Metadata and protection information: meta, mc, pi (hw/femu/nvme-pi.c), NoSSD and BBSSD only.
  • Streams: streams=on (hw/femu/nvme-streams.c).
  • Persistent Event Log: always on, kept in pel_file if set (hw/femu/nvme-pel.c).

Choosing a mode lists which of these combine.

Threads: one FEMU-FTL-Thread per controller, started at realize only when a namespace is BBSSD, ZNS or CSD, or the controller shares a BBSSD subsystem (femu_needs_ftl_thread()). It reads every poller's to_ftl[i] ring, calls femu_ftl_process_req(), adds the returned latency to the request's completion time and puts it on to_poller[i]. A controller whose namespaces are all NoSSD, OCSSD or KV has no FTL thread.

4. FTL​

BBSSD​

The BBSSD FTL in hw/femu/bbssd/ is a page-level device FTL. CSD uses it unchanged, and the CXL SSD runs a private instance of it. Its state is struct ssd (hw/femu/bbssd/ftl.h).

  • Geometry. Channels (nchs) hold LUNs (luns_per_ch), LUNs hold planes (pls_per_lun), planes hold blocks (blks_per_pl), blocks hold pages (pgs_per_blk) of secs_per_pg sectors of secsz bytes. With the defaults a page is 4 KiB. A line (superblock) is the same block index on every plane of every LUN, so there are blks_per_pl lines.
  • Write pointer. New pages are allocated from the current line in the order channel, then LUN, then plane, then page, so consecutive pages land on different channels and LUNs and can be programmed in parallel.
  • Mapping (mapping): page (default, a full table in memory), dftl (a cached mapping table of mapping_cache_mb whose misses cost NAND reads), hybrid and fast (log-block schemes). Code: hw/femu/bbssd/ftl-map*.c.
  • Write buffer (buffer_size, in pages): a DRAM buffer that accepts writes. Once it reaches buffer_thres_pcent, the next write that needs a slot programs a batch of the least recently written pages. Without power_loss it holds page numbers only, since the data is already in the memory backend. FUA writes and stream writes go straight to NAND, and Flush drains it. Code: hw/femu/bbssd/ftl-datapath.c.
  • Read cache (read_cache_mb, cache_evict): a DRAM cache of recently read pages (hw/femu/bbssd/ftl-cache.c).
  • Garbage collection. After every request the FTL thread runs one background GC pass if the share of lines in use has reached gc_thres_pcent. A write that finds it at gc_thres_pcent_high first runs GC in the foreground until it drops below. The victim line is chosen by gc_policy (greedy by default). GC reads, programs and erases occupy the same LUNs as host I/O. Code: hw/femu/bbssd/ftl-line-gc.c.
  • Placement extras: hot_cold_sep writes overwritten pages to separate lines, read_reclaim_limit and retention_limit_sec rewrite lines that were read too often or hold old data, and FDP maps reclaim units to lines.
  • Wear: each block counts its erases. SMART Percentage Used compares the average erase count with pe_cycles_rated, nand_bad_blocks lowers Available Spare, and ecc_step_ns makes reads of worn or old blocks slower.

Entry point: bb_ftl_process_req() in hw/femu/bbssd/ftl.c, which dispatches to ssd_read(), ssd_write(), ssd_trim() and the others in hw/femu/bbssd/ftl-datapath.c.

Other FTLs​

  • ZNS (hw/femu/zns/zftl.c): zones map onto blocks striped over zns_chnls_per_zone channels. There is no device GC; the host resets zones. Writes collect in per-zone SRAM write caches (zns_num_wc) and are programmed when a cache fills or is evicted. Zone Reset erases the zone's blocks.
  • KV (hw/femu/kvssd/kvssd-ftl.c): a hash index of keys, with values stored on NAND lines laid out with the BBSSD geometry and timing.
  • OCSSD has no device FTL. The host addresses channels, LUNs, blocks and pages directly.

Threads: FEMU-FTL-Thread runs the BBSSD, CSD and ZNS FTLs. The CXL SSD's FTL runs on its own femu-cxl-ftl worker; when an NVMe controller is linked with cxl_ssd=, both threads use the same FTL and take turns on one mutex.

5. NAND media timing​

hw/femu/nand/nand-media.c turns one NAND operation into a completion time. nand_media_op() takes a location (channel, LUN, plane, block, page), an operation (read, program, erase) and a start time. It keeps a busy-until time per LUN (BBSSD, CSD, KV) or per plane (ZNS): the operation starts when both the request and that unit are ready, and the unit stays busy until it ends. Units that differ run in parallel. A channel bus with command, data and status phases is modelled only when one of those phases has a non-zero time.

Read, program and erase times come either from flat properties or from built-in per-cell-type tables (hw/femu/nand/nand.h for BBSSD, CSD and KV; hw/femu/zns/zns.h for ZNS). OCSSD uses the older chip and channel timestamp model in hw/femu/timing-model/timing.c.

Timing model explains the rules and the properties that control them.

Threads: none of its own. It runs on whichever thread asks: the FTL thread for BBSSD, CSD and ZNS, the poller for KV and OCSSD, and the femu-cxl-ftl worker for the CXL SSD.

6. Memory backend​

For femu, the data lives in one host buffer of devsz_mb MiB (the full NAND capacity for BBSSD with op_pcent), allocated zero-filled at realize and locked into RAM when RLIMIT_MEMLOCK allows (init_dram_backend() in hw/femu/backend/dram.c). Namespaces are slices of it, packed back to back. Controllers that share namespaces through a subsystem share one buffer. The block modes read and write it with backend_rw(); KV keeps its values in a separate buffer of its own. Nothing is written to a file, so the contents are lost when QEMU exits. FEMU_MBE_INTERLEAVE controls its NUMA placement (see the environment variables).

For femu-cxl-ssd, the data lives in the QEMU memory backend object named by volatile-memdev, for example -object memory-backend-ram,id=cxlmem,size=4G. It must be a non-zero multiple of 256 MiB and at most 120 GiB. der=cylon needs a shared, preallocated hugetlbfs backend. The vCPU thread copies data to and from it directly. A femu controller linked with cxl_ssd= uses this backend instead of allocating its own.

Threads​

FEMU’s threads and the queues between them: vCPUs store doorbells, each poller owns its submission queues and its to_ftl/to_poller rings, and one FTL thread per controller times BBSSD, ZNS and CSD requests.

Figure: FEMU's threads and the queues between them: vCPUs store doorbells, each poller owns its submission queues and its to_ftl/to_poller rings, and one FTL thread per controller times BBSSD, ZNS and CSD requests.

Thread nameCreated byHow manyRuns
vCPU threads and the QEMU main loopQEMUone per vCPU, plus one main loopMMIO: controller registers, doorbells, admin commands; CXL accesses and their media wait; QMP and monitor commands
femu-pollernvme_start_dataplane() when the host enables the controller1, or ceil(queues / poller_ratio)I/O command fetch and execution, data copy, completion and interrupts
FEMU-FTL-Threadfemu_realize()0 or 1 per controllerBBSSD, CSD and ZNS FTL and NAND timing
femu-cxl-ftlfemu_cxl_start()1 per femu-cxl-ssd with ftl=onFTL and NAND timing for cache misses and write-backs
femu-cxl-ccafemu_cxl_cca_start()1 per femu-cxl-ssd with cca=oncache control commands from BAR5
femu-csd-cucsd_init()nr_cu per controller with a CSD namespaceCSD programs (Execute) and their compute-unit time

Where latency is charged and enforced​

Device or modeComputed inEnforced by
NoSSDnothing to computecompletion posted in the same sweep
BBSSD, CSD, ZNSFEMU-FTL-Thread adds the FTL's latency to the completion time; for a CSD Execute, the femu-csd-cu thread that ran the program adds its compute-unit time insteadthe poller posts once the time has passed
OCSSD, KVthe poller, while executing the commandthe poller posts once the time has passed
Host link, firmware CPU (pcie_bandwidth_mbps, pcie_prop_delay_ns, fw_cpu_ns)the poller, before queueing the completionsame
femu-cxl-ssdfemu-cxl-ftl returns the media time of a miss or write-backthe vCPU waits it out before the load or store completes

I/O walkthrough: one 4 KiB write on BBSSD​

One 4 KiB write on BBSSD over host time: the poller copies the data at fetch, the FTL thread books a 200 µs program on the LUN from that moment, and the poller holds the completion until it is due.

Figure: One 4 KiB write on BBSSD over host time: the poller copies the data at fetch, the FTL thread books a 200 µs program on the LUN from that moment, and the poller holds the completion until it is due.

One 4 KiB read on BBSSD: the guest buffer is filled at fetch and only the completion waits; below, the checks ssd_read() makes for each page and what each one costs.

Figure: One 4 KiB read on BBSSD: the guest buffer is filled at fetch and only the completion waits; below, the checks ssd_read() makes for each page and what each one costs.

Assume the defaults: femu_mode=1, one poller, buffer_size=0, 4 KiB pages, pg_wr_lat=200000 (200 us), no channel bus model.

  1. Submit. The guest NVMe driver writes the command into its submission queue in guest memory and writes the new tail to the queue's doorbell. The vCPU exits to QEMU, and nvme_process_db_io() stores the tail. Nothing else happens on the vCPU.
  2. Fetch and copy. The poller sees the new tail on its next sweep. nvme_process_sq_io() copies the command and stamps it with the current time t0. nvme_io_cmd() passes it to bb_io_cmd(), which calls nvme_rw(): the LBA range and size are checked, the PRPs are mapped, and backend_rw() copies the 4 KiB from guest memory into the memory backend at the namespace's offset. The data is now where later reads will find it.
  3. Hand to the FTL. The poller puts the request on to_ftl[1] and goes on with other queues.
  4. Time it. The FTL thread takes the request and calls bb_ftl_process_req(), which calls ssd_write(). If too few lines are free, foreground GC runs first. The one logical page is mapped to the next physical page at the write pointer, the old copy (if any) is marked invalid, and ssd_advance_status() asks the media layer for a program on that page's LUN. If the LUN is idle the program takes 200 us from t0; if it is busy until t1, it ends at t1 + 200 us. The LUN stays busy until then. If enough lines are now in use, one background GC pass follows; its NAND operations make the LUNs they use busy for later requests.
  5. Return. The FTL thread adds the latency to the request's completion time and puts it on to_poller[1].
  6. Hold. On a following sweep, nvme_process_cq_cpl() moves the request into the priority queue. It stays there until the host clock passes its completion time.
  7. Complete. The poller writes the completion entry into the guest's completion queue and, at the end of the sweep, raises the queue's interrupt. The guest sees the write complete about 200 us after the poller fetched it on an idle device.

With buffer_size set, step 4 puts the page in the write buffer instead and charges a DRAM access (pg_rd_lat / 16, 2.5 us with the defaults). The program happens on a later write that pushes the buffer past its threshold, and that write pays for it.

CXL load walkthrough​

The guest loads 8 bytes from a femu-cxl-ssd region at host physical address hpa. With the cache on (cache-pages > 0):

  1. Trap. Unless the page is mapped directly (below), the load exits to QEMU on the vCPU thread. The femu-cxl-media overlay routes it through the window, host bridge and endpoint decoders to a device address, and calls femu_cxl_access().

  2. Hit. If the page is in the cache, no media time is charged. The vCPU copies the 8 bytes from the memory backend and returns. The cost is the MMIO exit itself, paid on every access.

  3. Miss. The page must be brought into the cache:

    • The page is read from NAND: the vCPU drops the BQL, queues a read for the femu-cxl-ftl worker, and waits for the latency it returns. A page the FTL has never mapped costs nothing to read.
    • If the page's set is full, the policy evicts a victim. A dirty victim is written back, which costs a NAND program.
    • femu_cxl_delay() waits out the rest of the media time, sleeping until 100 us before the deadline and spinning for the rest, with the BQL dropped so other vCPUs keep running.
    • The vCPU copies the data from the memory backend. With prefetch-degree set, the next pages are inserted into the cache with no NAND read.

    A store follows the same path and marks the page dirty. A store miss reads the page first, as a write-allocate cache does.

What happens to the next access to the same page depends on der:

derAfter a miss or hitNext access to the page
off (default)nothingtraps again; a hit with no media time
memslotthe page is mapped as a one-page KVM memory slot alias (at most 1024 for all devices)goes straight to host memory with no exit and no time charged; the page counts as dirty, so its eviction costs a program
cylonthe page is mapped by writing the guest's EPT entry through a Cylon host kernelgoes straight to host memory; dirty state comes from EPT dirty bits

A direct mapping is removed when its page is evicted, flushed or invalidated, and the next access traps again. Both direct modes map pages only in the single-endpoint topology the design note describes; elsewhere accesses stay on MMIO. der=memslot fails realize under TCG. der=cylon fails realize without cylon-kernel-ack=on, and falls back to MMIO with a warning when the Cylon host kernel or the hugetlbfs backend is missing. The CXL SSD design note gives the details.