Skip to main content

The BlackBox FTL

Mirrored from the FEMU repository

This page is hw/femu/docs/design/ftl.md at FEMU 39a55eeb6 (2026-10-02), licensed GPL-2.0-or-later. Send corrections to the FEMU repository.

This chapter describes the flash translation layer (FTL) behind BlackBox mode (femu_mode=1): how it maps logical pages to NAND pages, how it organises NAND into lines and write pointers, how it collects garbage, and how its write buffer, read cache, wear and refresh models work. It is written for readers who want to know what the model does, where its time goes, and how to change it. To run the mode, read the BlackBox guide first; this chapter does not repeat its launch and guest steps.

The code lives in hw/femu/bbssd/. The same FTL serves CSD namespaces, which build it through the same ssd_init(). KV namespaces call ssd_init() for the geometry and NAND timing but run their own key-value FTL. A femu-cxl-ssd builds a private instance on its own worker thread.

NAND operation timing (per-LUN timelines, the channel bus, suspend, ECC tiers) is the subject of the NAND timing chapter in this directory, and of the timing model. This chapter covers what the FTL asks of the NAND, not how long each operation takes.

Place in the hierarchy​

guest NVMe driver
| SQ doorbell
v
+---------------------------+ copies the payload between guest memory
| femu-poller thread(s) | and the host memory backend, then hands
| nvme_rw(), nvme_dsm(),...| every I/O request to the FTL ring
+---------------------------+
| to_ftl[i] ^ to_poller[i]
v |
+---------------------------------------------------------------+
| FEMU-FTL-Thread: femu_ftl_thread() -> bb_ftl_process_req() |
| |
| +-------------+ +--------------+ +----------------------+ |
| | mapping | | write buffer | | read cache | |
| | page, dftl, | | (LPNs, LRU) | | (LPNs, timing only) | |
| | hybrid,fast | +--------------+ +----------------------+ |
| +-------------+ |
| +---------------------------------------------------------+ |
| | line management, write pointers, GC, read reclaim | |
| +---------------------------------------------------------+ |
+---------------------------------------------------------------+
| ssd_advance_status(): one read, program or erase
v
+---------------------------------------------------------------+
| NAND media layer (hw/femu/nand/nand-media.c) |
| per-LUN next-available times, optional channel bus |
+---------------------------------------------------------------+

Two things follow from this picture.

  • The FTL models time and placement, not data. The poller copies the payload to the host memory backend before the request reaches the FTL. The FTL tracks which NAND page each logical page would occupy and how long each operation takes, but it never stores the payload. The one exception is the power-loss model (power_loss=on), which runs the data copy on the FTL thread and keeps undo copies of buffered pages; see Power loss.
  • The FTL returns a duration. Each handler returns the time the request spends in the device, measured from the request's arrival time stime. The FTL thread adds it to req->expire_time, and the poller completes the request when the host clock reaches that time.

The FTL thread and its queues​

poller 1 FEMU-FTL-Thread poller 1
+---------+ to_ftl[1] +-------------------------+ to_poller[1] +-----------+
| SQ scan |------------->| |------------->| pq[1] |
+---------+ (rte_ring) | for i in 1..nr_pollers | | ordered by|
poller 2 | dequeue one request | | expire_ |
+---------+ to_ftl[2] | lat = femu_ftl_ | to_poller[2]| time; CQE |
| SQ scan |------------->| process_req(req) |------------->| posted |
+---------+ | expire_time += lat | | when now |
... | enqueue to_poller[i] | | >= expire |
+-------------------------+ +-----------+
|
| after every request:
v
background GC check (one victim)
  • One thread per controller. femu_ftl_thread() in hw/femu/femu.c starts at realize when a namespace is bbssd, ZNS or CSD (or the controller shares a bbssd subsystem), after every namespace is built. It is joined at device removal.
  • One pair of rings per poller. Each poller i feeds to_ftl[i] and reads to_poller[i]. The FTL thread visits the rings in turn and takes one request from each non-empty ring per pass, so no poller can starve another.
  • Every I/O command goes through the ring, including Flush, Dataset Management and commands that already failed in the I/O layer. A failed request is not processed: bb_ftl_process_req() returns 0 for it, and the poller posts the carried error status.
  • No lock protects the FTL state. The FTL thread is the only writer of the mapping, the lines, the write buffer and the caches. With Streams on, each request runs under n->streams_lock; on a shared subsystem namespace it runs under the subsystem's ns_lock. An NVMe controller linked to a femu-cxl-ssd runs its requests on its own FTL thread against the medium's FTL, under the medium's lock, which also serializes the medium's worker.
  • Admin commands that change FTL state pause the dataplane. nvme_pause_pollers() waits until the FTL thread is outside a request (the ftl_in_sweep flag), so the run-time switch (admin opcode 0xEF), the volatile write cache feature, Format and Sanitize change FTL state with nothing in flight.

bb_ftl_process_req() in hw/femu/bbssd/ftl.c dispatches on the opcode:

OpcodeHandlerWhat the FTL does
Readssd_read()translate each page; buffer hit, read cache hit or NAND read
Writessd_write()foreground GC, read reclaim, then buffer or program each page
Write Zeroesssd_write_zeroes()with the Deallocate bit: unmap; without it: program the range
Dataset Managementssd_trim()deallocate each range
Copyssd_copy()read every source range, then write the destination
Flushssd_buffer_destage(ssd, 0, ...)program everything the write buffer holds
I/O Management Sendssd_fdp_update_ruhs()FDP only: update reclaim unit handles

Read and Write latency is the largest per-page latency of the request: pages on different LUNs proceed in parallel, and pages on one LUN queue on its timeline. Copy costs the slowest source read plus the destination write. After the opcode handler, the thread runs one background GC step if the device is past the background watermark (Triggers).

Data structures​

All of the FTL's state hangs off struct ssd (hw/femu/bbssd/ftl.h), one per bbssd or CSD namespace.

Physical page address​

A struct ppa packs a NAND location into 64 bits:

63 62 51 50 44 43 40 39 32 31 16 15 0
+---+--------------+----------+------+----------+----------------+----------------+
|rsv| ch (12) | lun (7) |pl (4)| sec (8) | pg (16) | blk (16) |
+---+--------------+----------+------+----------+----------------+----------------+

The field widths bound the geometry: at most 4096 channels, 128 LUNs per channel, 16 planes per LUN, 65536 blocks per plane, 65536 pages per block and 256 sectors per page. bb_check_geometry() refuses larger values, and also refuses a geometry whose total sector count exceeds INT_MAX, because the derived totals in struct ssdparams are int. The FTL always addresses whole pages; the sec field is unused by the mapping.

NAND state​

ssd->ch[nchs]
.next_ch_avail_time
.lun[luns_per_ch]
.next_lun_avail_time <- the timeline every NAND op is gated on
.pl[pls_per_lun]
.blk[blks_per_pl]
.vpc, .ipc valid and invalid page counts
.erase_cnt P/E cycles of this block
.read_cnt reads since the last erase
.pg[pgs_per_blk].status FREE, VALID or INVALID

A page goes FREE -> VALID when it is programmed (mark_page_valid()), VALID -> INVALID when its logical page is overwritten, deallocated or moved by a log-block merge (mark_page_invalid()), and back to FREE when its block is erased (mark_block_free()). A page that line GC relocates stays VALID until its block is erased. Closing a stream's line marks its unwritten pages INVALID directly.

Mapping tables​

maptbl: logical page -> physical page rmap: physical page -> logical page
(one struct ppa per logical page) (indexed by ppa2pgidx(), one u64 each)

LPN PPA pgidx LPN
+-----+--------------------------+ +------------------------+-----+
| 0 | ch0 lun0 pl0 blk7 pg3 |---------->| ch0 lun0 pl0 blk7 pg3 | 0 |
| 1 | UNMAPPED | | ch1 lun0 pl0 blk7 pg3 | 5 |
| 2 | ch3 lun1 pl0 blk2 pg0 | | ... | ... |
| ... | ... | | (stale copy) | INV |
+-----+--------------------------+ +------------------------+-----+

maptbl is the forward map and rmap the reverse map that GC uses to find which logical page a valid physical page holds. Both have one entry per NAND page (tt_pgs), 8 bytes each: 32 MiB apiece for the default 16 GiB geometry. Every mapping scheme keeps these two tables as the source of truth.

Lines​

typedef struct line {
int id; /* the block index this line spans */
int ipc, vpc; /* invalid and valid pages over the whole line */
size_t pos; /* slot in the victim priority queue, 0 if none */
uint64_t close_time; /* when the line filled; age-based policies */
uint64_t stream_tag; /* Streams: the stream whose data it holds */
bool reclaiming; /* being rewritten by read reclaim */
...
} line;

struct line_mgmt keeps a free list (a FIFO), a victim priority queue keyed by valid page count, and a full list.

Address mapping​

The four BBSSD mapping schemes (mapping=): the full page table, DFTL’s cache of map pages, BAST’s log block per data block, and FAST’s shared log pool.

Figure: The four BBSSD mapping schemes (mapping=): the full page table, DFTL's cache of map pages, BAST's log block per data block, and FAST's shared log pool.

From LBA to logical page​

ssd_lpn_range() turns an LBA range into a range of logical page numbers (LPNs). A page is secsz * secs_per_pg bytes. The byte offset is the LBA times the LBA size, plus the namespace's offset in the backend:

LPN = (backend_offset + slba * lba_size) / page_size

So a namespace's LPNs start where its data starts in the backend. A controller with Namespace Management on uses offset 0, because each managed namespace has a private FTL. A request whose last LPN is at or beyond tt_pgs fails with LBA Out of Range.

Page mapping (mapping=page)​

The default. translate() returns maptbl[lpn]. A write invalidates the old physical page, if any, and points the LPN at the new one. A read of an LPN with no mapping costs no NAND time.

DFTL (mapping=dftl)​

DFTL keeps the page-level mapping on NAND and caches the translation pages that are in use. FEMU models the cost of that cache, not its contents: the flat maptbl stays the source of truth, and cmt_touch() in hw/femu/bbssd/ftl-map-cmt.c charges NAND time for misses.

host read/write of lpn
|
v
tp_id = lpn / lpn_per_tp lpn_per_tp = page_size / 8
| (512 for a 4 KiB page)
v
+-------------------------------+
| cached mapping table (CMT) | capacity = mapping_cache_mb / page_size
| CLOCK, one slot per TP | slots
+-------------------------------+
| hit: no cost; a write marks the slot dirty
| miss:
v
evict a slot by CLOCK
| dirty victim: program it on LUN (victim tp_id % tt_luns)
v
read the TP on LUN (tp_id % tt_luns), after the write-back
|
v
latency = write-back + read, charged to this request

With mapping_cache_mb left at 0, DFTL uses 4 MiB, which holds 1024 translation pages and so covers 2 GiB of logical space on a 4 KiB page. The translation-page traffic occupies the LUN timelines like any other NAND operation but is not counted as host or GC writes. Host reads, host writes and buffer write-backs are charged; GC relocation, deallocate, Write Zeroes and the FDP write path update maptbl without a CMT access.

Log-block mapping: BAST (mapping=hybrid) and FAST (mapping=fast)​

The two log-block schemes write every page through a separate LOG write pointer and model the merges a log-block FTL has to run. Like DFTL, they use maptbl for correctness, so a read always finds the newest copy. A logical block (LBN) is pgs_per_blk consecutive LPNs.

hybrid (BAST): one log per logical block fast (FAST): one sequential log
plus a shared random-write pool
pool of 16 logs SW log: one LBN, in-order run
+--------+--------+-----+--------+ +--------------------------+
| LBN 12 | LBN 40 | ... | free | | LBN 7: off 0,1,2,... |
| used 9 | used 256 | | +--------------------------+
+--------+--------+-----+--------+ RW pool: 16 blocks of pages,
| | any LBN, dirty LBN list
| +-- full: merge +--------------------------+
| | LBN 3, LBN 90, LBN 12... |
+-- pool exhausted: merge the fullest +--------------------------+
| pool full: merge two
merge: | dirty LBNs per pass
written in order (offset i at slot i)? v
yes -> switch merge: erase the old random merge: relocate each
data block, no copies live page of those LBNs
no -> full merge: relocate each live SW log full in order:
page of the LBN to DATA space switch merge, one erase
  • hybrid follows BAST (Kim et al. 2002): one log per logical block, from a pool of 16. Each program consumes a log slot, overwritten or not; neither an overwrite nor a deallocate frees a slot, only a merge does. After every programmed page, if a log is full or every log in the pool is taken, the fullest log is merged. A log written strictly in order switch-merges for the cost of one erase. Any other log full-merges: every live page of its LBN is read and programmed into DATA-class space, counted as GC writes. Hybrid counts its merges for log page C0h. The scheme models merge cost, not BAST's physical block placement: logs of different LBNs share physical lines, and line GC still runs underneath and can add copies.
  • fast follows FAST (Lee et al. 2007): a single sequential-write log for an in-order run that starts at offset 0 of an LBN, and a shared pool of 16 blocks' worth of pages for everything else. When the pool fills (or the dirty-LBN list reaches its 4095 limit), a reclaim relocates the live pages of up to two dirty LBNs, and repeats on later writes until the pool is no longer full. A full in-order sequential log switch-merges with one erase. The sequential log is released only when it fills in order; a broken run keeps holding it, and later sequential runs go to the shared pool. FAST reclaims once per request, not once per page, and its merge counts are not exported.

Merge reads, programs and erases are charged to the request that triggered the merge, and the relocated pages count in the write amplification factor.

Choosing a scheme​

mappingModelsExtra cost chargedExtra write pointerCounters
pagea full page map in DRAMnonenonenone
dftla page map on NAND with a cache of translation pagestranslation page read on a miss, program on a dirty evictionnonenone exported
hybridBAST log blocksswitch and full merges per logLOGC0h bytes 88 to 111
fastFAST log blocksbounded random merges, sequential switch mergesLOGnone exported

Streams need page or dftl. FDP supports only page.

Lines, superblocks and write pointers​

A line is a superblock​

A line is the block with the same index on every plane of every LUN of every channel. There are blks_per_pl lines, each nchs * luns_per_ch * pls_per_lun blocks wide. A write pointer fills a line in this order: channel first, then LUN, then plane, then page.

ch0 ch1 ch2 ... ch7
lun0 pg0 [ 0 ] [ 1 ] [ 2 ] [ 7 ] numbers are the order in
lun1 pg0 [ 8 ] [ 9 ] [10 ] [15 ] which the write pointer
... hands out pages of line N
lun7 pg0 [56 ] [57 ] [58 ] [63 ] (one plane per LUN)
lun0 pg1 [64 ] [65 ] ...
...
lun7 pg255 [16383] -> line full, take the next
free line

Consecutive pages land on different channels, then different LUNs, so a large write or a burst of small ones spreads over the whole device. With several planes, the pointer sweeps every channel and LUN on plane 0, then on plane 1, and so on, before the page index moves on.

Why there are several write pointers​

Each write pointer holds one line open and appends to it. Data written through one pointer shares lines, so the lines die together or not. Keeping data of different lifetimes apart is what makes GC victims mostly invalid. The FTL has these pointers:

PointerIn struct ssdTakesOpened
datawphost writes (only first writes with hot_cold_sep=on), GC relocations, read reclaim, merge relocationsat init
hothot_wpoverwrites of mapped pages, with hot_cold_sep=onon first use
loglog_wpevery host write under hybrid or faston first use
streamstream_wp[i]writes of open stream ion first use
stream GCstream_gc_wprelocated pages of a released streamon first use
reclaim unitper handleFDP placement; see FDPFDP init

When a hot or log pointer cannot get a line, the write falls back to the data pointer. Every open pointer pins a line, so bb_check_capacity() adds one reserved line per pointer the configuration can hold open (Capacity).

Hot/cold separation​

With hot_cold_sep=on and page or DFTL mapping, prepare_write() sends a write to the hot pointer when its LPN is already mapped. An overwrite is the cheapest predictor of another overwrite, so hot lines fill with pages that are invalidated together. GC relocations go through the data pointer, so pages that survive a collection join the cold data. The log-block schemes ignore the setting, but the extra reserved line is still counted. FDP refuses it.

Line state machine​

+-------------------------------------------------+
| |
v |
+---------+ a write pointer takes it +----------+ |
init ---->| FREE |--------------------------->| OPEN | |
| (FIFO) | | (curline)| |
+---------+ +----------+ |
^ | | |
| last page | | |
| programmed, | | |
| all valid | | |
| v | |
| +--------+ | |
| | FULL | | last page programmed,
| | (list) | | some invalid
| +--------+ | |
| first invalidation | | |
| v v |
| +----------------+ |
| | VICTIM | |
| | priority queue | |
| | keyed by vpc | |
| +----------------+ |
| selected by gc_policy | |
| v |
| +-------------+ |
+-------------------------| COLLECTING | |
relocate valid pages, | in no list | |
erase every block +-------------+ |
|
read reclaim: FULL or VICTIM -> COLLECTING (reclaiming=true) ---+
Streams: a released stream's partly written OPEN line has its
free pages marked invalid and joins VICTIM (or FREE if empty)

An open line belongs to no list; invalidations only lower its valid count. When a line fills, close_time is set. The free list is a FIFO: lines are taken from the head and returned to the tail. The FTL has no explicit wear levelling; the FIFO spreads erases over the free lines.

Garbage collection​

BBSSD line states and garbage collection: when a line moves between the free, written, full and victim sets, when background and foreground GC run, and how each gc_policy picks a victim.

Figure: BBSSD line states and garbage collection: when a line moves between the free, written, full and victim sets, when background and foreground GC run, and how each gc_policy picks a victim.

The code is in hw/femu/bbssd/ftl-line-gc.c.

Triggers and watermarks​

The two watermarks are stored as free-line counts:

gc_thres_lines = (int)((1 - gc_thres_pcent/100) * tt_lines)
gc_thres_lines_high = (int)((1 - gc_thres_pcent_high/100) * tt_lines),
at least 1 with hot_cold_sep, Streams, or a hybrid
or fast mapping (bb_gc_forced_lines())

The floor matters below 20 lines, where the default high watermark rounds to zero. A second write pointer can then take the last free line while the data pointer, which GC writes into, is nearly full, and GC would have nowhere to put a victim's pages (with Streams, GC refuses such a victim outright). One free line always holds a whole victim. The floor costs those small geometries one line of exposable capacity.

KindWhenWhereVictim filterHow many
Backgroundfree lines <= gc_thres_linesafter every request that reaches the FTL without an error, any opcodethe chosen line must have at least 1/8 of its pages invalidone line
Foreground (forced)free lines <= gc_thres_lines_highat the start of a write, and before every page that a write, a write-back of the buffer, or Write Zeroes without Deallocate programsnonerepeats until above the watermark or no victim

Forced GC runs per page, not once per command: one command can program more lines than the watermark keeps free, and a collection that starts after the last line is gone has nowhere to move pages to.

A GC step returns -1 when there is no victim, when the background filter rejects it, or when the victim's valid pages cannot all be moved; the forced loop then stops. With Streams on, a step also fails when it frees nothing, because lines of distinct retired streams cannot be combined, or when the victim holds valid pages and no line is free.

Victim policies​

gc_policy selects one entry of femu_ftl_policies[]:

gc_policyChoosesCost per choice
greedythe line with the fewest valid pages: the top of the priority queueO(log n)
randoma uniformly random line of the queue (rand()); a background step puts it back if it fails the 1/8 filterO(log n)
cost-benefitthe largest age x (1 - u) / 2u, with u = vpc / pgs_per_line and age = now - close_time, compared in 128-bit integers; a line with no valid pages always winsO(n) scan
fifothe line with the oldest close_time, whatever its valid countO(n) scan
d-choicethe fewest valid pages among 4 queue slots picked from the host clock; a slot can be picked twiceO(1) sample

Only lines in the victim queue are candidates. A line that closed with every page valid stays in the full list until something invalidates one of its pages. d-choice samples by the host clock, so its picks differ between runs. random uses the process-wide, unseeded rand() stream, so it repeats only when the order of collections, and of other rand() users, repeats.

Collecting a line​

do_gc(force)
|
v
victim = policy->select_victim_line(force) -- none? return -1
|
v
move: for ch, lun, pl, each page of block (ch, lun, pl, victim->id):
if VALID:
new = next page of the data pointer (stream pointer for
(takes a free line if it has none) Streams data)
none? requeue the victim, return -1
NAND read at the old page (GC_IO)
lpn = rmap[old]
maptbl[lpn] = new, rmap[new] = lpn
old page -> INVALID, rmap[old] = none
NAND program at the new page (GC_IO)
gc_write_pages++
|
v
erase: for ch in 0..nchs-1, lun in 0..luns_per_ch-1:
mark each plane's block free: erase_cnt++, read_cnt = 0
one multi-plane erase of the LUN's blocks (GC_IO, tplebsy between
planes)
|
v
line -> free list tail

The relocated page goes wherever the data pointer is, not to the victim's LUN: the model has no copyback. Every valid page is moved before any block is erased. If the data pointer runs out of lines part way, the victim goes back to the victim queue (or the full list, if nothing was moved from a full line) holding the pages it still has, and nothing is erased. The pages that were moved are already invalid in the victim, so its counts stay true and a later step moves only what is left. Writes that then find no line fail with Capacity Exceeded. The watermark floor, per-page forced GC and the capacity reserve keep a valid configuration from getting there.

How GC time is charged​

GC operations enter the NAND model with stime = 0, which the media bridge replaces with the current host time. Each one moves its LUN's next-available time forward, and the latency it returns is discarded. GC time therefore reaches the host only through the LUN timelines: a later host read or program on a LUN that GC is using starts when GC is done with it.

LUN 3 timeline |--host W--|--GC R--|--GC W--|--GC W--|--erase-------|--host R--|
^ ^
GC starts at "now" a read that
arrived here
waits until
the erase ends
  • In background GC the request that triggered it has already been timed, so it pays nothing. The requests that follow pay, on the LUNs GC occupied.
  • In foreground GC the write itself waits, because its own programs queue behind the GC operations on the same LUNs.
  • Admin opcode 0xEF code 2 turns GC timing off (enable_gc_delay): GC still moves the pages and erases the blocks, but issues no NAND operations. Code 1 turns it back on.

The LUN field gc_endtime is updated as GC programs and erases, and the channel field is never written; nothing reads either.

FDP reclaim units​

With Flexible Data Placement the FTL places data in reclaim units, one line each, through per-handle write pointers, and collects reclaim units instead of lines (do_gc_fdp_style() in hw/femu/bbssd/ftl-fdp.c). gc_strategy picks that victim policy; gc_policy and the other FTL knobs listed under Interactions are refused. The FDP chapter and the FDP guide describe it.

Write buffer​

The BBSSD write buffer with buffer_size=10 and buffer_thres_pcent=80: a write that finds the buffer at its watermark first programs the two oldest pages and pays for them.

Figure: The BBSSD write buffer with buffer_size=10 and buffer_thres_pcent=80: a write that finds the buffer at its watermark first programs the two oldest pages and pays for them.

The write buffer models DRAM in front of the NAND. It holds logical page numbers, not data: a write it accepts costs a DRAM access, and the program is charged later to whatever forces the page out. Repeated writes to a buffered page cost one program. The code is in hw/femu/bbssd/ftl-datapath.c.

host write, pages P..Q
|
v
forced GC if free lines <= high watermark; one read-reclaim step
|
v
buffer enabled, not FUA, not a stream write? ---- no ----+
| yes |
v v
for each page: drop buffered copies of
already held? -> move to tail, write hit these pages; if the cache
else if count >= watermark: was disabled, write back
write back `batch` pages from the head everything
(oldest first), cost charged to this |
write v
insert at tail program each page
| (NAND time)
v
latency = max(DRAM access, write-back cost)

watermark = max(1, (int)(buffer_size * buffer_thres_pcent / 100))
batch = max(1, buffer_size - watermark)
DRAM access = pg_rd_lat / 16 (1000 ns when pg_rd_lat is 0)
  • Structure. A tail queue in order of last write, LRU (write_buffer) and a GLib tree keyed by LPN (wb_tree) for lookup. Capacity is buffer_size pages.
  • Admission is per page. A command larger than the buffer cannot push it past its size: it writes back a batch whenever the buffer is at the watermark. If a write-back frees nothing because the device is out of lines, the write fails with Capacity Exceeded.
  • Write-back (ssd_buffer_destage()) takes pages from the head (the least recently written), runs forced GC as needed, charges the DFTL lookup and programs each page exactly as a direct write would. Hybrid merges after each page as usual; FAST merges once per batch. A page costs the same buffered or not; only the request that pays differs.
  • Reads. A read of a buffered page costs one DRAM access and does not reach NAND, even when the buffer has stopped accepting writes.
  • FUA and stream writes program directly. A direct write first drops any buffered copy of its own pages, so an older version is never programmed after the newer one.
  • Deallocate drops buffered copies before unmapping, so discarded data is never programmed later.
  • Flush reaches the FTL like any other command and writes back the whole buffer, charged to the Flush.
  • Volatile write cache. With vwc=1 the controller advertises a cache. When the host turns it off with feature 06h, the buffer stops accepting writes and the next write drains it. With vwc=0 the buffer still works, and the host has no way to turn it off; Linux sends no Flush to a controller that advertises no cache.

host_write_pages counts pages when the host writes them and nand_write_pages when they are programmed, so a buffer that absorbs overwrites makes the write amplification factor drop below 1.

Power loss​

With power_loss=on, the poller no longer touches payload: the FTL thread runs the whole I/O command (nvme_power_io()), so the data copy and the buffer change together. The first time a page enters the buffer, the FTL saves the page's previous bytes and per-LBA state as an undo record. When the page is programmed the record is dropped. Setting the QOM property simulate-power-loss (runtime properties) resets the controller, restores every buffered page from its undo record, empties the buffer and counts an unsafe shutdown.

In this mode a normal shutdown, turning the cache off and Sanitize first write back every namespace's buffer; Flush, Format, Dataset Management, Copy, Write Zeroes (except on a PI-formatted namespace) and Write Uncorrectable write back their own namespace's. If the device has no line to write to, they fail with Capacity Exceeded (a shutdown sets Controller Fatal Status instead) rather than claim the data is durable. An FUA write programs a whole NAND page, so buffered neighbours in that page become durable with it.

Read cache​

read_cache_mb adds a DRAM read cache in front of NAND reads (hw/femu/bbssd/ftl-cache.c). Like the write buffer it holds LPNs only.

host read of a mapped lpn (not in the write buffer)
|
v
lookup in an open-addressed hash over capacity slots
hit: cost pg_rd_lat / 16, no NAND read, no block read count
miss: insert (evict by cache_evict if full), then NAND read
  • Capacity is read_cache_mb MiB divided by the page size.
  • Only reads fill it. A host overwrite, a deallocate and a power cut remove the page's entry. A GC relocation does not, since the content is the same.
  • cache_evict picks the victim: clock (second chance), random (a fixed-seed generator, so runs repeat), lru (exact, by scan) or arc, a scan-resistant 2Q variant that evicts pages read once before pages read twice. It is not full ARC.
  • Hit and miss counts are kept in ssd->rcache but not exported.

Deallocate, Write Zeroes, Format and Sanitize​

  • Dataset Management deallocate (ssd_trim()) unmaps every LPN of every range: it drops the buffered copy, invalidates the physical page, clears maptbl and rmap, and drops the read cache entry. A log-block scheme does this through its trim() hook. A range past the device is skipped. The latency is trim_lat_ns times the number of ranges, or 0.
  • Write Zeroes with Deallocate unmaps the range in the same way and costs no time. Without Deallocate it programs the range directly, after forced GC if needed, and counts the pages as host writes.
  • Format and Sanitize call bbssd_deallocate_all() on BlackBox namespaces (not CSD), which unmaps every LPN of the namespace's FTL. The lines keep their wear.

ONCS bit 0x8 (Write Zeroes) and 0x100 (Copy) are off unless oncs sets them.

Wear, read reclaim and retention refresh​

Program/erase cycles​

Every erase increments the block's erase_cnt and the FTL-wide total_erases. SMART Percentage Used is

percentage_used = min(255, total_erases * 100 / (tt_blks * rated_pe_cycles))

where rated_pe_cycles is pe_cycles_rated, or the rating of nand_cell_type (SLC 100000, MLC 3000, TLC 1000, QLC 300), or 0, in which case the field reads 0. nand_bad_blocks marks a number of blocks as factory bad for SMART Available Spare only: 100 - bad * 100 / tt_blks. Placement ignores them. With several namespaces the controller reports the most worn namespace and the lowest spare.

With ecc_step_ns set, a block's erase count (one tier per 750 erases) and the age of its line add time to each read. That model belongs to the NAND timing chapter.

Read reclaim and retention refresh​

host NAND read of block B in line L
|
+-- read_reclaim_limit set and B.read_cnt >= limit --+
+-- retention_limit_sec set and now - L.close_time >= -+--> queue L
limit (one line at
a time)
next host write:
|
v
do_read_reclaim(): queued line, high watermark not reached,
at least 2 free lines, L full or in the victim queue?
| yes
v
take L out of its list, reclaiming = true, collect L like GC
read_reclaims++ or retention_refreshes++

Every NAND read the FTL issues, host, GC or merge, increments its block's read_cnt (GC reads are issued only while GC timing is on); an erase resets it. Reads answered by the write buffer, the read cache or a DFTL translation page do not count. The rewrite happens on a write, where relocation already costs something, so a read never stalls behind a whole line. One line is refreshed per write at most. A line that nothing reads is never refreshed, and a workload that only reads refreshes nothing: FEMU runs no background media scan. The request is dropped if GC collects the line first, or if the line is still open for writing when the next write checks it.

Fault insertion​

err_read_unc_ppm and err_write_fail_ppm turn a rate into a period, 1000000 / ppm, at least 1. Every Nth Read command completes with Unrecovered Read Error and every Nth Write command with Write Fault. The counters run per namespace FTL, so a run repeats exactly. The command is still timed and, for a write, still programmed and mapped; only its status changes. The injected counts feed SMART Media and Data Integrity Errors. Write Zeroes and Copy are never failed.

Over-provisioning and capacity​

The reserve​

bb_check_capacity() in hw/femu/bbssd/bb.c refuses a namespace that would leave GC no room. The reserve, in lines:

reserve = gc_thres_lines_high forced watermark, with its floor
+ 1 data pointer
+ 1 if hot_cold_sep hot pointer
+ 1 if mapping is hybrid or fast log pointer
+ streams.max + 1 if streams stream and stream GC pointers

usable = (blks_per_pl - reserve) * pages_per_line * page_size
refused when the namespace is larger than usable,
or when blks_per_pl <= reserve

Each namespace has its own FTL built from the whole geometry, so the check applies per namespace.

Sizing the namespace​

Without op_pcent, the backend is devsz_mb MiB, split across the namespaces, and the spare area is whatever the geometry has beyond that. With op_pcent, the backend is the raw NAND capacity and each namespace gets raw * 100 / (100 + op_pcent) / namespaces, rounded down to 512 bytes (an even split unless namespace_sizes is given); devsz_mb is ignored.

Worked example: the default geometry​

The default properties describe 8 channels x 8 LUNs x 1 plane x 256 blocks x 256 pages x 8 sectors x 512 bytes:

page = 8 x 512 = 4 KiB
NAND pages = 8 x 8 x 1 x 256 x 256 = 4,194,304
raw capacity = 4,194,304 x 4 KiB = 16 GiB
lines = blks_per_pl = 256
blocks per line = 8 x 8 x 1 = 64
pages per line = 64 x 256 = 16,384 (64 MiB)

background GC = (int)(0.25 x 256) = 64 free lines (75% used)
forced GC = (int)(0.05 x 256) = 12 free lines (95% used)
victim filter = 16,384 / 8 = 2,048 invalid pages

reserve = 12 + 1 = 13 lines
usable = 243 x 64 MiB = 15,552 MiB

devsz_mb=12288 (the launcher): 12 GiB = 192 lines of data
-> the data pointer holds a line from init, so free lines reach
64 and background GC starts as the last of the 192 lines is
opened, just before one full pass completes
op_pcent=25: 16 GiB x 100 / 125 = 13,107.2 MiB exposed
op_pcent=7: 16 GiB x 100 / 107 = 15,312.1 MiB exposed

host memory for the FTL: maptbl 32 MiB + rmap 32 MiB
+ page status 16 MiB + block and line arrays

With the default devsz_mb of 1024 the same geometry exposes 1 GiB, about 6% of the NAND, and GC barely runs. To give GC work on a small device, shrink the geometry instead; this one has 512 MiB of NAND and exposes about 410 MiB:

-device femu,femu_mode=1,nchs=2,luns_per_ch=4,blks_per_pl=64,op_pcent=25

Run-time controls​

Admin opcode 0xEF changes the FTL while the guest runs. Only a BlackBox controller accepts it. Codes 1 to 4 apply to every namespace with an FTL, with the dataplane paused; codes 5 to 7 act on the controller.

CDW10Effect
1GC operations take NAND time (the default)
2GC operations take no NAND time
3flat read, program, erase times back to 40 us, 200 us, 2 ms
4flat read, program and erase times to 0
5reset the poller completion counters
6, 7turn the per-request log lines on or off

Codes 3 and 4 set the built-in flat times, not the values given on the command line, and do not affect nand_cell_type tables. See changing timing at run time.

Statistics​

Vendor log page C0h​

nvme_collect_media_stats() in hw/femu/nvme-admin.c sums the FTL counters over the controller's bbssd, CSD and KV namespaces. log-pages-and-counters.md lists the offsets.

C0h fieldFTL counterMoves when
WAF x 1000(nand + gc) x 1000 / hostafter the first host write
host write pageshost_write_pagesWrite, the destination of Copy, and Write Zeroes without Deallocate, per page, buffered or not
GC write pagesgc_write_pagesGC, read reclaim and log-block merges relocate a page
NAND write pagesnand_write_pagesa host page is programmed
max block readslargest read_cntNAND reads; reset by erase
read reclaims, retention refreshesread_reclaims, retention_refreshesa queued line is rewritten
buffer reads and read hitssp.read_cnt, sp.read_hit_cnthost read pages, and those the buffer held
buffer writes and write hitssp.write_cnt, sp.write_hit_cnthost write pages, and those already buffered
hybrid switch, full merges, merge erasesstruct femu_map_hybridmapping=hybrid only

The buffer read and write counts move with or without a buffer, so their ratio is the hit rate. Telemetry log 07h captures the same 512 bytes.

SMART and Endurance Group​

FieldSource
Percentage Usedssd_percentage_used(), most worn namespace
Available Sparessd_available_spare(), lowest namespace; 20 or below sets the spare critical warning
Media and Data Integrity Errorsinjected read and write faults (and ZNS write faults)
Data units, host commandsthe pollers' host I/O counters, not the FTL
Endurance Group Media Units Written(NAND + GC write pages) x page size, in units of 10^9 bytes rounded up; needs a femu-subsys

Counters kept but not exported​

DFTL hits and misses (ssd->cmt), read cache hits and misses (ssd->rcache), FAST merge counts and the per-block erase_cnt stay inside the FTL. debug_ftl=on prints hybrid and FAST merge counts to stdout.

Parameters​

Every property is listed with its type, default and range in the property reference; this table says what each does inside the FTL and what it interacts with.

Geometry and capacity​

NAND geometry, mode and capacity

PropertyEffect in the FTLInteracts with
secsz, secs_per_pgpage size; LPN = byte offset / page sizethe LBA format; DFTL entries per translation page
pgs_per_blkpages per block and per line column; the log-block merge unitat most 512 with nand_cell_type
blks_per_plthe number of linesthe watermarks and the reserve are fractions of it
pls_per_lun, luns_per_ch, nchsline width and write stripingparallelism; tplebsy for multi-plane erase
devsz_mb, namespaces, namespace_sizesthe exposed capacitymust fit the usable lines per namespace
op_pcentexposes a fixed fraction of raw NANDoverrides devsz_mb; refused with cxl_ssd

Garbage collection, mapping and caches​

Garbage collection, mapping and caches

PropertyEffect in the FTLInteracts with
gc_thres_pcentbackground GC watermarkmust not exceed gc_thres_pcent_high
gc_thres_pcent_highforced GC watermark; sizes the reservelower values cost exposed capacity
gc_policyline victim selectionrefused with FDP
gc_strategyreclaim unit victim selectionFDP only
mappingL2P schemehybrid and fast reserve a line and refuse Streams; FDP needs page
mapping_cache_mbDFTL cache sizeused only with mapping=dftl
read_cache_mb, cache_evictread cache size and evictionhit time is pg_rd_lat / 16 at realize; 0xEF does not change it
hot_cold_sephot write pointer for overwritespage or DFTL only; reserves a line; refused with FDP
buffer_size, buffer_thres_pcentwrite buffer capacity in pages, and the write-back watermarkvwc, power_loss; refused with FDP
fdp_trim_erase_allFDP deallocate resets every reclaim unitFDP only
debug_ftlprints invalid page transitions and merge countsnone

Reliability and wear​

Reliability and wear

PropertyEffect in the FTLInteracts with
pe_cycles_ratedPercentage Used denominatoroverrides the nand_cell_type rating
nand_bad_blocksAvailable Spareplacement ignores it
ecc_step_ns, ecc_retention_secread time grows with erase count and line ageecc_retention_sec refused with FDP
err_read_unc_ppm, err_write_fail_ppmfixed-period command failurescounted in SMART media errors
read_reclaim_limitread count that queues a line for rewriteneeds host writes to act; refused with FDP
retention_limit_secline age that queues a line for rewritesame

NAND timing the FTL charges​

NAND timing

The FTL charges trim_lat_ns per deallocate range (refused with FDP) and issues a multi-plane erase per LUN during GC, whose inter-plane time is tplebsy. All the other timing properties are applied by the NAND media layer.

Controller properties the FTL reads​

PropertyWhere documentedEffect in the FTL
vwccontroller identitylets the host turn the write buffer off
oncssameWrite Zeroes and Copy reach the FTL only when enabled
power_losspower lossundo records and simulate-power-loss
streams, streams.maxsameper-stream write pointers; reserve streams.max + 1 lines

Interactions and refusals​

  • FDP refuses buffer_size, hot_cold_sep, read_reclaim_limit, retention_limit_sec, ecc_retention_sec, trim_lat_ns, a mapping other than page and a gc_policy other than greedy, with FEMU bbssd: <name> has no effect under FDP.
  • Unknown names for mapping, gc_policy and cache_evict are refused at realize.
  • Watermarks outside [1, 100], or a high watermark below the low one, are refused.
  • The write buffer is off when power_loss=on and vwc=0, and while the host has the advertised cache turned off.
  • DFTL cache size has no effect under page, hybrid or fast.

Examples, each checked by check-doc-examples.py:

-device femu,devsz_mb=1024,femu_mode=1,mapping=dftl,mapping_cache_mb=8
-device femu,devsz_mb=1024,femu_mode=1,hot_cold_sep=on,gc_policy=cost-benefit
-device femu,devsz_mb=1024,femu_mode=1,buffer_size=4096,buffer_thres_pcent=75,vwc=1
-device femu,devsz_mb=1024,femu_mode=1,mapping=hybrid

Validation status​

What the automated tests check:

AreaTestChecks
BAST mergesunit test test-femu-hybrid-oracle; qtests hybrid-oracle-*, hybrid-batch-occupancy, hybrid-destage-occupancy, hybrid-switch-trim-erase, hybrid-trim-occupancythe unit test checks the reference model; the qtests compare FEMU's switch, full merge and erase counts against it, including deallocate and the write buffer
Victim queueunit test test-femu-pqueuepriority queue operations, including random pop
NAND timingunit test test-femu-nand-mediathe media layer the FTL calls
C0h countersqtest media-countershost and NAND page counts and the WAF move with writes
Write bufferqtest buffer-counters, flush-without-vwc and power-loss-*hit counts; Flush (also with vwc=0), FUA, write-back, cache disable, shutdown and power-cut rollback
Streamsqtest streams-gc and the other streams-* casesstream placement and GC of stream lines
GC with no free lineqtests gc-no-destination, gc-no-destination-hot-cold, streams-gc-floor, gc-no-destination-fdpon a geometry whose forced watermark rounds to zero, random single-page and 64-page writes never fail, no mapping names an erased page and no valid page is orphaned (read through the qtest-only x-ftl-check property)
Format, Sanitizeqtests format-ftl, sanitizeafter Format, GC relocates nothing; Sanitize status and zeroed data (the FTL state is not checked)
Robustnessqtests io-fuzz, io-fuzz-fdp, config-refusedmalformed I/O, refused configurations
Start-updoc-examplesevery tagged example on this page and the BlackBox guide starts and moves one block

No automated test checks the behaviour of the random, cost-benefit, fifo and d-choice policies, dftl and fast mapping, hot/cold separation, the read cache, read reclaim, retention refresh, read and write fault insertion on BlackBox, or the wear and spare figures. They are covered only by the start-up examples and by guest runs during development; femu-test.sh (guest-side tests) checks that the counters move. FEMU's latencies and write amplification have not been calibrated against a specific commercial drive.

Limits​

  • Data placement is modelled per page; the FTL never tracks sectors within a page, so writes smaller than a page cost a whole page program.
  • GC relocations go through the data pointer (or a stream pointer for stream data), never back to the victim's LUN (no copyback), and GC runs on the FTL thread, one line at a time.
  • There is no explicit wear levelling, and bad blocks do not affect placement.
  • Read reclaim and retention refresh act only on reads followed by writes.
  • d-choice GC is not reproducible run to run, and random only when everything else that calls rand() repeats too.
  • An injected write fault still programs and maps the data.
  • FAST merge counts and DFTL and read cache hit rates are not exported.
  • Each namespace's FTL is built from the whole geometry, so host memory for the FTL grows with the namespace count.

Extending the FTL​

Adding a GC victim policy​

  1. Write static struct line *select_victim_line_&lt;name&gt;(struct ssd *ssd, bool force) in hw/femu/bbssd/ftl-line-gc.c. Candidates are the lines in ssd->lm.victim_line_pq; slots 1 to size - 1 of pq->d[] hold them.
  2. Apply the background filter: when !force and the line has fewer than pgs_per_line / 8 invalid pages, return NULL and leave the queue as it was.
  3. Remove the chosen line with pqueue_remove() (or pqueue_pop() for the top), set line->pos = 0 and decrement lm->victim_line_cnt. reclaim_line() expects a line that is in no list.
  4. Add { .name = "<name>", .select_victim_line = ... } to femu_ftl_policies[]. femu_ftl_policy_known() then accepts the name.
  5. Describe it in the gc_policy description in hw/femu/femu-props.c, regenerate the property reference, and add a qtest that tells the policy apart from greedy.

Adding a mapping scheme​

  1. Implement struct femu_mapping_ops in a new ftl-map-<name>.c: at least translate, prepare_write, commit_write and gc_relocate_commit. Keep maptbl and rmap correct: commit_write must invalidate the old page and set both tables, because GC and reads depend on them.
  2. Keep private state in ssd->map_priv, allocated in init and freed in exit.
  3. Set uses_log_class if writes go through the LOG pointer; the reserve then grows by one line and Streams are refused. Set uses_cmt to get the DFTL cost model.
  4. For merges, provide needs_reclaim and reclaim(ssd, budget). reclaim() charges its own NAND operations and returns the latency to add to the triggering request. reclaim_per_page runs it after every page instead of once per request. Allocate a destination before invalidating the source, so a full device never loses the only copy.
  5. Provide trim if the scheme keeps per-page state.
  6. Add the ops to femu_mapping_extra[] in hw/femu/bbssd/ftl-map.c, declare them in ftl-internal.h, and add the file to the FEMU source list in hw/femu/meson.build.

Run the unit tests and the qtests (testing guide), and check that the C0h counters move for the new code path: a counter that never changes in a mode means the mode is not reached.

Source map​

FileContents
hw/femu/bbssd/bb.cmode registration, bb_check_capacity(), FDP refusals, the 0xEF switch, teardown
hw/femu/bbssd/ftl.cssd_init(), bb_ftl_process_req(), Copy, counters, Percentage Used, Available Spare, ssd_free()
hw/femu/bbssd/ftl.hstruct ssd, struct ppa, lines, write pointers, the mapping and policy ops
hw/femu/bbssd/ftl-internal.haddress helpers, ssd_lpn_range(), GC watermark tests
hw/femu/bbssd/ftl-geom.cbb_check_geometry(), ssd_init_params(), NAND array allocation
hw/femu/bbssd/ftl-datapath.cread, write, write buffer, deallocate, Write Zeroes, power-loss rollback
hw/femu/bbssd/ftl-line-gc.clines, write pointers, Streams pointers, GC, victim policies, read reclaim
hw/femu/bbssd/ftl-map.cmaptbl, rmap, page and DFTL schemes, scheme registry
hw/femu/bbssd/ftl-map-cmt.cDFTL cached mapping table cost
hw/femu/bbssd/ftl-map-hybrid.cBAST log-block scheme and its C0h counters
hw/femu/bbssd/ftl-map-fast.cFAST log-block scheme
hw/femu/bbssd/ftl-cache.cread cache and its eviction policies
hw/femu/bbssd/ftl-media.cbridge to the NAND media layer, block read counts
hw/femu/bbssd/ftl-fdp.cFDP reclaim units, handles and their GC
hw/femu/bbssd/ftl-exp.cdebug tracing of marked pages, off unless FEMU_EXP_LOG or FEMU_DUMP_LPN is set
hw/femu/femu.cthe FTL thread, op_pcent sizing, simulate-power-loss
hw/femu/nvme-admin.cC0h, SMART and Endurance Group counters, Sanitize, the cache feature