Skip to main content

Namespace management, metadata and protection information

Mirrored from the FEMU repository

This page is hw/femu/docs/features/ns-management-and-pi.md at FEMU 39a55eeb6 (2026-10-02), licensed GPL-2.0-or-later. Send corrections to the FEMU repository.

This page covers two optional NVMe features of NoSSD and BlackBox controllers:

Both are off by default. Use them to test host software that manages namespaces, or that writes and checks per-block metadata and PI, without hardware that supports them.

Requirements​

  • Host and guest: see the mode table and requirements.md.
  • Namespace management: any guest kernel with the NVMe driver, and nvme-cli (create-ns, attach-ns).
  • Metadata through the block layer: a guest kernel with CONFIG_BLK_DEV_INTEGRITY=y. Without it, use passthrough commands.

Namespace management​

What it does​

With ns_mgmt=on, the controller advertises Namespace Management and Namespace Attachment and reports 256 as its namespace count. The namespaces from namespaces and namespace_sizes are created at boot, and together they define the capacity pool: the pool is the sum of their sizes, and at boot all of it is in use (unvmcap is 0). To create a namespace, the guest first deletes one, then allocates from the space it freed. A namespace the guest creates is detached until the guest attaches it.

It takes effect only when the controller and every boot namespace run the same mode, NoSSD or BlackBox. With another mode it is accepted and stays off.

Launch​

A BlackBox controller with two 2 GiB boot namespaces, a 4 GiB pool:

-device femu,devsz_mb=4096,femu_mode=1,ns_mgmt=on,namespaces=2

A NoSSD controller:

-device femu,devsz_mb=4096,femu_mode=2,ns_mgmt=on,namespaces=2

To share namespaces between controllers, set ns_mgmt=on on a femu-subsys and join it with subsys=. The subsystem comes first:

-device femu-subsys,id=shared,ns_mgmt=on -device femu,id=ctrl-a,subsys=shared,femu_mode=2,devsz_mb=64 -device femu,id=ctrl-b,subsys=shared,femu_mode=2

The first controller sets up the pool and the boot namespaces; later ones join with nothing attached. Shared namespaces in detail describes the shared model in full.

Configuration​

Properties: namespace management, streams and power loss, shared namespaces.

  • ns_mgmt: on a standalone controller, or on femu-subsys for shared namespaces. A controller with ns_mgmt=on cannot join a subsystem that does not have it.
  • bbssd_ns_limit (BlackBox only, default 4): the most namespaces that may be allocated at once, detached ones included. Each BlackBox namespace has its own FTL with the full NAND geometry, so this bounds host memory.
  • namespaces and namespace_sizes: the boot namespaces, whose sizes add up to the pool. Space that no boot namespace covers is not in the pool.

A shared subsystem takes NoSSD or BlackBox controllers that all have the same mode, meta, mc, pi, dpc, nlbaf, vwc and oncs, and no Streams, dps or namespace_modes.

Use it from the guest​

Check that the controller offers it, and find its controller ID:

sudo nvme id-ctrl /dev/nvme0 | grep -E '^(oacs|nn|cntlid|tnvmcap|unvmcap) '

oacs has bit 3 (0x8) set, nn is 256, and unvmcap is the free capacity in bytes.

With the BlackBox example above, delete boot namespace 2 to free 2 GiB, create a 1 GiB namespace of 512-byte blocks (LBA format 0) in it, attach it to this controller, use it, then remove it:

sudo nvme delete-ns /dev/nvme0 --namespace-id=2
sudo nvme create-ns /dev/nvme0 --nsze=2097152 --ncap=2097152 --flbas=0
sudo nvme attach-ns /dev/nvme0 --namespace-id=2 --controllers=0
sudo nvme ns-rescan /dev/nvme0
sudo nvme list
sudo nvme detach-ns /dev/nvme0 --namespace-id=2 --controllers=0
sudo nvme delete-ns /dev/nvme0 --namespace-id=2

create-ns prints the new namespace ID; use it in the commands after it. --controllers takes the cntlid from id-ctrl. nsze and ncap are in logical blocks and must be equal: an ncap below nsze fails with Thin Provisioning Not Supported, and a zero or larger one with Invalid Field. A BlackBox namespace is rounded up to whole NAND pages.

Limits​

Message or statusCause and fix
bbssd_ns_limit must be between 1 and 256 and cover the boot namespacesnamespaces above bbssd_ns_limit, or the limit out of range.
ns_mgmt=on does not support subsys; use a standalone controllerns_mgmt=on on a controller whose subsystem lacks it. Set it on the femu-subsys instead.
shared namespaces require homogeneous NoSSD or bbssd without Streams or default protectionA ZNS, KV, CSD or OCSSD controller, or Streams, dps or namespace_modes, in a shared subsystem.
shared namespace mode and capabilities must matchControllers in one shared subsystem differ in mode or format properties.
namespace management does not support FDPns_mgmt=on and fdp=on on one subsystem.
Namespace Insufficient CapacityThe pool has no free space (delete a namespace first), or a BlackBox namespace would not leave GC its free lines.
Namespace Identifier UnavailableAll 256 IDs are used, or bbssd_ns_limit is reached.
Thin Provisioning Not Supportedncap below nsze.
I/O Command Set Not SupportedA command set other than NVM; the guest can create only block namespaces.

Shared namespaces in detail​

With femu-subsys,ns_mgmt=on, the subsystem owns one namespace table, backend capacity pool and full-geometry FTL per bbssd namespace. This prevents controllers allocating the same NSID independently or deleting only a private copy. Realize the subsystem before its controllers. The first controller establishes the pool and boot namespaces; later controllers join that pool with no initial attachments. Namespace Attachment selects their active namespaces.

The subsystem copies the first controller's properties into an unrealized configuration object as its storage context, independently of PCI transport lifetime. This context has no queues or workers. Its namespace table and backend survive removal of that controller, including removal of every controller. They are released at subsystem teardown after its last controller has left. Namespace identity and the creation sequence belong to this context. Controllers must agree on the storage mode and format capabilities. Runtime construction uses the original storage configuration, including capacity, geometry, media options and the bbssd namespace cap; those properties on later controllers do not create additional storage.

Each controller keeps its own queues, pollers, FTL thread, attachment bitmap and namespace-change log. Pollers reach the common table through their controller; active lookup additionally checks that controller's attachment bitmap. A subsystem mutex serializes data and FTL processing. Existing pause/resume interfaces pause all controllers in a managed subsystem, so Format, Sanitize and lifecycle changes cannot race another controller's workers. Delete retires namespace references in every controller's request containers before releasing storage, without waiting for guest consumption of a full CQ. Detach retires only the selected controllers. Reset preserves attachments. Transport removal stops its workers and drops its attachments without releasing the subsystem's namespaces.

CNS 10h/11h describe the common allocated set. CNS 02h describes the issuing controller's attachments. CNS 12h lists attached controllers and CNS 13h lists eligible controllers, in ascending order with inclusive CNTID filtering. Attach/detach records changes and sends enabled Attached Namespace Attribute Changed notices on each affected controller. Delete does the same except that its issuer receives no notice. Allocated Namespace Attribute notices are not advertised. Create produces a detached namespace and no attached-set notice.

The model is opt-in on the subsystem. Existing standalone management and unmanaged subsystems retain their behavior. Only homogeneous NoSSD or bbssd is supported; FDP and Streams remain excluded. Full-geometry bbssd allocation and its namespace-count cap are retained. Scaled FTL geometry, allocated-namespace notices, persistent storage and other namespace modes are not supported.

With the shared example under Launch, only ctrl-a initially has the boot namespace attached. Controller lists report the assigned CNTLIDs; use Namespace Attachment to expose it through ctrl-b too. Create honors NMIC: private namespaces can move between controllers but cannot be attached to two simultaneously. Shared namespaces permit both attachments.

Metadata and protection information​

What it does​

meta=<bytes> adds that many metadata bytes to every logical block. The controller then offers each LBA format twice: formats 0 to nlbaf - 1 without metadata, and formats nlbaf to 2 * nlbaf - 1 with it. The namespace boots on the metadata version of lba_index. With the default nlbaf=5 and lba_index=0, that is format 5: 512-byte blocks with metadata.

The metadata travels either in a separate buffer (mc bit 1, 0x2) or interleaved with the data as extended LBAs (mc bit 0, 0x1). extended=1 boots on the interleaved layout.

With pi=on and meta of 8 or more, the controller also offers protection information types 1, 2 and 3, in the first or the last 8 bytes of the metadata. The guest selects a type with Format NVM, or with --dps when it creates a namespace. The controller generates and checks PI as the command asks (PRACT and PRCHK bits).

Launch​

8 bytes of separate metadata with PI available:

-device femu,devsz_mb=1024,femu_mode=1,meta=8,mc=2,pi=on

8 bytes of interleaved metadata, booting on the extended layout:

-device femu,devsz_mb=1024,femu_mode=1,meta=8,mc=1,extended=1

Configuration​

Properties: LBA formats, metadata and protection.

  • meta: metadata bytes per block. NoSSD and BlackBox namespaces only.
  • mc: which layouts the controller supports, separate (0x2), interleaved (0x1) or both (0x3). Required with meta.
  • extended: boot on the interleaved layout. Needs bit 0 of mc.
  • pi: offer PI types 1 to 3. Needs meta of 8 or more; with less it is accepted and offers nothing.
  • nlbaf: at most 8 with meta, since each format is offered twice.
  • dpc and dps describe PI without metadata support; leave them at 0 and use pi.

Use it from the guest​

Look at the formats and the protection capabilities:

sudo nvme id-ns /dev/nvme0n1 -H | grep -E 'LBA Format|dpc|dps|mc'

Reformat to 512-byte blocks with 8 bytes of separate metadata and PI type 1 in the last 8 bytes (--pil=0), then check that the namespace changed:

sudo nvme format /dev/nvme0n1 --lbaf=5 --pi=1 --pil=0 --ms=0 --force
sudo nvme id-ns /dev/nvme0n1 | grep -E '^(flbas|dps) '

Write and read one block and let the controller generate and check the PI (--prinfo=0xf: PRACT plus the guard, application tag and reference tag checks; the reference tag of LBA 0 is 0):

head -c 512 /dev/urandom > blk.bin
sudo nvme write /dev/nvme0n1 -s 0 -c 0 -z 512 -d blk.bin --prinfo=0xf --ref-tag=0
sudo nvme read /dev/nvme0n1 -s 0 -c 0 -z 512 -d out.bin --prinfo=0xf --ref-tag=0
cmp blk.bin out.bin

With CONFIG_BLK_DEV_INTEGRITY=y, Linux generates and verifies PI for ordinary block I/O on such a namespace by itself.

Limits​

MessageCause and fix
meta/extended need a matching metadata capability (mc)mc must have bit 0 when extended=1, and bit 1 otherwise.
meta: at most 8 block sizes (nlbaf), each is also offered with metadataLower nlbaf.
meta: namespace N runs a mode without metadata support (block or no-SSD only)Every namespace must be NoSSD or BlackBox.
meta: not supported with placement (fdp)Metadata and FDP do not combine.
meta: protection information (dpc, dps) is not supporteddpc or dps with meta; use pi.
dps needs 8 bytes of metadata and a matching dpcUse pi=on instead of dps.
power_loss requires bbssd, buffer_size > 0, and no metadata, namespace management or subsystemmeta or pi with power_loss.

At run time, a Format NVM that asks for a PI type when pi is off, a PI type with less than 8 bytes of metadata, or a layout mc does not allow fails with Invalid Format. On a type 1 namespace, a command whose reference tag is not the low 32 bits of its LBA fails with Invalid Protection Information. A check against stored PI that fails returns the matching End-to-end Guard, Application Tag or Reference Tag Check Error.

Verify​

  1. Namespace management: oacs has bit 3 set, and a created and attached namespace appears in nvme list with the size you asked for.
  2. Metadata: nvme id-ns /dev/nvme0n1 shows a non-zero ms for the current format, and dpc is 0x1f with pi=on.
  3. PI: the write and read above succeed, and a read with --ref-tag=1 for LBA 0 fails with Invalid Protection Information.

Troubleshooting​

  • ns_mgmt=on but oacs has no Namespace Management bit. The controller or a boot namespace is not NoSSD or BlackBox, or the modes differ. The property is silently off in that case.
  • create-ns fails with Insufficient Capacity. The pool is the boot namespaces, and they use all of it at boot. Delete a namespace first.
  • A new namespace does not appear. It is detached after creation. Attach it with the controller ID from id-ctrl, then run nvme ns-rescan.
  • The namespace shows a size of 0 after formatting with metadata. The guest kernel cannot use that format through the block layer, usually because it lacks CONFIG_BLK_DEV_INTEGRITY or the format interleaves the metadata. Use passthrough commands on /dev/ng0n1, or a separate-buffer format on a kernel with integrity support.

Related issues: #121.