Skip to main content

Debugging

Mirrored from the FEMU repository

This page is hw/femu/docs/guides/debugging.md at FEMU 39a55eeb6 (2026-10-02), licensed GPL-2.0-or-later. Send corrections to the FEMU repository.

Where FEMU reports problems, how to get more output, and how to run it under a debugger. For answers to common problems, see troubleshooting.

Where messages go​

FEMU prints its messages on QEMU's console:

PrefixStreamMeaning
[FEMU] Log:stdoutinformation, such as a vendor command taking effect
[FEMU] Err:stderran error FEMU recovered from, such as a CSD program that failed to load
[FEMU] FTL-Log:, [FEMU] FTL-Err:stdout, stderrthe BlackBox FTL
[Misao] ZFTL-Log:, [Misao] ZFTL-Err:stdout, stderrthe ZNS FTL
[FEMU] FDP-Log:, [FEMU] FDP-Trace:stderrFDP setup; placement and reclaim traces with FEMU_FDP_DEBUG
qemu-system-x86_64: -device femu,...: MESSAGEstderrthe device refused its options and QEMU did not start

run-blackbox.sh, run-zns.sh and run-csd.sh copy the console to build-femu/log, and run-blackbox-fdp.sh to /tmp/femu-fdp.log. run-nossd.sh, run-whitebox.sh and run-cxlssd.sh print only to the terminal. Through the tee pipe, standard output is buffered, so stdout lines can reach the log later than stderr lines. Put stdbuf -oL in front of ./qemu-system-x86_64 on the launcher's sudo line to write them at once.

Check the guest's side too. A guest that gives up on the device logs it in dmesg:

sudo dmesg | grep -i nvme

FEMU defines no QEMU trace events and does not use QEMU's -d log categories, so -d and -trace show nothing from FEMU itself. They still help with the QEMU code around it. -d guest_errors,unimp -D qemu.log logs bad register accesses that QEMU's MSI-X and CXL code detect. Add them to the QEMU command line in the launcher:

-d guest_errors,unimp -D qemu.log

Run FEMU under gdb​

Do not use gdb-run.sh; it starts QEMU's stock nvme device, not FEMU. Instead, run a launcher's own command line under gdb. From build-femu/, make a copy of the launcher that starts QEMU through gdb, and run it:

sed -e 's|\./qemu-system-x86_64|gdb -ex "handle SIGUSR1 nostop noprint pass" --args ./qemu-system-x86_64|' \
-e 's/ 2>&1 | tee .*$//' run-blackbox.sh > gdb-blackbox.sh
bash gdb-blackbox.sh

At the gdb prompt, set breakpoints and start QEMU:

(gdb) break femu_realize
(gdb) run

KVM uses SIGUSR1 to kick vCPU threads, which is why gdb is told to pass it on. The guest's serial console shares the terminal with gdb; press Ctrl-C to get back to the gdb prompt.

To attach to a QEMU that is already running instead, on the host:

sudo gdb -p "$(pgrep -x qemu-system-x86)" -ex "handle SIGUSR1 nostop noprint pass"

When QEMU crashes, thread apply all bt in gdb prints every thread's stack. The launchers pass -name NAME,debug-threads=on, so info threads shows FEMU's thread names (femu-poller, FEMU-FTL-Thread, femu-csd-cu, femu-cxl-ftl, femu-cxl-cca); add it to a QEMU command line of your own (performance tuning).

Debug builds and compile-time switches​

femu-compile.sh builds with -O2 -g. gdb works on that build, but many variables are optimized out. To single-step, configure an unoptimized build yourself from build-femu/:

../configure --enable-kvm --target-list=x86_64-softmmu --enable-slirp \
--disable-libnfs --disable-libiscsi --disable-curl \
--enable-debug --enable-debug-info
make -j"$(nproc)"

The debug messages (Dbg:) have no run-time switch. These macros, passed with --extra-cflags, compile them in:

MacroEffect
FEMU_DEBUG_NVMEcontroller debug messages ([FEMU] Dbg:), and the per-command log that vendor command 0xEF with CDW10 6 and 7 turns on and off
FEMU_DEBUG_FTLBlackBox FTL debug messages ([FEMU] FTL-Dbg:), and the FTL invariant checks of BlackBox and ZNS
FEMU_DEBUG_ZFTLZNS FTL debug messages ([Misao] ZFTL-Dbg:)
FEMU_FTL_ASSERTthe FTL invariant checks only, without the messages

For example:

../configure --enable-kvm --target-list=x86_64-softmmu --enable-slirp \
--disable-libnfs --disable-libiscsi --disable-curl \
--extra-cflags="-DFEMU_DEBUG_FTL -DFEMU_DEBUG_NVME"
make -j"$(nproc)"

The invariant checks (ftl_assert) are compiled out of a normal build. A mapping or bound error there corrupts FTL state quietly instead of stopping QEMU. When a BlackBox or ZNS run produces wrong data or impossible counters, rebuild with -DFEMU_FTL_ASSERT first. The debug messages print on the I/O path, so use them with small workloads.

The sanitizer build adds AddressSanitizer and UndefinedBehaviorSanitizer on top of the FTL checks. Run the failing workload on it to find memory errors.

Run-time switches​

FEMU reads a few environment variables (environment variables):

VariablePrints to stderr
FEMU_FDP_DEBUG (any value)FDP placement and reclaim traces
FEMU_EXP_LOG=1 with FEMU_SECRET=TEXTwrites, overwrites, deallocations, GC moves and erases of pages whose data contains TEXT
FEMU_DUMP_LPN=Na hex dump of logical page N on every read not served from the write buffer
FEMU_KV_SELFTEST (any value)the result of a KV FTL self-test run once at start

The launchers other than run-cxlssd.sh start QEMU with sudo, which drops your environment. run-blackbox.sh passes FEMU_EXP_LOG, FEMU_SECRET and FEMU_DUMP_LPN through, so this works:

FEMU_EXP_LOG=1 FEMU_SECRET=MARKER ./run-blackbox.sh

For the others, add the variable to the launcher's sudo line, for example sudo FEMU_FDP_DEBUG=1 ./qemu-system-x86_64 \ in run-blackbox-fdp.sh.

On a running BlackBox controller, the vendor admin command 0xEF switches GC time and NAND time on and off and prints poller counts (timing model).

Inspect a running device​

QEMU's monitor shows what QEMU built. With the QMP socket the launchers other than run-cxlssd.sh create (build-femu/qmp-sock), on the host:

printf '%s\n' '{"execute":"qmp_capabilities"}' '{"execute":"query-pci"}' |
sudo socat - UNIX-CONNECT:./qmp-sock

qom-list and qom-get on /machine/peripheral/<id> read a device's properties and counters (runtime properties). run-nossd.sh names its controller nvme0 and run-cxlssd.sh names its device cxlssd. The other launchers give the controller no id=, so add one to address it by name.

Common crash reports​

SymptomWhat it meansWhat to do
qemu-system-x86_64: -device femu,...: MESSAGE and QEMU exitsA property value was refused at startRead the message; the mode guide's "Limits and refusals" table explains it
[FEMU] Err: cannot pin the N MiB memory backendNot a crash. QEMU runs without sudo and the memory lock limit is too lowIgnore it, or raise ulimit -l (requirements)
failed to allocate N bytes, then QEMU stops with a trap or abort and may dump coreThe host has less free memory than the device needsLower devsz_mb or free memory (host sizing)
Guest dmesg shows nvme nvme0: I/O ... timeout, then a controller resetThe guest's NVMe driver waited too long for a completion and reset the controllerCheck the QEMU console for an error at the same time. If there is none, check that the pollers have cores (performance tuning). If it repeats on an idle host, report it
The device is missing in the guest, and the console shows [FEMU] Err: linesFEMU refused a command or a configuration the guest usedRead the console; for ZNS, see the ZNS troubleshooting
QEMU exits with a segmentation fault or an assertionA FEMU bugRun it under gdb or the sanitizer build, and report it with the steps below

Older reports of crashes in FDP garbage collection (#186, #189, #191) and of controllers disabled under load (#22, #30) were made against older FEMU versions. Reproduce on current master before you report one.

Reporting a bug​

Open an issue with:

  • the FEMU commit (git log -1 --oneline) and how you built it;
  • the full QEMU command line, or the launcher and every change you made to it;
  • the guest kernel (uname -r) and the commands you ran in the guest;
  • the QEMU console output and the guest's dmesg;
  • for a crash, the thread apply all bt output from gdb, or the sanitizer report.

Report a security problem privately, as SECURITY.md says.

Related issues: #18, #22, #30, #105; discussion #133.