Performance
What makes Warm64 fast, what to expect, and the few choices that change it.
What to expect
Measured on an Apple M3:
| Just-In-Time | Interpreter | |
|---|---|---|
| Hot loops | 1,000 to several thousand million instructions a second | about 30 million |
| md5sum, recursive calls, sorting | about half native speed | |
| A systemd boot | 150 to 200 million instructions a second, on average |
Operating systems, measured with the Just-In-Time engine:
| System | To a shell or login |
|---|---|
| Alpine Linux (netboot kernel and initramfs) | a few seconds |
| FreeBSD 15.1 | about 40 seconds |
| Debian 13 | about 100 seconds |
| Arch Linux ARM | about 2 minutes |
Choices that matter
Keep the default CPU model
-cpu max advertises pointer authentication, MTE and SVE, and a kernel that uses them boots about 14% slower. Use the default cortex-a76 unless the guest needs those features.
Stay at or under 3 GiB on one core
Up to 3 GiB, Warm64 uses its 32-bit core. Above it, a single-core machine switches to the 64-bit core, which runs about a third slower: every memory address is 64 bits wide, in the engine and in the compiled code. If the guest doesn't need the memory, don't give it.
Run it in a worker, and keep the tab visible
In a browser, a worker compiles the JIT's code at once where a page's main thread can't, and a hidden tab gets a fraction of the CPU. See Embedding in a web page.
Kernel parameters don't matter
mitigations=off, nokaslr and the like make no measurable difference.
Counters
x-accel-stats (QMP) and info jit (the monitor) show how much ran in compiled code:
console.log(await vm.qmp("x-accel-stats"));
// { "human-readable-text": "TCG: guest instructions N, compiled blocks entered N, instructions interpreted N" }