How the JIT works

How Warm64's Just-In-Time engine turns Arm64 code into WebAssembly, and why it is built the way it is.

Warm64's Just-In-Time engine runs guest code in two ways. Code that runs rarely is interpreted, one instruction at a time. Code that runs often is compiled to WebAssembly, which the browser's or Node's own compiler turns into native code for the host's CPU.

From blocks to modules

The interpreter runs straight-line blocks of guest instructions. When a block has run 128 times, it's queued to be compiled. Queued blocks are lowered to WebAssembly and compiled together, sixteen to a module, which spreads the host's cost of compiling a module over many blocks.

Each compiled block is a WebAssembly function that works directly on the CPU's state in the core's memory: registers and flags are fields at fixed offsets. The core installs the function in its function table and, from then on, calls it instead of interpreting the block.

Staying in compiled code

Leaving compiled code to find the next block costs more than most blocks take to run, so the compiler avoids it:

  • Linked blocks. A block that ends in a branch to another compiled block jumps straight to it with a WebAssembly tail call.
  • Regions. A hot loop, and the calls made inside it, compile as one region: a single function with the loop's structure, which runs start to finish without leaving compiled code.
  • Registers in locals. Within a block or a region, the guest registers it uses most live in WebAssembly locals, which the host's compiler keeps in its own registers.
  • Flags only when read. Arm instructions set condition flags that most code never reads. The compiler works a flag out only where something reads it, so a compare and a branch become a compare and a branch.

Memory

Every guest memory access goes through the guest's page tables. The interpreter walks them; compiled code looks the page up in a direct table instead, which maps a guest page straight to its place in the core's memory, and falls back to the full translation only when the page isn't there.

In a loop that steps through memory, the compiler goes further: it translates the addresses once for each page the loop touches, not once per access, by running the loop in two versions. A slow version finds the next page boundary; a fast version runs up to it with no checks at all.

Correctness first

Compiled code must do exactly what the interpreter does. Every lowering is tested instruction for instruction against the interpreter, and whole operating-system boots are run in lockstep with it, comparing the machines as they go. Code the compiler can't express exactly, such as some system instructions, simply stays interpreted.

When the guest rewrites its own code, as a kernel does when it patches itself, the compiled blocks for that code are dropped and compiled again from the new instructions.

What bounds it

WebAssembly is a portable machine, and the host's compiler decides the final code. The costs that remain are:

  • Block boundaries. Entering and leaving a compiled block costs a few dozen host instructions. Code with many small functions, such as recursive calls, crosses boundaries often.
  • Memory checks. An access the compiler can't prove safe still checks its page.
  • Working set. Compiled code is larger than the guest code it replaces, so code that jumps between many blocks runs out of the host's instruction cache sooner.

Even so, compute-heavy code runs at around half native speed. See Performance.

See also