# WebAssembly
**Run WebAssembly plugins inside your program, with a millisecond start and
nothing to install.**
- **Fastest start measured.** The first call runs 1.0 ms after launch; a JIT
pays 2 to 20 ms before it runs anything.
- **Cheap calls.** A module loads in about ten microseconds, and a call into it
costs around a hundred nanoseconds.
- **Small and portable.** wasm3 is 550 KB of portable C with no executable
pages to allocate at runtime.
- **Bounded guests.** A gas budget or `interrupt` stops a module that would
otherwise run forever.
- **WASI commands run as-is** through `run`.
| At a glance | |
|---|---|
| Version | wasm3 **0.9.1** (commit `c036c43`) |
| Licence | MIT, copy in `wasm/vendor/LICENSE` |
| Links | Statically, from `wasm/lib/wasm3.a` |
| Builds with | `just wasm` |
| Threads | One thread for the whole package |
## Quick start
```odin
vm := must(wasm.open())
defer wasm.close(&vm)
mod := must(wasm.load(vm, transmute([]byte)must(path.read("add.wasm"))))
add := must(wasm.find(mod, "add"))
out := must(wasm.call(add, i32(19), i32(23)))
fmt.println(out[0].(i32)) // 42
```
> [!WARNING]
> wasm3 is not thread-safe, even with one runtime per thread. Know this before
> reaching for `jm:flow`: keep `jm:wasm` on one thread, and `interrupt` is the
> only call meant to come from another.
Under the hood: the threading measurement
The limit is not merely per runtime. Eight threads, each with its own
environment, runtime and copy of a two-instruction module, still produce
spurious traps — a stack overflow, an out of bounds access — roughly five
times in sixteen hundred calls. That was reproduced in C against this archive,
with no Odin involved, so it is wasm3's own state and not the binding's.
`jm:wasm` is therefore a one-thread package. `interrupt` crosses threads
safely because it does nothing but set a flag. `just test` runs this package's
tests with the test runner on a single thread for the same reason.
## Speed
An interpreter trades raw speed for size and start time. The table shows
milliseconds of work for four modules, fastest of three runs, lowest first.
*First run* is what the whole process costs at `run(1)`.
| engine | fib | mandel | memsum | sort | first run |
|---|---:|---:|---:|---:|---:|
| wasmtime 49.0.1, Cranelift JIT | 84 | 178 | 52 | 158 | 2.9ms |
| Node 26.8.1, V8 JIT | 93 | 316 | 80 | 309 | 15.6ms |
| wazero 1.12.0, Go compiler | 125 | 193 | 130 | 380 | 2.0ms |
| WAMR 2.4.5, default (JIT) | 160 | 370 | 71 | 301 | 18.7ms |
| **jm:wasm** (wasm3 0.9.1) | 820 | 897 | 1200 | 1140 | 1.0ms |
| wasm3 0.9.1, its own CLI | 920 | 937 | 1367 | 1236 | 1.1ms |
| WAMR 2.4.5, `--interp` | 2643 | 6528 | 2806 | 6437 | 20.2ms |
What to take from it:
- **The bindings cost nothing.** wasm3's own CLI comes out a few percent
*slower* on the same modules, so the table measures the interpreter, not the
crossing.
- **A JIT is faster at work.** The interpreter is five to ten times slower on
arithmetic and twenty times on memory traffic. That is the price of the
550 KB of portable C and the one millisecond to first execution.
- **Among interpreters, wasm3 holds up.** WAMR's classic interpreter, the other
embedded standard, is three to seven times slower on the same modules. WAMR's
default mode is not an interpreter at all, which is why it looks fast in the
fourth row.
The per-call costs, from `just bench` on this machine:
| module | load | first | call | work |
|---|---:|---:|---:|---:|
| fib | 14.5µs | 2.8µs | 138ns | 219µs per unit |
| mandel | 10.5µs | 7.4µs | 118ns | 481µs |
| memsum | 10.8µs | 5.3µs | 106ns | 2.3ms |
| sortbench | 9.9µs | 7.9µs | 102ns | 9.5ms |
Under the hood: how the benchmarks were run
`just bench` times jm:wasm against the four workloads in
`tools/wasm-bench/workloads`, compiled from the C beside them by
`just bench-build`. Each exports `run(i32) -> i32`: the argument scales the
work linearly and the result is a checksum, so the same module can be run
through another engine and checked.
The harness reports the load, the first call — which is where wasm3 compiles
the body — an empty call, and the workload. It subtracts `run(2n) - run(n)`
so that only what scales with the work is left.
The engine comparison runs the same four modules at a fixed size — `run(4096)`,
`run(2048)`, `run(512)` and `run(128)` — through every engine that could be
made to run them. Each run is one command that loads the module, calls `run`
once and exits.
Every engine returned the same checksum for every workload, which is what
makes the columns comparable. The figures are the fastest of three runs on an
otherwise idle machine and repeat to within about fifteen percent. The orders
of magnitude are the point; the last digit is not.
"The bindings cost nothing" is the same measurement against a C program
driving `wasm/lib/wasm3.a` directly: 800ms against 923 on fib, 1166 against
1153 on memsum. The two are inside each other's noise.
Under the hood: provenance and build options
`wasm/vendor/` holds **wasm3 0.9.1** (commit `c036c43`): the interpreter's
`source/` tree minus the two files this build does not compile. Those are the
uvwasi backend, which needs libuv, and the meta-WASI one, which is for running
wasm3 inside wasm3. wasm3 is MIT, and `wasm/vendor/LICENSE` is its copy.
`just wasm` compiles it once into `wasm/lib/wasm3.a`, which is gitignored and
rebuilt when any vendored source changes. As with SQLite, `foreign import`
resolves that archive relative to the package directory, and `odin check`
never opens it, so `just check` still type-checks all three targets on one
machine with no archive built.
The build takes wasm3's defaults, which already have bytecode validation and
gas metering on, and adds one option: `d_m3HasWASI`, without which `run` has
no `_start` to call.
## Tested by fuzzing
`wasm/fuzz` is the suite: six properties over generated modules, damaged ones
and bytes that were never a module. A small Wasm encoder builds each case, so
no toolchain is needed. `just fuzz "wasm -for=1m"` runs it, and
[Fuzzing](fuzzing.md#what-it-found) records what it found.
## See also
- [Fuzzing](fuzzing.md): the property suites and what they found
- [Packages](packages.md): every package in jm
- [SQLite](sqlite.md): the opposite threading model, a connection per worker