Skip to content

fix(wasm): reserve enough analysis space for dense exports - #134

Open
BridgeAR wants to merge 3 commits into
mainfrom
BridgeAR/2026-10-06-cjs-wasm-memory
Open

BridgeAR wants to merge 3 commits into
mainfrom
BridgeAR/2026-10-06-cjs-wasm-memory

Conversation

@BridgeAR

@BridgeAR BridgeAR commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Dense shorthand exports can exhaust analysis capacity because a 12-byte record can occur every two code units. Grow Wasm memory before writing beyond capacity.
A 296-file Mocha, Babel, and Terser corpus retains 2.19 MiB of Wasm linear memory.

@guybedford guybedford left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, the analysis is right: Slice is 12 bytes and shorthand literal exports (a,) can produce one every two code units, so 8 * len is exactly the worst-case bound, and the tests reproduce the trap on main.

Rather than scaling the fixed reservation though, I'd prefer to bound-check in the allocator and grow on demand, which is what es-module-lexer does now: sa() snapshots analysis_limit = __builtin_wasm_memory_size(0) << 16, and each _add* calls an inline ensureAnalysisCapacity(sizeof(Slice)) that falls into a noinline growAnalysis doing __builtin_wasm_memory_grow when analysis_head + size > analysis_limit. The JS side can then keep len * 4 (or lower) as a floor. Wasm memory never shrinks, so with 8 * len every parse of a large bundle permanently pins 8x its size for a case that almost never occurs; growing on demand keeps retained memory proportional to actual results, at the cost of one compare per recorded export. The JS wrapper only touches wasm.memory.buffer before parseCJS, so in-wasm growth detaching the buffer is safe.

This needs a wasm rebuild, so happy to help with that if useful. The tests here are good as the regression either way.

Comment thread src/lexer.js Outdated
Dense object exports can trap in WebAssembly because the host reserves
four bytes per UTF-16 code unit, below the result-list requirement. A
result node needs 12 bytes and can occur every two code units. Reserve
eight bytes per unit, including the source. This doubles the input-based
capacity allowance. Application measurements show no consistent CPU
regression.

Signed-off-by: Ruben Bridgewater <ruben.bridgewater@datadoghq.com>
Dense shorthand exports produce a 12-byte record for every two code units,
so a source-sized analysis allowance can overflow. Reserving for that
worst case retains excess memory on sparse inputs.

A 1,048,596-code-unit sparse input retains 4,259,840 bytes instead of
8,454,144 bytes. Dense shorthand parses cost 2.18-3.99% more time (Node
22.23.3, V8 12.4.254.21-node.57, seven alternating trials, trimmed means,
two fresh processes).
@BridgeAR
BridgeAR force-pushed the BridgeAR/2026-10-06-cjs-wasm-memory branch from f45d1a0 to 07b764e Compare October 11, 2026 09:33
@BridgeAR
BridgeAR marked this pull request as ready for review October 11, 2026 09:48
The proportional analysis allowance wastes linear memory on ordinary modules with few exports.
Parsing a 296-file Mocha, Babel, and Terser corpus retains 2.19 MiB of Wasm linear memory instead of 4.25 MiB.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants