Skip to content

Repository files navigation

Industrial Workers of the World logo

wobbly

Parallelize map, filter, reduce, and forEach over large arrays using Web Workers.

The logo is the Industrial Workers of the World logo — the IWW are also known as "the wobblies".

bun add wobbly-js   # or: npm install wobbly-js

Before you use this, read SHOULD-YOU-USE-A-WORKER.md. It is written to talk you out of it. Most work that looks like it needs threads needs a better algorithm instead — the project this was built for got a 4.3× win by fixing their maths and then didn't need a worker at all. And the membrane means objects are usually a loss. wobbly wins on numeric data, generators, and genuinely heavy per-item work. Measured, with the losses shown.

The four data paths

How your data reaches a worker decides almost everything. Cost to hand over 1M items:

Your data Mechanism Cost Needs
Array of objects structured clone 968ms 🐌 nothing
Array of numbers structured clone 23ms nothing
TypedArray transfer (pointer move) ~1.6ms nothing
SharedArrayBuffer-backed nothing — in place 0ms COOP/COEP headers

Objects cost about 1µs each to clone, which is ~600× a numeric array of the same length. That is not a tax you can optimise away; it is what postMessage costs for a rich object graph.

So: wobbly is at its best on numeric data, and at its worst on big arrays of objects. With objects, a 10-worker fan-out only pays off if your callback takes more than roughly 1µs per item — which real work often does (parsing, regex, crypto, geometry), but a cheap field transform never will. If you're moving a large object graph just to do something trivial to it, you will lose.

The best workload: send the recipe, not the data

The table above says the membrane is what costs you. The way to win, then, is to not send data at all.

A deterministic generator — procedural terrain, a noise field, mesh or texture generation, a simulation rolled forward from a seed — has a tiny input by construction. You send a recipe, and the worker generates the data on the other side:

// in:  a few numbers.  out: a 15KB Float32Array, transferred back.
const tiles = await new AsyncArray(recipes).map(function buildTile(r) {
  const out = new Float32Array(r.verts * 3)
  // …generate from r.seed / r.cx / r.cz…
  return out
})

Nothing is copied in — the input is tens of bytes. Nothing is copied out — the result is a TypedArray and gets transferred. The membrane cost is zero, and it stays zero however big the output is.

If you're tempted to ship a lookup table (a seeded permutation table, say) as context: don't. Send the seed. It's a pure function of the seed, so each worker can rebuild it once, in microseconds. Shipping the table is shipping a cache of something cheaper to recompute than to transmit.

This class of work is where a worker pays for itself unambiguously — much more reliably than "heavy callback over a big array," which, as the benchmarks below show, is a narrower win than it sounds.

For this workload, throughput is the wrong metric

If you're generating frames, nobody perceives your average. They perceive the one frame that took 60ms. So the number that matters is not tiles/sec — it's how long the main thread was blocked, and how soon the first needed result arrived.

Generating 24 terrain tiles (10-core Mac, onPartial to take results as they land):

serial wobbly
worst main-thread block 10.09ms — a dropped frame 0.56ms
time to first usable tile 10.09ms (all or nothing) 0.39ms
wall clock, all 24 tiles 9.2ms 2.45ms

The bottom row is the throughput win, and it's the least interesting one. The top row is the product: the burst stops blocking your render loop. Note also that serial has no "first tile" — you get everything at the end or nothing, which is why onPartial is a tail-latency feature, not a convenience.

Two things follow, and wobbly does both:

  • Warm the pool when you're idle, with warmWorkerPool(). Worker startup is part of your worst frame; spawning lazily means paying it during the first burst, exactly when you're already busy.
  • Idle is free. Workers park on an event; nothing polls or spins between bursts.

Does this actually make things faster?

Yes — and how much depends almost entirely on whether you hand it a TypedArray.

Every item has to reach the worker somehow. A plain Array of numbers is copied by structured clone, element by element: for 10M numbers that alone costs ~227ms. The same data as a Float64Array is handed over by transfer — a pointer move — in ~0ms. wobbly does this automatically when the input is a TypedArray.

Measured on a 10-core M-series Mac:

Workload Serial wobbly (Array) wobbly (Float64Array)
Heavy — primality of 100k integers near 1e12 3557ms 767ms (4.6×) 552ms (6.4×)
Trivialn > 0.5 over 10M floats 74ms 238ms (0.3× 🐌) 47ms (1.6×)

The bottom-left cell is the classic worker trap: a cheap callback over a big plain array, where you pay to copy 10M numbers and get no work to spread in return. With a TypedArray that trap disappears — even a trivial predicate comes out ahead, because there's nothing left to pay.

So: if your data is numeric, keep it in a TypedArray and wobbly is close to free. If it's an array of objects, it must be cloned, and you need genuinely heavy per-item work to come out ahead.

(The heavy row's Array column uses the default half-pool; the Float64Array column adds .withWorkers(10) to take every core. See Worker pool.)

Quick start

import { AsyncArray } from 'wobbly-js'

const isPrime = (n) => {
  for (let i = 2, s = Math.sqrt(n); i <= s; i++) {
    if (n % i === 0) return false
  }
  return n > 1
}

// Big integers, so the primality loop actually has work to do.
const numbers = Array.from({ length: 1e5 }, (_, i) => 1e12 + i)

const primes = await new AsyncArray(numbers).filter(isPrime)

AsyncArray splits the array into one contiguous chunk per worker and reassembles the results in input order, so map and filter return exactly what the built-in methods would.

spawn() — a stateful worker, no build step, no eval

AsyncArray drags data to the work. spawn() puts the work where the data is — and that is the only thing that reliably pays (see SHOULD-YOU-USE-A-WORKER.md).

It hosts your module in a long-lived worker. new Worker(url) demands a same-origin script, which is why a CDN-loaded library can't use workers at all; a Blob URL inherits the document's origin, so it may dynamic-import() your module by absolute URL — even cross-origin. import() is not eval, so unlike AsyncArray this needs no 'unsafe-eval'. Verified in Chromium under a CSP that refuses it.

// main thread
import { spawn } from 'wobbly-js'

const gm = await spawn('https://cdn.example/game-master.js')

gm.on('spawn', (npc) => world.place(npc)) // recipes out
gm.send('player', { pos, action }) // events in
const stats = await gm.call('stats') // request/response
// game-master.js — a real module. Import anything. State is RESIDENT.
import { serve } from 'wobbly-js/worker'

const world = { npcs: [] } // never crosses the membrane
let beats = 0

serve({
  player(event, { emit }) {
    beats++
    const npc = decide(world, event) // your agent, your VM, your rules
    if (npc) emit('spawn', npc) // a recipe, not data
  },
  stats: () => ({ beats, npcs: world.npcs.length }),
})

The two shapes this is for

1. An agent with resident state. A story graph, a planner, a simulation brain. Its world model lives in the worker and never crosses. Only a small event goes in and a small recipe comes out. Pair it with a fuel-metered VM (tjs-lang's AgentVM) and you get something JS cannot otherwise do: cooperative preemption. Give the agent 2ms of fuel, it yields, you resume it — so it can never stall your frame. wobbly's own AsyncArray cancellation can only terminate a worker, because a JS hot loop cannot be interrupted.

2. Processing where the data lives. The worker fetches its own data — from the network, from IndexedDB — crunches it, and returns a compact result. The raw bytes never touch the main thread.

serve({
  async report({ url }) {
    const raw = await fetch(url).then((r) => r.arrayBuffer()) // fetched HERE
    return summarise(raw) // only this crosses
  },
})

Any TypedArray you return is transferred, not copied — so a big result is free.

Handlers may be async. A thrown error rejects the corresponding call(). If the worker fails to start, the error names CSP, because that is nearly always what it is.

Your callback has no closure

This is the one rule that matters, and everything else follows from it.

Your callback is serialized with Function.prototype.toString(), shipped to each worker, and rebuilt there with new Function(). It arrives with no closure — every variable it captured lexically is simply gone. So this throws ReferenceError: helper is not defined:

const offset = 3
const helper = (n) => n * 2

// Broken: neither `offset` nor `helper` exists inside the worker.
await asyncNumbers.map((n) => helper(n) + offset)

Instead, pass what the callback needs to withContext(). It becomes the callback's this:

// Works. Note it's a `function`, not an arrow — arrows cannot be re-bound.
function addOffset(n) {
  return n + this.offset
}

const shifted = await asyncNumbers.withContext({ offset: 3 }).map(addOffset)

The context must survive JSON.stringify() — plain data only, no functions, no class instances. If your callback needs a helper function, inline it into the callback body.

withContext() returns a new AsyncArray and does not mutate the original, so a context can't leak into unrelated operations on the same data.

Array items are copied by structured clone, so they must be cloneable too: plain objects, numbers, strings and the like — not DOM nodes, not functions.

TypedArrays are the fast path

If your data is numeric, hand wobbly a TypedArray and the copy cost essentially vanishes. There's no flag to set — it's automatic:

const data = Float64Array.from(readings)

const loud = await new AsyncArray(data).filter((n) => n > threshold)
//    ^ a Float64Array. Filter preserves the container type.

Chunks are transferred to the workers rather than cloned, and results are transferred back. Your array is never detached — wobbly slices off its own copy of each chunk first, so the array you passed in is still there and still usable afterwards.

filter gives you back the same TypedArray type it was given, which is safe: the survivors are input elements, so nothing is converted.

map is the one place you have to choose. By default it returns a plain array:

const halves = await new AsyncArray(Int32Array.from([1, 2, 3])).map(
  (n) => n / 2
)
// [0.5, 1, 1.5] — a plain array

That's deliberate. TypedArray.prototype.map would coerce every result back into the input's element type, so 1 / 2 in an Int32Array would silently become 0. If you want the result transferred back as a TypedArray — and you accept the coercion — say so with into:

const doubled = await new AsyncArray(data).map((n) => n * 2, {
  into: Float64Array,
})
//    ^ a Float64Array, transferred, not cloned

SharedArrayBuffer: the zero-copy path

Transferring a TypedArray is fast, but wobbly still has to slice each chunk out of your array and allocate somewhere to put the results. A shared array removes even that: the workers read and write your memory, in place. Nothing is copied in either direction, and nothing is allocated per call.

Hand wobbly a TypedArray backed by a SharedArrayBuffer and it happens automatically:

const input = new Float64Array(new SharedArrayBuffer(n * 8))
const output = new Float64Array(new SharedArrayBuffer(n * 8))

// Workers write straight into `output`. Resolves to `output` itself.
await new AsyncArray(input).map(transform, { out: output })

out is the arena: allocate it once, reuse it every frame. That's the difference between a render loop that allocates 80MB per pass and one that allocates nothing.

Measured on 10M Float64s with moderate per-item work, 10 workers: serial 313ms, transferred 60ms, shared 51ms. The remaining copy was ~16% of the parallel time — and on a single worker with a cheap callback, where the copy dominates, it's closer to 40%.

Most people should not use this

SharedArrayBuffer requires cross-origin isolation — the COOP and COEP response headers — and that cost is much bigger than it looks:

  • You may not be able to set headers at all. GitHub Pages, for one, can't.
  • COEP breaks cross-origin embeds that don't send CORP: images, iframes, analytics, third-party scripts.
  • If you are a library, you cannot demand this of your host application. Requiring COOP/COEP of someone else's app is a bigger ask than unsafe-eval.

So shared memory is for applications that control their own headers and have genuinely large numeric arrays — DSP, image processing, simulation. If you're shipping a library, or you have a few numbers of input and a buffer of output, use the transfer path instead. It needs nothing, and for you it will be just as fast.

Everything degrades on its own: out works whether or not memory is shared (zero-copy when it is, filled from transferred results when it isn't), so the same code deploys everywhere. You never have to branch on sharedMemoryAvailable().

A shared input is never transferred — that would detach your own memory. filter still copies on the way back, because the survivor count isn't known in advance.

Reduce: read this before you use it

A parallel reduce is not a serial reduce. wobbly folds the array in chunks, so unless you say how the chunks merge, your reducer must be associativef(f(a, b), c) must equal f(a, f(b, c)).

Adding is associative. Math.max is associative. Subtracting is not, and a non-associative reducer gets you a confidently wrong answer rather than an error:

const subtract = (acc = 0, n) => acc - n

;[10, 1, 2, 3].reduce(subtract) // 4
await new AsyncArray([10, 1, 2, 3]).reduce(subtract) // something else entirely

So for anything beyond a trivially associative fold, tell wobbly how to combine two partial results:

const counts = await new AsyncArray(harvest).reduce(
  // Runs in a worker: fold one item into a tally.
  (counts = {}, item) => {
    counts[item.fruit] = (counts[item.fruit] || 0) + 1
    return counts
  },
  {
    // Runs on the main thread: merge two workers' tallies.
    combine: (a, b) => {
      for (const [fruit, n] of Object.entries(b)) a[fruit] = (a[fruit] || 0) + n
      return a
    },
  }
)

Folding an item into an accumulator and merging two accumulators are different operations, and whenever the accumulator is a different shape from the items they are always different. combine is how you say so.

Two conveniences: combine runs on the main thread, so unlike every other callback here it has a normal closure and can use anything in scope. And you don't need withContext just to feed it — the fruit list in that example is only needed by the merge, which already has it.

There's no seed value, so the accumulator arrives as undefined on the first item. Give it a default — counts = {} above. Reducing an empty array throws, exactly as Array.reduce does.

One caveat even for "safe" reducers

Floating-point addition is not truly associative. Summing a million random floats serially and through wobbly gives answers that differ in the last few bits:

serial   499490.8326655442
parallel 499490.8326655383

That's ~6e-9 on a value near 5e5 — irrelevant for a progress bar or a statistic, potentially not irrelevant for money or a checksum. It is inherent to splitting the sum, not a bug wobbly can fix. If you need bit-exact reproducibility, don't parallelize the reduce.

The older this.final style still works

If you omit combine, wobbly merges the partials with the reducer itself, setting this.final to true for that pass so the callback can branch. This is retained for compatibility, but combine says the same thing more clearly and keeps the two operations honestly separate.

Progress

Pass a second argument to any operation to receive progress in the range 0…1, throttled to roughly one call per percent. Counting is automatic; your callback doesn't need to cooperate.

const results = await asyncNumbers.map(expensiveFn, (progress) => {
  progressBar.value = progress
})

Streaming results as they land

By default you wait for the slowest worker. onPartial hands you each worker's results the moment they arrive, so you can start using them:

const tiles = await new AsyncArray(tileSpecs).map(buildTile, {
  onPartial: (chunk, startIndex) => {
    // fires ~4 times for 4 workers, first one long before the last
    chunk.forEach((tile, i) => uploadToGpu(tile, startIndex + i))
  },
})

startIndex is where the chunk begins in the input — for a map, that's also its offset in the final array. The promise still resolves with everything at the end; onPartial is a head start, not a replacement.

For a generation workload (24 terrain tiles from a noise field, say) this turns "nothing for 2.4ms, then everything" into "first tiles at 0.5ms" — the difference between filling a frame budget progressively and stalling on the slowest tile.

If your callback builds a TypedArray per item, those buffers are transferred back rather than cloned, so this pattern is cheap:

const heightfields = await new AsyncArray(specs).map((spec) => {
  const out = new Float32Array(66 * 66)
  // …fill it…
  return out // transferred, not copied
})

Cancellation

Pass an AbortSignal. The promise rejects with signal.reason.

const controller = new AbortController()

const pending = asyncNumbers.filter(expensiveFn, { signal: controller.signal })
cancelButton.onclick = () => controller.abort()

try {
  const results = await pending
} catch (e) {
  if (e.name === 'AbortError') return // user cancelled
  throw e
}

JavaScript has no preemption — you cannot politely interrupt a worker mid-loop — so aborting terminates the workers doing the work and replaces them with fresh ones. The work genuinely stops rather than running on in the background. Any side effects your callback already performed inside those workers stand.

Errors

If your callback throws inside a worker, the returned promise rejects with that error, and the affected workers are replaced so the pool stays healthy.

await asyncNumbers.map(() => {
  throw new Error('kaboom')
}) // rejects: wobbly worker failed: kaboom

Why there's no worker file to set up

new Worker(url) needs a same-origin script URL — a separate file, which means bundler configuration, an extra asset to serve, and a deployment story, before you can run one line of code off the main thread. That friction is why most people never bother with workers at all.

wobbly builds its worker from a Blob URL and rebuilds your callback inside it with new Function(). There's nothing to configure and nothing to serve: import it and go. That's the whole point of the library.

The price is a CSP grant. Verified in Chromium, wobbly needs two directives:

script-src 'unsafe-eval';   # to rebuild your callback inside the worker
worker-src blob:;           # to spawn the worker at all

Grant both and wobbly works under a Content-Security-Policy — tested. Grant neither and it fails cleanly (it rejects, with an error that names CSP; it does not hang).

Note the order of failure, which is not what you'd guess: under script-src 'self' the blob worker is refused first, before new Function() is ever reached. So worker-src blob: is the directive people miss.

If you can't grant those — a hardened enterprise CSP, or a Chrome MV3 extension — you need a real worker file and something like comlink, and you'll be doing the build setup wobbly exists to avoid. That's the trade, made deliberately.

Worker pool

Workers are expensive, so wobbly keeps a pool. It's created lazily on first use — importing the library on its own does nothing — and holds one worker per navigator.hardwareConcurrency.

Each operation claims half the pool by default. That's deliberate: it lets two concurrent operations overlap rather than deadlock behind each other, at the cost of any single operation using only half your cores. It's also why the benchmarks above show ~5× on 10 cores, not ~10×.

If one operation is all you're running, take the lot — or leave headroom if you're not:

await new AsyncArray(data).withWorkers(10).map(fn) // use every core
await new AsyncArray(data, { workers: 2 }).map(fn) // stay out of the way

And you can size the pool itself, or tear it down:

import { configureWorkerPool, terminateWorkerPool } from 'wobbly-js'

configureWorkerPool({ size: 4 }) // must be before the first operation
terminateWorkerPool() // on teardown; the next operation builds a fresh pool

API

new AsyncArray(array, options?) wraps an array — options is { workers? }. Every operation is async.

Method Resolves to
map(fn, options?) a new array of results, in input order
filter(fn, options?) the items for which fn returned truthy, in input order
reduce(fn, options?) the reduced value — read the caveat
forEach(fn, options?) undefined
withContext(obj) a new AsyncArray binding obj as the callback's this
withWorkers(n) a new AsyncArray claiming n workers per operation

options is { onProgress?, signal? } — plus combine? for reduce — or a bare function as shorthand for { onProgress }. Neither withContext nor withWorkers mutates the receiver.

TypeScript users can type the bound this with the exported WobblyContext:

import { AsyncArray, type WobblyContext } from 'wobbly-js'

function addOffset(this: WobblyContext, n: number) {
  return n + this.offset
}

Limits

Worth knowing before you adopt it:

  • No strict CSP. new Function() needs unsafe-eval, and the worker needs worker-src blob:. No Chrome MV3 extensions. See why.
  • Parallel reduce is not serial reduce. Non-associative reducers are silently wrong unless you pass combine. See above.
  • Browser (or Bun) only. It needs Worker, Blob, and navigator.hardwareConcurrency. Node has worker_threads, not the web Worker global, so it isn't supported.
  • forEach can't touch your app. It runs in a worker with no DOM and no access to your main-thread state, so it's only useful for side effects within the worker.
  • Items and context must be serializable — structured clone and JSON.stringify respectively.
  • The whole array is copied, so expect peak memory of roughly double.

Development

bun install
bun test           # run the suite
bun run typecheck  # tsc --noEmit
bun run format     # prettier --write .
bun run pack       # test + typecheck + build to dist/

License

Apache-2.0 — see LICENSE.

About

use web workers without server-side ceremony

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages