Skip to content

Repository files navigation

@lickle/bin

Binary codecs for JavaScript. Define a schema, encode/decode Uint8Array.

Build Status Build Size Version Downloads

Install

npm install @lickle/bin

Encode / decode

import * as b from '@lickle/bin'

const User = b.struct({
  id: b.uint32(),
  name: b.utf8(),
  active: b.bool(),
})

const bytes = b.encode(User, { id: 1, name: 'ada', active: true })
const [user, pos] = b.decode(User, bytes)

Infer types:

type User = b.Codec.Decoded<typeof User> // decode output
type UserIn = b.Codec.Encodable<typeof User> // encode input (may be wider, e.g. optional fields)

measure(codec, value) returns the encoded byte length without retaining the bytes. encodeInto(codec, value, buf, offset?) writes into an existing Uint8Array and returns the new write position; throws RangeError if the buffer is too small.

Primitives

b.bool()
b.uint8()    b.int8()
b.uint16()   b.int16()    // pass `true` for little-endian
b.uint32()   b.int32()
b.float32()  b.float64()
b.bigUint64() b.bigInt64()
b.varint()                       // LEB128 unsigned, 1–5 bytes, range [0, 2^32 − 1]
b.utf8()                         // varint length-prefixed
b.utf8({ size: 20 })             // fixed-width, throws when encoded length ≠ size
b.utf8({ size: 20, pad: true })  // fixed-width, zero-pads short strings
b.bytes()                        // varint length-prefixed Uint8Array
b.bytes({ size: 16 })            // fixed-width, throws when input length ≠ size
b.bytes({ size: 16, pad: true }) // fixed-width, zero-pads short inputs
b.literal('v1')                  // zero-byte constant

Discriminated unions

// Keyed — variant chosen by which key on the input is set (others undefined).
b.keyed({ a: b.u8(), b: b.utf8() })

// Tagged — every variant is a struct with a `literal`/`constant` field at `key`.
b.tagged('kind', [
  b.struct({ kind: b.literal('circle'), r: b.u32() }),
  b.struct({ kind: b.literal('square'), w: b.u32(), h: b.u32() }),
])

keyed.fixed(...) / tagged.fixed(...) produce constant-size codecs (tag + max(variant.s)), zero-padding shorter variants — required for lens random-access. The dynamic forms encode tightly (no padding).

keyed throws on encode if more than one variant key is set (undefined values are ignored).

Composites

b.struct({ x: b.uint8(), y: b.uint8() })

b.tuple(b.uint8(), b.utf8())

b.array(b.uint16(), 4) // fixed-length, no length prefix
b.list(b.utf8()) // varint length
b.list(b.utf8(), { length: b.uint8() })

b.optional(b.utf8()) // value | undefined
b.nullable(b.utf8()) // value | null

Tagged unions (discriminated by a struct field):

const Msg = b.tagged('type', [
  b.struct({ type: b.literal('text'), body: b.utf8() }),
  b.struct({ type: b.literal('img'), url: b.utf8(), w: b.uint16() }),
])

Recursive schemas:

type Tree = { v: number; c: Tree[] }
const Tree: b.Codec<Tree> = b.lazy(() =>
  b.struct({
    v: b.uint32(),
    c: b.list(Tree),
  }),
)

intersect(a, b) merges two struct shapes; on key collision, b's field wins both at the type level and on the wire. fields(struct, m) toggles per-field optionality — either via an exhaustive boolean map ({ a: true, b: false }) or selectively ({ optional: ['a'] } / { required: ['b'] }).

Length-prefix overflow

Custom length codecs cap the maximum collection size. Encoding a value whose length exceeds the codec's range throws RangeError:

const c = b.list(b.uint8(), { length: b.uint8() })
b.encode(c, new Array(300)) // throws: length 300 out of range [0, 255]

On decode, list bounds the claimed length by the remaining bytes (1-byte minimum per item) before allocating, so an attacker-controlled length prefix cannot trigger a runaway allocation. reader(buf, { trust: true }) opts out of this guard along with all other bounds checks.

Transformations

b.imap(
  b.uint32(),
  (n) => new Date(n * 1000),
  (d) => d.getTime() / 1000,
)
b.fallback(b.uint8(), 0) // returns default if decode throws
b.json<{ a: number }>() // utf8 + JSON.parse/stringify

imap strips composite shape metadata, so a lens built over an imap'd codec does not drill into the underlying struct/tuple/array — only read/write/slice are exposed.

Lens — random access without full decode

const L = b.lens(User)
const buf = b.encode(User, { id: 1, name: 'ada', active: true })

L.id.$read(buf) // 1
L.active.$write(buf, false)
L.$read(buf) // { id: 1, name: 'ada', active: false }

Lens base accessors are $-prefixed ($read, $write, $slice, $codec, $offset, $size) so they never collide with struct field names like read/write/size.

Arrays and lists:

const Pts = b.array(b.struct({ x: b.uint8(), y: b.uint8() }), 8)
const LP = b.lens(Pts)
LP.at(3).x.$read(buf)
LP.at(3).x.$write(buf, 42)

const Names = b.list(b.utf8())
const bound = b.lens(Names).bind(buf)
bound.at(0)
bound.toArray()

.bind(buf) — preferred for repeated access

Every lens node exposes .bind(buf, base?) which returns a parallel tree whose $-accessors close over the buffer. Bind once, then $read() / $write() / $slice() (and at(i), field accessors, etc.) take no buffer argument:

const L = b.lens(User).bind(buf)
L.$read() // whole struct
L.id.$read() // field
L.active.$write(false)

const Pts = b.lens(b.array(Point, 8)).bind(buf)
Pts.at(3).x.$write(42) // .at(i) returns a bound child

const Names = b.lens(b.list(b.utf8())).bind(buf)
Names.length // BoundList helpers
Names.at(0)
Names.$slice() // bound base accessors are present too

Bound nodes also expose absolute $offset (already includes base) and the same $size/$codec as the unbound lens. Binding once and reading many times is the recommended pattern.

In-place writes must preserve byte length. Writing a sized field at a known offset always works. Writing an unsized field (e.g. utf8) or a whole composite throws RangeError if the new value would have a different encoded length than the existing bytes — anything else would shift downstream data. Use encode for resizing rewrites.

Reader / writer

const w = b.writer()
b.write(w, b.uint8(), 7)
b.write(w, b.utf8(), 'hi')
const bytes = w.flush()

const r = b.reader(bytes)
const [n, s] = b.read(r, b.uint8(), b.utf8())

b.writer() allocates a 256-byte buffer that doubles on demand. The other two forms wrap a fixed-capacity buffer:

b.writer(size) // preallocate exactly `size` bytes
b.writer(buf) // wrap an existing Uint8Array
b.writer(buf, { trust: true }) // ...and skip the bounds check

When a composite encoder asks for more room than is available, writer(size) and writer(buf) throw RangeError('writer: out of space …') (instead of producing an opaque DataView error later). { trust: true } makes grow() a true no-op — use it when you've already proven capacity.

The Uint8Array view is the standard interchange currency for an ArrayBuffer slice:

b.writer(new Uint8Array(ab, offset, length))
b.reader(new Uint8Array(ab, offset, length), { trust: true })

encodeInto(codec, value, buf, offset?) writes into an existing Uint8Array and throws the same clear RangeError if the buffer is too small.

Reader modes:

b.reader(bytes, { trust: true }) // skip bounds checks (only for buffers you produced)

Decode failures throw plain RangeErrors. For path-aware diagnosis use b.annotate(codec, bytes) or b.layout(codec, bytes) — the walker catches per-field decode failures, emits type: 'error' slots, recovers past sized failures, and appends a FAILED at <path> (offset N): <message> line for dynamic failures.

Wire format

Codec Bytes
bool 1 byte (0 = false, non-zero = true on decode)
uint8 / int8 1 byte
uint16 / int16 2 bytes, big-endian (little-endian if le = true)
uint32 / int32 4 bytes, big-endian
float32 / float64 4 / 8 bytes, IEEE 754, big-endian
bigUint64 / bigInt64 8 bytes, big-endian
varint LEB128 unsigned, 1–5 bytes (range [0, 2^32 − 1])
literal(v) 0 bytes
utf8() varint length, then UTF-8 bytes
utf8({ size: N }) exactly N bytes; decode strips trailing null bytes
bytes() varint length, then raw bytes
bytes({ size: N }) exactly N bytes
optional(c) 1-byte flag (0 = absent), then c if present
nullable(c) 1-byte flag (0 = null), then c if present
array(c, N) N × c concatenated, no length prefix
list(c) varint length, then N × c
tuple(c1, …) each element in order
struct({ k: c, … }) each field in declaration order
keyed({ k: c, … }) 1-byte variant index (configurable via { tag }), then variant payload
tagged(t, variants) 1-byte variant index (configurable via { tag }), then variant struct

bytes() decoded values are zero-copy views into the source buffer — mutating either aliases the other. Treat decoded Uint8Array as read-only or copy before mutating.

License

MIT © Dan Beaven

About

A tiny, efficient utility for defining binary data schemas and performing encoding/decoding of JavaScript objects to and from Uint8Array

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages