Repository navigation
bench: make current macOS evidence independently reproducible #180
Description
Activity
- addedhelp wantedExtra attention is neededExtra attention is neededfoundationOpen-source foundation readiness and its evidenceOpen-source foundation readiness and its evidence
on Sep 7, 2026 amishabenramani commented
on Sep 15, 2026 MemberMore actions2026-09-15 public benchmark-owner refresh
One macOS atom named by the KEL-129 work is now already backed by public raw evidence and should not be reimplemented here:
gyldlab/keld-benches#21merged (b7137c1c2107e0f622eb350819d32605be7a781f). It pins Keld4fbf94bbb755854058067b986877177f00b25a39and records a real cross-process Rust↔Rust kipc echo over a Unix-domain socket on Apple M4 / macOS 26.5.1 arm64: 20 independent sessions × 100,000 calls for each of the small and 1,024-byte tiers, with 40 raw session documents committed. The PR also preserves a forged-token negative control and explicitly limits the evidence to macOS.Important scope for this help-wanted issue:
- Reuse those raw IPC results as the existing
ipc.rtt.macos.library-armevidence; do not create another IPC runner or reconstruct its medians from prose. - PR docs(perf): superseded by #44 after branch rename #21 itself names a remaining result-schema/registry mismatch for block-bootstrap-by-session evidence. Do not silently make the schema claim more than it validates.
- This issue is still open because IPC is only one atom. A complete public run still needs product-shaped artifact/process provenance for the other claimed macOS metrics (for example first paint/RSS/package/update/dev-loop where applicable), with exact hardware/OS/tool/build/cache/census/sample metadata.
- CI compilation is not a benchmark and the Rust↔Rust IPC result is not a Bun/full-product latency or overall framework-performance claim.
Public source/evidence reference: gyldlab/keld-benches#21. Keep any private research notes optional; an external contributor must be able to complete the public task from public repos and recorded evidence alone.
- Reuse those raw IPC results as the existing
amishabenramani commented
on Sep 16, 2026 MemberMore actions2026-09-16 reproducibility reconciliation
The public evidence split is now explicit:
- IPC atom: keep reusing merged
gyldlab/keld-benchesPR docs(perf): superseded by #44 after branch rename #21. It already provides the macOS Rust↔Rust KIPC fixture, 40 raw session documents, negative control, hardware/OS/source pins, and block-bootstrap summary. It remains a library/wire-path measurement, not Bun/full-product evidence. - Product paint/RSS atom: the existing KEL-64 macOS harness has been preserved rather than rewritten, but it lives on the
keld-benchesfeature lineage (agent/kel-64-macos-perf-oracle, currentlyd079c672abfacee6eab22354840838e9560e5ec1), not on publicmain. Historical PR fix(wv/macos): default-deny camera/mic via keld-guard #6 was merged into that feature branch. - Current public
keld-benchesMEASUREMENTS.mdstill marksmacos/keld/hello/Not yet measured, and that fixture explicitly says it has no first-paint, RSS, or shipped-size claim.
KEL-129's latest handoff records that everything possible without physical-machine control is complete. The next valid evidence requires the real Apple M4 Mac mini foreground measurement session using the preserved harness/recipe. Historical summary-only/traced diagnostic values, Windows Keld numbers, and CI runs must not be promoted into this macOS product row.
So this issue stays open. Next completion check: publish a fresh immutable Mac product result (exact source/artifact hashes, environment, process census, raw samples/statistic, and uncertainty) from the existing harness; then reconcile the public report. No second runner is needed.
- IPC atom: keep reusing merged
Problem
The public project needs a reproducible path from benchmark claim to OS-qualified fixture, built artifact, process census and raw samples. Existing macOS runner work and summaries are not a substitute for a complete published run.
Acceptance
Coordinate with existing KEL-64/KEL-129 work. Reuse gyldlab/keld-benches and its OS-qualified harness; preserve unmerged work instead of writing another runner. Recover raw sample/manifest provenance where available, or publish a clearly new run without reconstructing old medians. Separate census, workload, queue/copy, clock, statistic and artifact; include Bun and engine helpers where the metric requires them. Report cache/build/tool/hardware versions and uncertainty. CI compilation is not a benchmark. No superiority/budget verdict until comparable evidence supports it.
Public code and evidence: https://github.com/gyldlab/keld-benches. Foundation #167; performance remains a measurement outcome.