Skip to content

"Reasoning with sampling" to improve efficiency and quality #9

Description

@naasking

Very cool project! While reading the clever self-embeddings approach, it occurred to me that some other recent work [1] was complementary. Rather than only scoring the quality after the sequence has been generated, "reasoning with sampling" computes a probability distribution of the output while it's generating, and backtracks and regenerates blocks on the fly if it sees falling confidence. This should let you filter out some bad candidates without having to complete and then score them via the geometric lens, thus a) improving end to end efficiency, and b) improving overall quality; having to evaluate fewer bad candidates means evaluating more possibly good candidates!

If I understand correctly, you've already made modifications to llama.cpp to extract the self-embeddings, so adding the MCMC over the logits seems like a natural evolution.

[1] Reasoning with Sampling: Your Base Model is Smarter Than You Think, https://arxiv.org/abs/2510.14901

Activity

  1. self-assigned this
    on Mar 27, 2026
  2. itigges22 commented on Apr 18, 2026

    @itigges22
    Collaborator

    Added to V3.2 Roadmap!

  3. itigges22 commented on May 12, 2026

    @itigges22
    Collaborator

    @naasking — wanted to flag that I closed the internal duplicate (#40) and pointed it at your original report, so this is now the canonical ticket for the "Reasoning with Sampling" work.

    Status as of V3.1.0: still tracked, not yet implemented. The V3.1.0 roadmap got streamlined to OS support + accelerator support + formal 9B benchmarks for V3.1.1; in-generation MCMC over logits remains exploratory. The paper is referenced in docs/SOURCES.md so the design context is preserved.

    The Geometric Lens has moved closer to where this complementary approach would plug in — V3.1.0 ships lens-as-PRM (PC-206 + PC-207): per-step C(x) + G(x) scoring during candidate generation, with a severe-score short-circuit (gx_min < 0.05 fires a corrective immediately). That's still post-token, but the per-step plumbing is now there. An in-generation MCMC layer would sit upstream of it.

    If you want to pick this up or pair on a prototype, leave a note here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions