Repository navigation
AI-SLOP: Develop best current practises for Open Source maintainers #178
Description
Activity
Great initiative!
I suggest that right after "Document the Problem" but before "Develop Detection Guidance" you add "Collect Existing Practices, Guides, and Ideas: Collect existing practices, guides, and ideas for countering AI slop, along with any information on their success or lack thereof" - basically, before creating guidance, let's capture what is already working (or not).
As noted in the meeting, we can also link to (free) things which might help write patches / accelerate - copilot is in the
resource guide here along with some more https://github.com/cloudcommunity/Free-for-Open-SourceWorth noting GH is looking at this space: https://github.com/orgs/community/discussions/185387
Reacted by Chris de Almeida and Niko FöhrReacted by Sebastian Grüner and Piotr P. KarwaszAbsolutely - I think we definitely need best practices here and some way to guide and help maintainers to navigate this problem - and also possibly work with GitHub to add some tooling to support it.
I would think that some of the ideas here might be turned into some kind of automation that GitHub could implement - they are now discussing of different ideas they might have - and I think it might be good to reach out to them to also cooperate with someone like this working group - extrernal to GitHub - that might also serve as some kind of a sibling to product GitHub team by providing an external input ?
I am happy to propose it in the discussion there - been discussing it in the past with different people in GitHub? I did not make it to the Maintainer's UnConference after FOSDEM (I simply collapsed after all 6 days of running around and the GitHub Maintainers dinner lasting - with after-party - till 3 am did not help), but I know few people were there - like @ppkarwasz - so maybe some of that was already discussed there ? I know the ai-slop case was being discussed there.
Reacted by Kevin YangI was at the Maintainer's UnConference but I was in a different group (also arrived late for--huh--similar reasons as you couldn't make it). Maybe @abbycabs can help make the right connections here?
Well, @taladrane is Github, so we already have extremely good connections.
Thank you for all the feedback, keep it coming!
Reacted by Jarek Potiuk, Tobie Langel and nyoI was at the Maintainer's UnConference but I was in a different group (also arrived late for--huh--similar reasons as you couldn't make it). Maybe @abbycabs can help make the right connections here?
Learning for the future.. Do the social AFTER the unconference 😄
written in response to discussions at GVIP: the-tcpdump-group/tcpdump-htdocs#52
Rather than PoC videos (which bring their own issues) Apache Tomcat has been been asking for the source code (web application and client) for a minimal PoC when we suspect AI slop. It has been very effective so far. The bit we need to get better at is spotting the slop in the first place but the more slop we see, the easier it is to spot.
Reacted by Art ManionPoC||GTFO comes to mind (as noted above). The biggest concern for my small team is the combination of AI slop masking a non-vulnerability or a very low impact issue. Very expensive to sort these out and be confident that indeed they are non-vulnerabilities or very low impact.
The key element of a good report is reproducability or PoC. If AI helps someone translate or communicate, fine, as long as the concise PoC is also part of the report (and not hidden in the slop). IMO any reporting policy or criteria should focus on this as the primary requirement.
That's a great initiative. As an old security nerd, I'm huge believer in the value of Bug Bounty programs. And that we should give a chance and encourage everybody to report a security research report. One of the important educational initiative would be to pivot from publicly shaming bug bounty programs (as "poisoned by AI slop") to promoting the best practices and tools on how to minimize the "rotting" effect of AI reports.
9 remaining items
A updated the draft of the working group's AI-slop guidance a few weeks ago but failed to post a comment here.
It's written as two companion parts:
- Guide for Maintainers — practical steps for projects receiving AI-assisted vulnerability reports: reducing low-quality submissions before they arrive (threat models, a strong SECURITY.md, clear submission requirements), triage signals and a reusable response template, AI contribution policies, and protecting maintainer time and wellbeing.
- Guide for Finders — responsibilities for researchers using AI tools: understanding the project before submitting, independently verifying findings, disclosing AI use, and what a high-quality report looks like.
https://docs.google.com/document/d/1csseaiMVQeILSPjx3BvpCBH88PifgPf_ebXVKD5DIOs/edit?tab=t.0
Please leave comments and suggestions directly in the doc. Thanks to everyone who's contributed so far!
Do we want to finalize this so we have this as a basis before wrapping up:
Document existing practices for responding to AI-Slop - #179
AI Contribution/Disclosure Policy Template(s) - #180
Community Survey on AI-Slop Impact - #181
Blog Post for Q2 Vulnerability Coordination – Community Involvement on AI-Slop - #182I work the researcher side of this — authorization vulnerability research with an AI-assisted pipeline — so let me offer what the slop-versus-signal line looks like from inside a pipeline, since "reduce slop, not ban AI" depends on maintainers being able to draw it.
In my experience the discriminator is not the writing style, and not whether AI was involved. It is whether the report survives four checks that a valid finding passes and slop almost never does:
- A working reproduction, not a source-pattern guess. A real report includes a live PoC against a shipped release. Slop stops at "this pattern looks dangerous."
- A complete source-to-sink trace with file and line, and an author who can explain the whole data flow when asked. Slop cites a sink without showing it is reachable.
- The finding is in a shipped release, not only on main.
- The obvious false-positive shapes are already ruled out. This is the strongest tell. Most AI-generated slop pattern-matches a dangerous-looking sink and skips whether it is reachable by an under-privileged actor, actually read by a security gate, or simply intended behavior for a role that already has the capability. A valid report shows those were checked; slop shows they were not, because the model never traced past the sink.
That last point is usable as detection guidance without any AI-detector, which the issue rightly notes is unreliable. You are not detecting AI; you are checking whether reachability and authorization were traced. A report that asserts impact with no reachability trace is low-signal whether a human or a model wrote it.
I keep the rejected candidates from my own pipeline with the reason each was closed, and the recurring shapes are a short list: intended-by-design, not-in-a-shipped-release, writable-but-never-read, guard-one-layer-up, wrong-table-for-the-permission. If a labeled catalogue of those shapes would help the Best Practices document or the detection-guidance deliverable, I am happy to contribute it. The rejects are the scarce half — most people do not keep them — and they are exactly what shows a maintainer where a report went wrong.
@Santoshkumarpuppala - I really like your other points! However, I disagree with "The finding is in a shipped release, not only on main."
If it's in main, it may not be an exploitable vulnerability in shipped software, but if nothing changes it will be. Also, if it's OSS, there are people who deploy "from tip of main" instead of what officially shipped (perhaps they shouldn't do this but this happens). So the process may be different, but I don't agree that this indicates slop.
Thanks!
Reacted by Mihai Maruseac@david-a-wheeler — that's a fair correction, and I'll take it. You're right: a flaw on main is a real vulnerability in waiting, people do run from tip-of-main, and "only on main" doesn't make a report slop.
Let me restate what I was actually reaching for, because the version detail matters more than the release boundary I drew. The tell isn't that a bug lives on main — it's a report that asserts shipped-software impact without saying where the issue is reachable: which version or branch, default configuration or not, and what it was actually run against. A finding reported honestly as "present on main, not yet in a release" is valid and useful. The same finding dressed up as "exploitable in production" with none of that stated is the shape I was pointing at. So the check is better written as: the report is explicit about the affected version/branch and the conditions under which it's reachable — not that it has to be in a release.
That also lines up with your point that the process is different. A main-only finding is usually a pre-release fix conversation with the maintainer rather than a CVE against a shipped version, and a report that says so is doing the reader a favor. Which affected-version and reachability fields a report states up front is probably worth capturing in the detection guidance either way.
There's now something meta going on in this thread.
Reacted by Mihai Maruseac, Kris Borchers and Craig JellickThere's now something meta going on in this thread.
I was thinking the exact same thing
I'm okay with meta :-). I'll take truth wherever I can get it.
Since the guides may be close to being finalised, here is the catalogue rather than another offer to write one. These are the recurring reasons candidates in my own pipeline get closed as not-a-vulnerability, in rough order of how often they come up. The first is the largest by a wide margin.
-
Already entitled. The actor who reaches the sink already has the capability by design — an admin role invoking an admin function, a library trusting its caller. Tell: the report never states what privilege the actor started with.
-
Not reachable in the stated configuration. Real code path, but gated behind a non-default setting, a cloud-only branch, or an unreleased branch. Restated per David's correction above: the defect is not that the code is on main, it is that the report does not say which version, branch and configuration it was reached in.
-
Guard one layer up. The check exists, but on a router, middleware or decorator the report did not read, so the handler looks unprotected in isolation. Tell: the trace starts at the handler.
-
Writable but never read. A field is settable or injectable, but nothing security-relevant consumes it. Tell: the report proves the write and asserts the consequence.
-
Wrong source for the permission. Rarer than the four above, but distinctive enough to be worth naming: the role or tenant is read from a different table or claim than the one the report inspected. Tell: privilege escalation claimed from a field that is not the authoritative one.
The usable form for the Guide for Maintainers is a response, not a taxonomy: for each shape, the one question that resolves it. Which privileges did your test account hold. Where does the request enter before it reaches this function. What reads this value. Where does that permission actually come from. What did you run it against.
Limits worth stating: this is one researcher's set, weighted to web-application authorization findings, and the labels are mine rather than adjudicated by the maintainers who received the reports. It is a starting vocabulary, not a measured distribution.
Happy to write this up properly for the Google Doc, or wherever the guides are being assembled — say which and I'll do it.
-
I run an independent security-research practice (North Echo Security Research) that leans heavily on AI assistance and discloses through vendor PSIRT/CERT channels. So I'm on the side of the table this thread is worried about: I generate a lot of AI-assisted analysis. The discipline framework below is what we use to keep that output from becoming the slop you're describing, and most of it turns into intake requirements a project or program could enforce.
The premise I've found useful: the fix for AI slop isn't detecting AI involvement; it's a validation bar that unvalidated output structurally cannot clear. "Was a tool or a model involved?" is close to unfalsifiable and, in my experience, the wrong question. "Did the submitter demonstrate this against real, shipped code with reproducible behavior?" is answerable, and it filters slop regardless of how the lead was generated. Each discipline below is here because it's a specific place a plausible-but-wrong machine finding dies.
Evidence bar
- Dynamic proof is mandatory; a source-code reachability argument is a lead, not a finding. The most common failure mode of AI analysis is a confident chain read off source that never fires at runtime. Requiring the effect to be reproduced on a live, exact-shipped build (not a dev branch or emulator) evaporates most of it.
- Reproduce N times from clean state, or it doesn't exist. One-shot results are frequently artifacts of dirty or lucky state.
- Positive and negative controls on every test; void the run if a control fails. A positive result with no negative control is indistinguishable from a broken harness that could only ever report success, which is exactly what an over-eager auto-runner produces. Reproducers should fail closed in both directions and assert on machine-greppable markers.
- Verify the runner's own tally, especially when a model wrote or drove it. Don't accept the summary line; confirm the precondition assertion actually fired every run.
Claim discipline
- An evidence ladder with no inference across the gaps. Gate claims from "static prediction" up through "attacker-reachable entry point," "controlled security-relevant effect (a primitive)," and only then "execution." Slop collapses this ladder: it reads source and asserts RCE, borrowing reachability, attacker-control, and execution it never showed. Forcing each rung to be independently demonstrated, with no state inferred from an adjacent component, makes that leap impossible to smuggle through. Name the primitive (a crash, an SSRF, a write); don't inflate it to impact.
- Prove it through the real attacker path: no aliases, no admin-side shortcuts. The most seductive AI PoC quietly assumes the hard step using privileges the threat model's attacker wouldn't have.
- Upstream-state gate before writing a word. Read the affected code in current upstream main, find the fixing commit if any, and establish whether it's unfixed, already-fixed, or fixed-for-an-unrelated-reason. A report about a bug the maintainer fixed weeks ago proves the submitter didn't read the tree, and nothing else in the report survives that impression. It's also where a large fraction of machine findings correctly die as already-known or intended-by-design.
- Compute CVSS; lead the conservative vector; never estimate. Scope-change math is non-linear; pattern-matching to a number is how everything becomes a 9.8.
Self-review and accountability
- A rebuttal audit before submission: enumerate the reviewer's likely dismissals (intended behavior, unsupported config, needs credentials, not attacker-controlled, already-known, severity overstated) and pair each with the evidence that answers it. Unanswered rebuttals are tracked as open gaps, not buried in prose. This institutionalizes the adversarial scrutiny a model omits about its own output.
- A human gates every external send; the assistant drafts and never transmits. Every candidate-to-confirmed promotion is a human decision. No model output reaches a maintainer without a person having verified it.
- Declare AI assistance honestly and honor the project's AI policy. Where a project declines AI-derived code, we ship a declared bug report with prose remediation and no patch. Honest compliance, not technical compliance.
The through-line: make the work reproducible enough that where it came from stops being the interesting question. When a report ships a guarded reproducer that fails closed both ways (and, where the fix is small, a patch), the conversation moves from "is this real / is this automated" to "is this code correct," which is the conversation maintainers actually want.
We publish our full standards here: https://northecho.dev/standards/. If it's useful, I'm happy to contribute this to the policy-template / best-practices deliverable too. Most of it translates directly into intake requirements a program can enforce at submission time, which moves the burden off the maintainer and onto the submitter, where it belongs.
Reacted by Art ManionI think threat modeling (or something like it, "use modeling?) is a key part of the solution here. Create a doc for agents to use to evaluate incoming reports.
https://news.apache.org/foundation/entry/security-scanning-at-foundation-scale
This is one such tool, but I really think the concept has legs:
https://github.com/alpha-omega-security/threat-model@SecurityCRob @taladrane Hey Madison and CRob, I updated the draft with comments and feedback we have received so far. At least I think I got it all. There are still some parts that I did not delete but we might once we talk, which you mentioned that you wanted to do this week. Just let me know when and we can jump on a call. Josh BTW - Here is a link to the doc again: https://docs.google.com/document/d/1csseaiMVQeILSPjx3BvpCBH88PifgPf_ebXVKD5DIOs/edit?usp=sharing
I built a small public tool that overlaps with the policy-template and LLM-friendly-documentation parts of this work, so I wanted to share the approach for critique.
Repo Policy Score - https://repopolicyscore.com/methodology checks the public guidance on a repository’s default branch for deterministic policy signals: where contribution rules live, whether AI-assisted work is addressed, who remains responsible, what must be tested, what an agent may do autonomously, when a human must take over, and who owns the rules. It shows the evidence behind the result and the gaps it found.
It does not try to detect whether a report or patch was written by AI, assess code quality or vulnerability severity, certify security, or enforce a policy. The scan is read-only and limited to public repository data.
I am sharing this as a working implementation, not a proposed standard. I would especially value feedback on two questions:
Which policy controls are missing or wrongly framed for vulnerability-report and security-contribution workflows?
Which findings should never result in suggested wording because they require an explicit maintainer decision?If it would help the working group, I can also provide the control-to-evidence mapping in a reviewable format.
Open source projects are increasingly facing a wave of low-quality, AI-generated vulnerability reports and contributions—commonly referred to as "AI-slop." This issue aims to develop best current practices for open source maintainers to help them detect, manage, and mitigate the impact of AI-slop on their projects while still benefiting from legitimate AI-assisted security research.
Problem Statement
The rise of AI tools has created a significant challenge for open source maintainers:
Goals
Key Themes from Existing Public Discussions
What Projects Are Doing
Policy Elements from Existing Projects
Key principles emerging from LLVM, Selenium, and Django policies:
Recommendations for Platforms
Platforms accepting vulnerability reports should consider:
Open Questions
Related Efforts
Proposed Deliverables
How to Contribute
We welcome input from:
Please share:
References
Blog Posts & Articles
Project Policies & Changes
Examples & Data
Talks & Events