Agent skills for EU medical device regulation that say where the rules stop.
claude plugin marketplace add Ecgecer/eu-mdr-skills
claude plugin install device-claims@eu-mdr-skills
/device-claims:device-claims-review
Not using Claude Code? Paste any file from dist/ into any chat, each one
carries a skill and its verbatim statute text in a single document. Nothing else needed.
Full install and usage · worked example · what it does not
do
- A benchmark, if you build regulatory AI
- See it before you install it
- The method, if you want to copy it
- Verify the claim yourself, in one command
- Scope limits, stated up front
- Which one do you want
- Install
- Use
- Using it outside Claude Code
- Evals
- Licence and attribution
The verify badge is not decoration. It runs scripts/verify-sources.py, which
re-fetches every statute source and diffs each quoted passage against it, so a green
badge means the text in this repo still matches the law it claims to quote, as of the
last run. It runs again every Monday.
Claude already knows MDR. Across this repo's 30 eval cases, the no-plugin baseline recalls implementing rule 3.3 verbatim, applies Rule 11's escalations correctly, and refuses to classify a non-device. Knowledge is not the gap.
The gap is confident over-reach. In those same measurements it cited HWG § 11(1) no. 2 against a medical device, when only nos. 7 to 9, 11 and 12 reach devices, in 3 of 3 runs. Told a page was French-market only, it correctly dropped German law and then asserted French advertising rules it cannot cite, in 3 of 3. It told a Dutch manufacturer to translate a Declaration of Conformity into German on the authority of MDR Art. 19(4), when MPDG § 8(1) accepts German or English, in 3 of 3. It answered "Class IIa, plan for a notified body" to a question the rule it cited does not settle.
Every one of those is plausible, well-reasoned, and wrong in a way you cannot see from the answer. Not a hallucinated rule. A real rule applied one step past where it reaches.
These skills pin every finding to verbatim statute text, state what they do not carry, and stop rather than conclude past their own boundary.
| Skill | Answers | Carries |
|---|---|---|
device-claims |
"Can we say this in our copy?" | MDR Art. 7, IVDR Art. 7, HWG, UWG |
mdr-classification |
"What class is our software?" | Annex VIII impl. rules 3.1–3.7, Rule 11 |
mdr-transition |
"How long can we still sell this legacy device?" | Art. 120(3)–(3d) as amended by 2023/607 |
mpdg-germany |
"Does Germany want more than MDR?" | MPDG §§ 4, 8, 73 |
scope-statement |
"What should this report say it didn't check?" | nothing, domain-general |
Every skill ships an eval suite measured against a no-plugin baseline, and every suite publishes the cases where the skill adds nothing.
Measured across 5 suites, 30 cases:
| Suite | Mean delta | Cases measured | Content |
|---|---|---|---|
device-claims |
+0.84 | 10 | MDR/IVDR Art. 7 plus HWG and UWG |
mpdg-germany |
+0.62 | 5 | German national law only |
scope-statement |
+0.30 | 3 | domain-general |
mdr-classification |
+0.27 | 7 | EU-level only |
mdr-transition |
+0.20 | 5 | EU-level only |
12 of those 30 cases measure a delta of 0.00, the skill changes nothing. They are published case by case, because a suite that reports only what it earns is not reporting.
This table is spliced in from the stored run data by scripts/report-evals.py and CI
fails if it drifts. It was kept off this page for a while because hand-typed figures are
how the repo published three wrong ones. The answer turned out to be generating it, not
hiding it. Per-case numbers, including every case where the skill adds nothing:
device-claims ·
mdr-classification ·
mdr-transition ·
mpdg-germany ·
scope-statement.
The pattern across the measured cases: positive delta wherever the model would over-apply, over-conclude, invent an authority or reach for boilerplate; zero wherever it already had what it needed. These skills do not add knowledge. They add the discipline to stop.
That is a fact about a frontier model, not about the skills. Re-run pinned to Haiku 4.5, the same suites invert: the knowledge cases that earn nothing against Opus earn +1.00 each, and the boundary cases that earn +1.00 against Opus earn +0.33 or nothing. The reference text transfers; the refusal discipline does not. MODELS.md has the per-case numbers and what they do not show.
Not legal advice. Drafting and risk-triage aids, not a substitute for a regulatory professional or a Fachanwalt.
benchmark/ exports every eval case as a portable, Apache-2.0
benchmark with measured baseline difficulty. It answers one question about any model
or tool, not just this one:
Does it apply EU medical device regulation to products, markets and audiences the provision it cites does not reach?
Of the 30 cases, 13 are ones Claude failed in every run with no reference material and no web access, 10 it passed in every run, and 7 it passed only sometimes. All three groups are published, because a benchmark that hides its easy cases overstates itself.
Those hardest cases are what the benchmark is for, telling a manufacturer to translate a Declaration of Conformity its member state accepts in English, citing a German advertising item that does not reach devices, asserting French advertising rules it cannot cite once told German law does not apply, manufacturing findings on clean copy.
Almost only Claude has been tested. One non-Claude run exists. Gemini 3.5 Flash,
baseline arm, six cases, with every verdict and its reason in benchmark/judgments/.
Four failure modes reproduced, two did not, and no skill arm ever ran. GPT, Llama,
Mistral and Gemini Pro are untested, and the
benchmark says so rather than generalising from one model, which would be the exact
failure it measures. Results from another model are welcome as a PR.
A worked example. One prompt, run with the skill and without it, both outputs verbatim from the harness.
The short version: asked to review consumer copy for a Class IIa blood-pressure monitor, the model without the skill cites HWG § 11(1) Nr. 2 against the physician endorsement. That provision is real and it described it accurately, but the closing sentence of § 11(1) gives medical devices only nos. 7, 8, 9, 11 and 12. No. 2 does not reach devices.
Act on it and you pull a lawful endorsement off a product page. Nothing in the answer signals it is wrong. That happened in 3 of 3 runs; with the skill, the correct answer happened in 3 of 3.
METHOD.md is how to build a regulatory skill worth trusting, for any regulation. Measure before you build; pin the text and publish its edges; let the skill refuse; test against a baseline or you are measuring the model; publish the cases where you added nothing; make the claim executable.
It also lists every trap we walked into, graders that punish correct reasoning, a freshness guard that could not see drift because it had been told what to look at, and a confident prediction that measurement destroyed.
CONTRIBUTING.md turns that into what a new skill has to ship.
Every provision these skills apply is stored verbatim, with its source URL and retrieval date. That is easy to assert and worth nothing unless you can check it, so checking it is one command:
python3 scripts/verify-sources.py
It re-fetches each source and confirms every quoted passage still appears, character for character after whitespace and quote-glyph normalisation. Elided quotes are verified fragment by fragment, because the joined string is not what the source says.
OK.../references/hwg.md (6 fragments across 4 quotes match)
OK.../references/uwg.md (5 fragments across 2 quotes match)
UNVERIFIED.../references/mdr-ivdr-art7.md
fetch failed: HTTP Error 403: Forbidden
check by hand: https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32017R0745
search the page for: "Article 7 Claims In the labelling, instructions for use..."
It does not pass silently for what it could not read. EUR-Lex blocks non-browser clients, so MDR and IVDR text is reported as UNVERIFIED with the exact string to search for. A verifier that reported success for a source it never fetched would be the same defect these skills exist to prevent.
Why the text is pinned locally at all. EUR-Lex is not a dependable read-through source. Over one afternoon it returned an HTTP 202 stub to scripted clients, then a 403, and then began redirecting every document URL, including ones that had served the full text an hour earlier, to the Official Journal index, where it displays:
EUR-Lex is temporarily not fully available.
That is an outage, not a block on any particular client, which is rather the point: the authority can be unavailable for reasons that have nothing to do with you, at a moment you did not choose. A skill that fetches the law when asked inherits that. A skill carrying the text, dated, with a command to re-check it when the source returns, does not.
If a quote drifts, it says where:
DRIFTED.../references/hwg.md (1 of 6 fragments no longer match)
quoted: "Unzulässig ist eine irrefuehrende Reklame. Eine Irreführung..."
on page up to: "...Unzulässig ist eine irref"
Exit 0 when every fetched source matches, 1 on drift, 2 if nothing could be fetched.
| Reference | Skill | Source | Retrieved |
|---|---|---|---|
hwg.md |
device-claims |
gesetze-im-internet.de, heilmwerbg |
2026-09-09 |
mdr-ivdr-art7.md |
device-claims |
EUR-Lex, Regulation (EU) 2017/745 (MDR), CELEX 32017R0745 | 2026-09-09 |
uwg.md |
device-claims |
gesetze-im-internet.de, uwg_2004 |
2026-09-09 |
annex-viii-software.md |
mdr-classification |
EUR-Lex, Regulation (EU) 2017/745, CELEX 32017R0745, Annex VIII | 2026-09-09 |
art120-amended.md |
mdr-transition |
Publications Office, Official Journal L 80, 20.3.2023, p. 24 | 2026-09-10 |
mpdg.md |
mpdg-germany |
gesetze-im-internet.de, Medizinprodukterecht-Durchführungsgesetz (MPDG) | 2026-09-10 |
Every one of them is re-fetched and diffed on each push and again every Monday, which is
what the verify badge reports. An earlier version of this table marked the two EU-Lex
references "not auto-verifiable. EUR-Lex blocks scripted clients". That was true until
the MDR and IVDR text was rerouted through the Publications Office, and the column stayed
wrong afterwards, understating what the repo checks. It is generated now.
The skills use two citation tiers: [verified] for provisions in those files, and
[verify] for anything else. They are instructed to refuse rather than supplement from
model knowledge.
Statute text only. No case law, no MDCG guidance, no notified-body practice, no national enforcement decisions. UWG in particular is heavily shaped by BGH and OLG case law this skill does not assess.
That limit is deliberate. Statutory text can be verified against a free primary source and re-checked by anyone; case law cannot, without paid research access. A tool that claims case-law coverage it cannot verify is worse than one that draws the line.
This is a drafting and risk-triage aid, not legal advice, and not a substitute for a Fachanwalt für Medizinrecht or Wettbewerbsrecht.
| If you are asking | Use | It carries |
|---|---|---|
| "Can we say this in our copy?" | device-claims |
MDR Art. 7, IVDR Art. 7, HWG, UWG |
| "What class is our software?" | mdr-classification |
Annex VIII impl. rules 3.1–3.7, Rule 11 |
| "Does Germany want more than MDR?" | mpdg-germany |
MPDG §§ 4, 8, 73 |
| "What should this report say it didn't check?" | scope-statement |
nothing, domain-general |
Each is independent. Install only what you need; together they cost about 660 tokens always-on.
claude plugin marketplace add Ecgecer/eu-mdr-skills
claude plugin install device-claims@eu-mdr-skills # advertising claims
claude plugin install mdr-classification@eu-mdr-skills # software classification
claude plugin install mdr-transition@eu-mdr-skills # Article 120 legacy transition
claude plugin install mpdg-germany@eu-mdr-skills # German additions to MDR
claude plugin install scope-statement@eu-mdr-skills # bound a compliance claim
/device-claims:device-claims-review # review advertising copy
/mdr-classification:software-classification # classify software under Annex VIII
/mdr-transition:legacy-transition # how long a legacy device may be sold
/mpdg-germany:german-additions # what Germany adds on top of MDR
/scope-statement:scope-statement # bound a compliance result
Each skill establishes its own anchors before answering, and asks rather than guesses.
device-claims wants the intended purpose as assessed and the audience (Fachkreise vs.
Publikum, which gates HWG § 11). mdr-classification wants the intended purpose and
whether the product is qualified as a device at all. mpdg-germany wants confirmation
the German market is actually in play.
The substance is plain markdown. Only the packaging is Claude-specific.
| Tool | What to use |
|---|---|
| Claude Code | Install the plugin (above) |
| Codex, Cursor, anything reading AGENTS.md | AGENTS.md. It names the load order |
| Gemini CLI | GEMINI.md |
| ChatGPT, Gemini web, Claude.ai, any chat | Upload or paste the matching file in dist/, each bundles one skill with all its references |
| Anything else | The four source files in device-claims/skills/device-claims-review/ |
GEMINI.md and the bundle are generated from the canonical skill:
python3 scripts/build-portable.py # rebuild
python3 tests/check-portable-fresh.py # fail if stale
Never edit them by hand. The freshness check exists because a drifted bundle would have someone reviewing against an older rule while the repo claimed otherwise, the exact failure this project is meant to prevent.
AGENTS.md lists the five invariants any port must preserve.
Seven cases in device-claims/evals/, including three
false-positive controls:
| Case | Checks |
|---|---|
01-limb-d-intended-purpose-drift |
Catches a true claim that breaches Art. 7(d) |
02-hwg11-wrong-audience |
Does not apply § 11 to a Fachkreise audience |
03-hwg11-item-scope |
Does not cite § 11(1) no. 2 against a device (only nos. 7-9, 11, 12 apply) |
04-limb-c-omission |
Catches an omission where nothing is false |
05-uwg6-comparison |
Treats comparative advertising as lawful-if-compliant, not banned |
06-no-case-law-supplement |
Refuses to state a BGH holding as fact |
07-clean-copy-control |
Reports no findings on clean copy rather than inventing one |
claude plugin eval ./device-claims runs them where the eval harness is enabled.
Apache-2.0. The review workflow, claim-block format, citation tiering and non-lawyer
approval gate derive from
anthropics/claude-for-legal
(product-legal/skills/marketing-claims-review), Apache-2.0. See NOTICE for
what was changed.