Launched July 28, 2026
SDKProof

SDKProof

How ready is your SDK for AI coding agents?

FreeAI AgentsAI Coding2,900 impressions#5 of its week2 comments

Comments

>log in to comment
  • Kalpit Rathore[maker]· 2mo ago

    @zain Thanks, & that's the direction I want to take it. Right now I only have the other axis, and half by accident. When Opus 5 shipped I re-ran the whole board & two libraries jumped 90 to 100 (Vercel AI SDK 7 and Zod 4). So I can watch the score move per model, but not yet per library version. Per-version is the more useful one for teams, I agree. That's the version where you can point at a specific upgrade & say this is what it cost you. It needs me to keep the old lockfiles around & re-run on each major, which I'm not doing yet but should be. One question since you're clearly thinking about this properly. Would you want it as a history on the library page, or as something that tells you when it moves? Those are pretty different builds & I'd rather build the one people actually want.

  • Zain Sheikh· 3mo ago

    Letting the compiler decide instead of an AI judging another AI is a clean signal. Do you plan to track scores per library version over time so teams can see when an upgrade breaks AI accuracy?

AI coding assistants keep writing your library's old API. When a library ships a new major version, the model still writes the previous version's code from memory — it looks right, but it doesn't compile against the version you actually have installed. SDKProof measures exactly how often that happens. It gives an AI model real coding tasks for a library, then type-checks the generated code against the real installed package. Compiles = pass. No AI judging another AI — the compiler decides. Scores so far: Prisma 80, Vercel AI SDK 90, Zod 90, Next.js 92, TanStack Query 100. The pattern: the more recently a library changed, the more the AI gets wrong. Open source — run it on any TypeScript package.

SDKProof evaluates how well AI coding assistants handle updated TypeScript library APIs by compiling generated code against real packages.

for
Developers building or testing AI coding assistants for TypeScript libraries.
pricing
open source
license
MIT
Kalpitrathore/sdkproof 5 0HTMLupdated 2 months ago

Key features

6 features of SDKProof
  • Drift detection — Compares two published npm versions and shows API changes without installing anything.
  • Real coding tasks — Runs AI models on library-specific tasks and feeds results to the TypeScript compiler.
  • Compiler-based scoring — Counts a task as passed only if the generated code compiles with tsc.
  • CLI interface — Run with a single npx command, e.g., npx sdkproof <package>.
  • Open-source — The tool is free and can be run on any TypeScript package.
  • Scorecards per library — Shows pass rates for libraries like Prisma, Next.js, TanStack Query, etc.

Use cases

  • Validate an AI assistant’s output against the latest version of a TypeScript SDK
  • Detect breaking API changes that cause AI-generated code to fail
  • Generate a scorecard for multiple libraries to compare model performance
  • Automate regression testing of AI-written code after a major library release

SDKProof FAQ

How does SDKProof decide if an AI answer is correct?+

It runs the generated TypeScript code through the real installed package and the TypeScript compiler; only compiled code counts as a pass.

Do I need an API key or install the target library?+

No, SDKProof reads package versions from npm and performs type-checking without requiring an API key or manual installation.

What command runs the evaluation?+

Use the CLI: npx sdkproof <package> or npx sdkproof driftPackageName to compare versions.

Which AI model was used for the published scores?+

Claude Opus 5 was used for the example tasks and scorecards.

Is SDKProof limited to specific libraries?+

It works with any TypeScript package published on npm; the site shows examples for eight libraries.

Summarized by DevHunt from sdkproof.dev · Sep 28, 2026. Details may change; check the official site.