Public healthcare data that's usually unusable — parsed, joined, and queryable. Look it up, or ask Claude.
trove builds open-source lookup tools, parsers, and Claude skills on top of public-domain healthcare datasets that are widely cited but rarely usable in their raw form. The full data, parsers, and skills are MIT-licensed at github.com/cbetz/trove.
New · findingHospitals report charity care to CMS and the IRS. The numbers don't match.
Yale New Haven told the IRS $113.1M and CMS $35.6M — same hospital, same year. Of the 99 hospitals where the comparison is clean, 22 differ by more than 50%. Read the finding →
Hospital reporting
See what 1,295 nonprofit U.S. hospital systems told CMS (Worksheet S-10) vs. the IRS (990 Schedule H) about charity care, side by side — ranked by the size of the gap, with a link to each hospital's actual 990.
FDA · 2021–2024FDA drug approvals
Look up any FDA novel drug approval from 2021–2024 and find the approval package — sponsor, application number, key dates, the FDA-approved label, and a deep link to every document FDA released (medical review, statistical review, pharmacology, chemistry). Then ask Claude to read them.
More areas coming. Suggestions welcome.
Each area ships with a Claude Code skill that translates natural-language questions into queries over the area's published data. The skills are bundled as a single Claude Code plugin in Anthropic's community marketplace:
/plugin marketplace add anthropics/claude-plugins-community
/plugin install trove@claude-community
Listed in Anthropic's community plugin marketplace.
Skill details and example prompts:
Public-domain healthcare data is famously messy. CMS publishes 100,000+ row long-skinny CSVs; the IRS publishes 990s as XML in bulk ZIPs; the FDA scatters approval reviews across hundreds of PDF directories. trove's job is to do the parsing, joining, and packaging so the data is browsable and queryable rather than something only people with a Python environment and free time can use.
Each area is also a Claude skill — meaning you can install it, ask questions in natural language, and get answers grounded in the actual underlying data rather than what an LLM half-remembers from training.
Plain-language explainers for the datasets: