Stars
RIFT: RubrIc Failure mode Taxonomy, a taxonomy for systematically characterizing failure modes in rubric composition and design
NEJM-Bench: A Benchmark for Evaluating LLMs on Clinical Image Challenges
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Implementation of OpenAI's HealthBench evaluation framework
Implementation of OpenAI’s FrontierScience benchmark for evaluating LLM performance on expert-level scientific reasoning tasks.
Medical Sphere https://medicalsphere.ai/
Label Studio is a multi-type data labeling and annotation tool with standardized output format
Acceptance rates for the major AI conferences
Conference schedule, top papers, and analysis of the data for NeurIPS 2023!
⚡ A Fast, Extensible Progress Bar for Python and CLI
Data and software for building the ACL Anthology.
Query your Apple Health data with natural language 💬 🩺
Foundation Models for Genomics & Transcriptomics
A collection of fact-checked news articles
Python port of the R Bioconductor `seqLogo` package
MIMIC Code Repository: Code shared by the research community for the MIMIC family of databases
A command line utility to convert 23andme raw data files to VCF format
Does Causal Coherence Predict Online Spread of Social Media? (Presented at SBP-BRiMS 2019)
A very simple framework for state-of-the-art Natural Language Processing (NLP)
Repository for Causal News Corpus (LREC 2022) and RECESS (IJCNLP-AACL 2023)