Lists (4)
Sort Name ascending (A-Z)
Stars
Ontology-anchored, git-native data modeling for dbt-core teams
Practice Databricks coding skills with hands-on exercises. Import into Databricks Free Edition, write code, run assertions, check pass/fail. Covers Delta Lake, Spark SQL, PySpark, Auto Loader, meda…
Difference-in-Differences causal inference in Python. Callaway-Sant'Anna, Synthetic DiD, Honest DiD, event studies. sklearn-like API, validated against R.
Marketing skills for Claude Code and AI agents. CRO, copywriting, SEO, analytics, and growth engineering.
This repository shares some usefull information about A/B testing and entry level knowledge.
Lists of company wise questions. Every csv file in the companies directory corresponds to a list of questions on leetcode for a specific company based on the leetcode company tags. Updated as of 20…
End-to-end Data Lakehouse project built on Databricks, following the Medallion Architecture (Bronze, Silver, Gold). Covers real-world data engineering and analytics workflows using Spark, PySpark, …
LTVision is an open-source library from Meta, designed to empower businesses to unlock the full potential of predicted customer lifetime value (pLTV) modeling.
Meridian is an MMM framework that enables advertisers to set up and run their own in-house models.
Render real-time, dynamic charts in LiveView applications
A schema for defining and standardising behavioural tracking events based on UX components in any digital product.
Free, simple, and intuitive online database diagram editor and SQL generator.
1 Line of code data quality profiling & exploratory data analysis for Pandas and Spark DataFrames.
Turns Data and AI algorithms into production-ready web applications in no time.
This repository contains everything you need to become proficient in Scikit learn
Deploying Anomaly Detection model in BigQuery for GA4 Data
Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Join the course here 👇🏼
ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.
Machine Learning From Scratch. Bare bones NumPy implementations of machine learning models and algorithms with a focus on accessibility. Aims to cover everything from linear regression to deep lear…
A scikit-learn-compatible library for estimating prediction intervals and controlling risks, based on conformal predictions.
LightweightMMM 🦇 is a lightweight Bayesian Marketing Mix Modeling (MMM) library that allows users to easily train MMMs and obtain channel attribution information.
Time series easier, faster, more fun. Pytimetk.
Explain complex systems using visuals and simple terms. Help you prepare for system design interviews.
Replace Splunk in your small company with this one weird trick!
Statistical Rethinking Course for Jan-Mar 2023
Pre-built metrics for Stripe data. Check out metrics for other SaaS tools at https://hub.houseware.io/
Open-source, low-code AutoML platform for Python. PyCaret 4.0: sklearn-native engine + React control plane.