Stars
A no-fluff and highly practical masterclass that reignites engineering curiosity and helps SDE-2, SDE-3, and above become great at designing, implementing, and shipping scalable, fault-tolerant, an…
Light, fluffy, and always free - The AWS Local Emulator alternative
The open-source communication infrastructure for agents and products
An ETL Data Pipelines Project that uses AirFlow DAGs to extract accessories and jewelry data from PostgreSQL Schemas and the shoes data from a CSV file, load them in AWS Data Lake, transform them w…
Open-source platform for creating safe, isolated production sandboxes for API, integration, and E2E testing.
Practical Hands on Exercises for MLOps
Open-source, secure environment with real-world tools for enterprise-grade agents.
[Paper List] Papers integrating knowledge graphs (KGs) and large language models (LLMs)
Replace 'hub' with 'ingest' in any GitHub URL to get a prompt-friendly extract of a codebase
A JSON-like data structure (a CRDT) that can be modified concurrently by different users, and merged again automatically.
ERDDAP is a scientific data server that gives users a simple, consistent way to download subsets of gridded and tabular scientific datasets in common file formats and make graphs and maps. ERDDAP i…
Curated list of project-based tutorials
The Patterns of Scalable, Reliable, and Performant Large-Scale Systems
All notes and materials for the CS229: Machine Learning course by Stanford University
Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
An MCP server using the AviationStack API to fetch real-time flight data including airline flights, airport schedules, future flights and aircraft types
Virtual whiteboard for sketching hand-drawn like diagrams
Open source, composable payments platform | PCI compliant | SaaS and Self-host options | Enables connectivity to multiple payment, payout, fraud, vault and tokenization providers | Uplifts authoriz…
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
The official home of the Presto distributed SQL query engine for big data
Data science Python notebooks: Deep learning (TensorFlow, Theano, Caffe, Keras), scikit-learn, Kaggle, big data (Spark, Hadoop MapReduce, HDFS), matplotlib, pandas, NumPy, SciPy, Python essentials,…
This Guidance demonstrates how to deploy a generative artificial intelligence (AI) model provided by Amazon SageMaker JumpStart to create an asynchronous SageMaker endpoint with the ease of the AWS…
This Guidance shows how Amazon Bedrock Data Automation streamlines the generation of valuable insights from unstructured multimodal content such as documents, images, audio, and videos through a un…