- Bangalore
- in/sagar-sumit
Stars
Michelangelo AI: Uber's end-to-end machine learning platform.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
Apache Fluss is a streaming storage built for real-time analytics.
Monitoring and insights on your data lakehouse tables
The native Rust implementation for Apache Hudi, with C++ & Python API bindings.
A repository where sample Hudi code, tips/tricks, etc. will be hosted.
Apache XTable (incubating) is a cross-table converter for lakehouse table formats that facilitates interoperability across data processing systems and query engines.
A composable and fully extensible C++ execution engine library for data management systems.
All the things about TPC-DS in Apache Spark
The official home of the Presto distributed SQL query engine for big data
Notes on books I read, talks I watch, articles I study, and papers I love
Upserts, Deletes And Incremental Processing on Big Data.
Open source, privacy focused client side library for the creation and monetisation of online audiences.
Papers & presentation materials from Hugging Face's internal science day
A list of NLP(Natural Language Processing) tutorials
Companion webpage to the book "Mathematics For Machine Learning"
Google Cloud Platform Certification resources.
NLP 101: a resource repository for Deep Learning and Natural Language Processing
Thinking in tensors, writing in PyTorch (a hands-on deep learning intro)
Financial Sentiment Analysis with BERT
The "Python Machine Learning (1st edition)" book code repository and info resource
⛔️ DEPRECATED – See https://github.com/ageron/handson-ml3 or handson-mlp instead.
Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.
comparing stand up comedians using natural language processing
Data science teaching materials