Skip to content
View gmittal's full-sized avatar

Block or report gmittal

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results
Swift 14,324 1,689 Updated Sep 24, 2026

LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.

C++ 2,432 137 Updated Sep 23, 2026

Doing simple retrieval from LLM models at various context lengths to measure accuracy

Jupyter Notebook 2,385 246 Updated Jun 8, 2026

Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.

Python 10,740 1,103 Updated Jul 1, 2024

A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).

Python 910 51 Updated Sep 30, 2025

MLX: An array framework for Apple silicon

C++ 28,537 2,276 Updated Sep 23, 2026

The official Porsche Design System repository, offering fundamental UXI guidelines and a library of reusable web components to enable designers and developers to build consistent, intuitive, and hi…

TypeScript 658 63 Updated Sep 23, 2026
Python 32 7 Updated Jan 9, 2025

A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.

Python 2,295 251 Updated Oct 16, 2025

Inference Llama 2 in one file of pure 🔥

Mojo 2,127 139 Updated Sep 20, 2026

Python pdb for multiple processes

Python 83 9 Updated May 24, 2025

The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.

Python 9,013 631 Updated May 3, 2024

Inference code for CodeLlama models

Python 16,249 1,934 Updated Aug 12, 2024

[EMNLP 2022] Training Language Models with Memory Augmentation https://arxiv.org/abs/2205.12674

Python 192 13 Updated Jun 14, 2023

Supercharge Your LLM Application Evaluations 🚀

Python 15,835 1,725 Updated Feb 24, 2026

A dataset of 100,000 minutes of driving video compressed using a VQ-VAE.

Jupyter Notebook 378 80 Updated Aug 8, 2026

Tools for building GPU clusters

Shell 1,475 362 Updated Sep 23, 2026

LLMs for your CLI

Python 1,375 79 Updated May 29, 2024

A Data Streaming Library for Efficient Neural Network Training

Python 1,562 206 Updated Jun 25, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 92,588 22,597 Updated Sep 24, 2026

The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.

Python 65,930 7,748 Updated Sep 23, 2026

Write scalable load tests in plain Python 🚗💨

Python 28,179 3,248 Updated Sep 21, 2026

CUDA on non-NVIDIA GPUs

Rust 14,877 941 Updated Sep 22, 2026

It's React, but in Python

Python 8,147 332 Updated Jul 14, 2026

Implementation of Flash Attention in Jax

Python 230 25 Updated Mar 1, 2024

The official implementation of “Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training”

Python 1,000 57 Updated Jan 30, 2024

A simulation framework for RLHF and alternatives. Develop your RLHF method without collecting human data.

Python 844 66 Updated Jul 1, 2024

Tevatron - Unified Document Retrieval Toolkit across Scale, Language, and Modality. Demo in SIGIR 2023, SIGIR 2025.

Python 750 131 Updated Jul 18, 2026

The Modular Platform (includes MAX & Mojo)

Mojo 29,873 3,182 Updated Sep 23, 2026

The Official Python Client for Lamini's API

Python 2,533 151 Updated Apr 7, 2025
Next