Lists (10)
Sort Name ascending (A-Z)
Starred repositories
Real-time system audio translation for macOS — translate any audio (YouTube, podcasts, meetings) live on screen with OpenAI or Google Gemini. On-device speech recognition; optional low-latency real…
A framework for generating realistic LLM serving workloads
Code search MCP for Claude Code. Make entire codebase the context for any coding agent.
A reactive notebook for Python — run reproducible experiments, query with SQL, execute as a script, deploy as an app, and version with git. Stored as pure Python. All in a modern, AI-native editor.
GitHub Copilot extension for JupyterLab
🪐 🔧 Model Context Protocol (MCP) Server for Jupyter.
A CLI tool to switch and manage Codex accounts
Implementing DeepSeek R1's GRPO algorithm from scratch
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
Official implementation of ARDY: Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation (SIGGRAPH 2026).
A high-performance and light-weight router for vLLM large scale deployment
FlashKDA: high-performance Kimi Delta Attention kernels
aiha-lab / pim-iree
Forked from aiha-lab/ireeCompiler and runtime implementation for PIM device.
A retargetable MLIR-based machine learning compiler and runtime toolkit.
This repository is to provide graph mode execution with GPU+PIM in runtime.
Anthropic's original performance take-home, now open for you to try!
Build a compiler to solve Anthropic's interview challenge.
SW Library for Samsung PNM (including functional simulator)
🛰️ Track token usage across AI coding agents from your terminal. 🏅 Global leaderboard with trillions of tokens tracked.
An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing coding agent traces.
vLLM wall-clock emulation pluging to allow running the real vllm serve and benchmark without gpu with simple mocking