Skip to content
View elllusion's full-sized avatar

Block or report elllusion

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Self-hosted native-Rust runtime for real-time voice agents. Own the stack: one binary in your own VPC or air-gapped, no hosted control plane. pipecat-compatible pipeline, in-process SIP/RTP, single…

Rust 92 10 Updated Aug 3, 2026

GLM-5.2, a 744 billion parameter mixture of experts model, in a pure C inference engine: quantized to int4, experts streamed from disk, deployed and benchmarked. Generates in 16 GB of RAM.

C 35 10 Updated Jul 15, 2026

A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

C 4,218 658 Updated Aug 7, 2026

⚡ Co-optimized LLM compression, static/dynamic runtime acceleration, and agent harness tuning for memory-constrained on-device agents.

TeX 4 1 Updated Aug 9, 2026

AI coding platform for teams

TypeScript 4,296 625 Updated Aug 8, 2026

The most RAM efficient harness

Rust 16,635 1,876 Updated Aug 10, 2026

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python depend…

C++ 1,248 163 Updated Aug 9, 2026

The local UI to run and train text and diffusion models, including Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, FLUX and more.

Python 69,774 6,298 Updated Aug 10, 2026

Gemma open-weight LLM library, from Google DeepMind

Python 5,641 1,009 Updated Aug 5, 2026

Generator Bootcamp Material: Learn Chisel the Right Way

Jupyter Notebook 1,148 314 Updated Sep 10, 2024

A template project for beginning new Chisel work

Shell 707 204 Updated Feb 24, 2026

llama.cpp fork with additional SOTA quants and improved performance

C++ 3,022 409 Updated Aug 9, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 31,590 7,763 Updated Aug 10, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,211 1,513 Updated Aug 8, 2026

An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a high-performance GPU ke…

Python 71 23 Updated Aug 9, 2026

Contract-Aware RTL Code Generation Agents with Temporal Tracing, Slicing and Formal Verification

SystemVerilog 23 4 Updated Jan 28, 2026

ARM64 ELF Virtual Machine Protection System

Go 477 170 Updated Mar 26, 2026

AI 驱动的多领域安全分析 Agent 平台 —— 让 LLM 端到端完成安全分析

Python 56 10 Updated Jul 29, 2026

中国专利.skill:从项目文档到交底书编写(挖点·查新·脱敏成文)+ 专利通俗解读(叙事·图谱·Obsidian 私库)。

Python 4,839 623 Updated Jul 24, 2026

FSA: Fusing FlashAttention within a Single Systolic Array

Scala 190 19 Updated Apr 15, 2026

A paper list of spiking neural networks, including papers, codes, and related websites. 本仓库收集脉冲神经网络相关的顶会顶刊以及CNS论文和代码,正在持续更新中。

814 78 Updated Mar 24, 2026

Backward compatible ML compute opset inspired by HLO/MHLO

MLIR 683 214 Updated Aug 5, 2026

The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.

C++ 1,881 721 Updated Aug 6, 2026

An open-source, JEDEC JESD270-4A-compliant HBM4 memory subsystem (controller + PHY-shim + DFT + RAS + security wrapper) tightly coupled to an open RISC-V-native LPU accelerator. Apache-2.0 RTL, CER…

SystemVerilog 22 3 Updated Jun 14, 2026

Framework providing operating system abstractions and a range of shared networking and memory services for common modern heterogeneous platforms.

SystemVerilog 441 116 Updated Aug 3, 2026

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

C 23,519 2,533 Updated Aug 9, 2026

A custom C++ routine to identify logic gates in the layout extracted netlist (SPICE) of digital circuits and generate gate-level Verilog netlist, in the presence of logic gate defintions from the s…

C++ 34 12 Updated Aug 21, 2024
Python 172 52 Updated Dec 4, 2022
Next