Stars
- All languages
- ANTLR
- Arduino
- Assembly
- Batchfile
- BibTeX Style
- Bluespec
- C
- C#
- C++
- CMake
- CSS
- Common Lisp
- Cython
- Dockerfile
- Emacs Lisp
- Go
- HTML
- Haskell
- Java
- JavaScript
- Jupyter Notebook
- Just
- Kotlin
- Lean
- Lua
- MATLAB
- Makefile
- Nim
- PHP
- Pascal
- Pony
- PowerShell
- Python
- QML
- R
- Riot
- Roff
- Ruby
- Rust
- Scala
- Shell
- Swift
- SystemVerilog
- Tcl
- TeX
- TypeScript
- Typst
- V
- VHDL
- Vala
- Verilog
- Yacc
- Zig
fossi-foundation / gf180mcu-pdk
Forked from google/gf180mcu-pdkPDK for GlobalFoundries' 180nm MCU bulk process technology (GF180MCU).
DeepSeek Harness: Everything is a Plugin.
Open-source MCP server that lets an AI drive a live session of Altium Designer, and optionally KiCad or EasyEDA Pro, editing the open design in place while you watch. 300+ tools: schematic, PCB, li…
DeepSeek-V4-Flash-0731 on 4x CMP 170HX (sm_80): 98 tok/s decode, ~5300 tok/s prefill. DSpark speculative decoding under pipeline parallelism, which vLLM does not support upstream.
The most comprehensive community resource for the NVIDIA CMP 170HX.
terminatorul / NvStrapsReBar
Forked from xCuri0/ReBarUEFIResizable BAR for Turring GTX 1600 / RTX 2000 GPUs
Silicon-proven INT8 systolic NPU (8×8 MAC array) taped out on SkyWater 130nm via LibreLane. Features a custom 32-bit ISA, UART–APB host interface, and fused streaming datapath. Validated on chest X…
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Skills for Real Engineers. Straight from my .agents directory.
Pi extension: a persistent second model that reviews the main agent's work each turn and injects concise advice inline.
Implement a simple PDM to PCM conversion on an FPGA platform.
A minimal, open-source Image Signal Processor (ISP) for AMD FPGA, implemented in Verilog.
📚 Novle setting | 小说书源及软件整理 爱阅书香 / 香色闺阁 / 阅读(含字体、净化规则、TTS配置)
Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode
Fast and memory-efficient exact attention (for turing architecture only)
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Flash Attention 2 implementation for Turing GPUs
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, su…
Academic Research Skills for Claude Code: research → write → review → revise → finalize
A machine learning accelerator core designed for energy-efficient AI at the edge.
Open-source 24/7 Cowork app for OpenClaw, Hermes, Claude Code, Codex, OpenCode and 20+ more CLI Agent | Customize your assistants | Team them up|Star if you like it!
Open-source, community-driven agent harness
LLM inference in C/C++ optimized for Strix Halo
A high-performance and light-weight router for vLLM large scale deployment
An open-source, JEDEC JESD270-4A-compliant HBM4 memory subsystem (controller + PHY-shim + DFT + RAS + security wrapper) tightly coupled to an open RISC-V-native LPU accelerator. Apache-2.0 RTL, CER…
An OpenSkills agent skill for SpinalHDL development with correctness-first guidance from official docs and VexRiscv.
AI-agent Skill for generating polished HTML slide decks: editorial magazine and Swiss layouts, image prompts, social covers, and a WebGL/low-power presentation runtime.
AI academic PPT workflow skill for Codex / ChatGPT: paper-to-editable PowerPoint, planning table, mockup family, and PPTX generation.
A terminal workspace with batteries included
The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-request decode with support of FP8 weight