Highlights
Lists (1)
Sort Name ascending (A-Z)
- All languages
- Apex
- Astro
- Bikeshed
- Blade
- C
- C#
- C++
- CSS
- Cuda
- Dart
- Dockerfile
- Elixir
- GDScript
- Go
- Go Template
- HCL
- HTML
- Java
- JavaScript
- Julia
- Jupyter Notebook
- Kotlin
- Lean
- Lua
- MATLAB
- MDX
- Makefile
- Markdown
- Objective-C
- PHP
- PLSQL
- PLpgSQL
- Perl
- PostScript
- PowerShell
- Pug
- Python
- QML
- R
- Ren'Py
- Rich Text Format
- Ruby
- Rust
- SCSS
- SMT
- Sass
- Scala
- ShaderLab
- Shell
- Solidity
- Svelte
- Swift
- TSQL
- TeX
- TypeScript
- Vue
- Zig
- templ
Starred repositories
Agent skill for impressive 3D visuals using Blender + image gen + subagent critic
Optimal stopping package for LLM evaluations
Professional video production workflows for Codex—from story brief and media analysis through editorial, audio, captions, motion graphics, Adobe automation, QC, and delivery.
Reproducibility code and ten-run results for CentaurBench: augmenting vs. automating real-world work tasks (DIAL, UC Berkeley Haas).
Collection of evals for Inspect AI
Local-first LLM summarization and evaluation workbench with BYOK privacy controls
DeepSeek Harness: Everything is a Plugin.
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
An alignment auditing agent capable of quickly exploring alignment hypothesis
bloom - evaluate any behavior immediately 🌸🌱
Repository for "User awareness in frontier models: who's asking shifts what models say"
Framework for generating behavioral evaluations of frontier AI models.
Experimental implementation of DeepSeek v4 flaash in llama.cpp
🪄 Flint is a visualization language that lets AI agents reliably create expressive, good-looking charts from simple, human-editable chart specs.
A minimal autonomous code-agent CLI built with Qwen Code
ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.
Open-source observability tool that uses AI agents to self-heal your software
Build your own AI SRE agents. The open source toolkit for the AI era.
Secure environments for developers and their agents
Fully automatic censorship removal for language models
AgentENV (AENV) is a distributed platform for running agent environments at scale.
MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts
ControlArena is a collection of settings, model organisms and protocols - for running control experiments.
Inspect: A framework for large language model evaluations