Skip to content
View apolmig's full-sized avatar
👨‍🏭
👨‍🏭

Block or report apolmig

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Agent skill for impressive 3D visuals using Blender + image gen + subagent critic

JavaScript 1,541 165 Updated Sep 9, 2026

METR Task Standard

TypeScript 197 38 Updated Feb 3, 2025

Optimal stopping package for LLM evaluations

Python 21 4 Updated Sep 13, 2026

Professional video production workflows for Codex—from story brief and media analysis through editorial, audio, captions, motion graphics, Adobe automation, QC, and delivery.

Python 6 1 Updated Aug 28, 2026

Reproducibility code and ten-run results for CentaurBench: augmenting vs. automating real-world work tasks (DIAL, UC Berkeley Haas).

Jupyter Notebook 3 Updated Aug 27, 2026

Collection of evals for Inspect AI

Python 682 445 Updated Sep 24, 2026

Local-first LLM summarization and evaluation workbench with BYOK privacy controls

TypeScript 3 1 Updated May 5, 2026

Finance terminal, in your terminal.

TypeScript 2,208 137 Updated Sep 24, 2026

Run Inspect AI evals in the cloud

Python 79 50 Updated Sep 24, 2026

DeepSeek Harness: Everything is a Plugin.

TypeScript 235,042 28,275 Updated Sep 24, 2026

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Python 16,818 3,256 Updated Sep 24, 2026

An alignment auditing agent capable of quickly exploring alignment hypothesis

Python 1,346 228 Updated Sep 19, 2026

bloom - evaluate any behavior immediately  🌸🌱

Python 1,408 171 Updated May 7, 2026

Repository for "User awareness in frontier models: who's asking shifts what models say"

Python 14 5 Updated Aug 6, 2026

Framework for generating behavioral evaluations of frontier AI models.

Python 34 6 Updated Jul 15, 2026

Experimental implementation of DeepSeek v4 flaash in llama.cpp

C++ 24 4 Updated Apr 30, 2026
Python 30 7 Updated Jan 7, 2026

🪄 Flint is a visualization language that lets AI agents reliably create expressive, good-looking charts from simple, human-editable chart specs.

TypeScript 4,276 238 Updated Sep 23, 2026

A minimal autonomous code-agent CLI built with Qwen Code

TypeScript 829 86 Updated Aug 12, 2026

ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities designed to evaluate AI agents' ability to develop exploits.

C 1,088 139 Updated Aug 6, 2026

Open-source observability tool that uses AI agents to self-heal your software

TypeScript 1,452 117 Updated Sep 24, 2026

Build your own AI SRE agents. The open source toolkit for the AI era.

Python 11,221 1,642 Updated Sep 24, 2026

Secure environments for developers and their agents

Go 16,674 1,586 Updated Sep 24, 2026

Fully automatic censorship removal for language models

Python 32,297 3,629 Updated Sep 22, 2026

AgentENV (AENV) is a distributed platform for running agent environments at scale.

Rust 3,534 318 Updated Sep 24, 2026

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

Python 1,153 134 Updated Sep 20, 2026

ControlArena is a collection of settings, model organisms and protocols - for running control experiments.

Python 241 141 Updated Aug 24, 2026

Inspect: A framework for large language model evaluations

Python 2,854 751 Updated Sep 24, 2026
Next