Skip to content

Repository files navigation


  ███╗   ███╗ ██████╗ ██████╗ ███████╗██████╗  █████╗ ████████╗ ██████╗ ██████╗
  ████╗ ████║██╔═══██╗██╔══██╗██╔════╝██╔══██╗██╔══██╗╚══██╔══╝██╔═══██╗██╔══██╗
  ██╔████╔██║██║   ██║██║  ██║█████╗  ██████╔╝███████║   ██║   ██║   ██║██████╔╝
  ██║╚██╔╝██║██║   ██║██║  ██║██╔══╝  ██╔══██╗██╔══██║   ██║   ██║   ██║██╔══██╗
  ██║ ╚═╝ ██║╚██████╔╝██████╔╝███████╗██║  ██║██║  ██║   ██║   ╚██████╔╝██║  ██║
  ╚═╝     ╚═╝ ╚═════╝ ╚═════╝ ╚══════╝╚═╝  ╚═╝╚═╝  ╚═╝   ╚═╝    ╚═════╝ ╚═╝  ╚═╝

BYOK AI content moderation microservice. Bring your own API key, plug it in, get a verdict on every message. That simple.

Python FastAPI PostgreSQL Docker Claude OpenAI Gemini pytest


Architecture

01-diagram-export-6-10-2026-4_27_00-PM

What it does

Any app sends POST /moderate with a user ID and message. The service runs it through your connected AI provider (Claude, GPT-4o, or Gemini), returns a verdict, and tracks repeat offenders across a persistent strike pipeline. Five strikes and the user is permanently flagged — no more moderating them, no more wasted inference calls.

The key design decision: you supply the API key. The service encrypts it with Fernet symmetric encryption and stores it server-side. Your AI usage stays on your billing account. The service just handles the routing, session management, strike tracking, and admin tooling.


Technical decisions worth knowing

Provider abstraction (Strategy pattern) — all three AI providers implement a single AIProvider ABC with one method: moderate(message) → ModerationResult. At request time, provider_factory.py resolves the correct provider from the user's stored credentials, decrypts the key, and hands back the right implementation. Swapping providers doesn't touch the moderation logic.

Fernet encryption for stored keys — provider API keys are never stored in plaintext. Every key goes through cryptography.Fernet before hitting the database. Decryption only happens inside the request cycle, never logged, never returned.

JWT + server-side session table — access tokens are signed with HS256 but also hashed (SHA-256) and stored in a sessions table. Logout physically revokes the hash, so a stolen token can't be replayed after logout. Standard JWT-only auth doesn't give you this.

Strike pipeline designUserViolation tracks strike count per app per user. ViolationLog records every individual flagged message with full context. These are separate tables intentionally — you can query aggregate counts without scanning the full log, and you can audit the full log without touching counters.

Structured output enforcement — the system prompt forces JSON-only responses and strips markdown fences in all three provider implementations. The fallback on parse failure is safe=True (fail open) to avoid false positives on provider glitches.


Quick start

git clone https://github.com/kisugez/moderator.git
cd moderator
cp .env.example .env
# add ANTHROPIC_API_KEY, ADMIN_SECRET, FERNET_SECRET, JWT_SECRET to .env
docker compose up --build

Generate the secrets you need:

# FERNET_SECRET
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"

# JWT_SECRET
python -c "import secrets; print(secrets.token_hex(32))"

API

Method Path Auth Description
GET /health None Service health check
POST /moderate X-API-Key Moderate a message, track strikes
GET /admin/violations X-Admin-Secret Full violation log with filters
GET /admin/users X-Admin-Secret User strike records
PATCH /admin/users/{id}/unban X-Admin-Secret Reset user strikes
POST /auth/register None Create account
POST /auth/login None Get JWT token
POST /providers/connect Bearer JWT Encrypt + store provider key
GET /providers/me Bearer JWT List connected providers

Moderate a message

curl -X POST http://localhost:8000/moderate \
  -H "Content-Type: application/json" \
  -H "X-API-Key: mod_live_yourkey" \
  -d '{"user_id": "user-42", "message": "Hello world"}'

Safe response:

{"safe": true}

Strike response:

{
  "safe": false,
  "strike_count": 2,
  "flagged": false,
  "reason": "threat / violent language",
  "severity": "high",
  "warning": "Warning 2/5: threat / violent language"
}

Permanently flagged:

{
  "safe": false,
  "flagged": true,
  "warning": "Your account has been permanently restricted."
}

Strike system

Count State Behaviour
1–4 Active Warning returned, message blocked, strike recorded
5 Flagged Account permanently restricted — instant block on all future messages, no inference call made

The permanently-flagged check runs before the AI call. A user at strike 5 costs you zero tokens.


mod-cli

An interactive terminal CLI ships with the service — no Docker needed for local testing.

download

Full CLI reference in docs/quickstart.md.


Tests

# local
pytest --cov=app -v

# full docker integration
docker compose -f docker-compose.test.yml up --build --abort-on-container-exit

Test coverage includes auth flows (register → OTP verify → login → logout), strike accumulation and flagging logic, provider auth and Fernet secret handling, and API key lifecycle.


Project structure

moderator/
├── app/
│   ├── main.py                  # FastAPI app, lifespan, router registration
│   ├── config.py                # Pydantic settings, env validation
│   ├── database.py              # Async SQLAlchemy engine + session factory
│   ├── middleware/auth.py       # API key, JWT, admin secret validators
│   ├── models/
│   │   ├── db.py                # SQLAlchemy ORM models
│   │   └── schemas.py           # Pydantic request/response schemas
│   ├── routers/
│   │   ├── moderate.py          # POST /moderate — core pipeline
│   │   ├── admin.py             # Admin endpoints
│   │   ├── auth.py              # Register, verify OTP, login, logout
│   │   └── providers.py         # Provider connect/disconnect/switch
│   └── services/
│       ├── ai.py                # AIProvider ABC + ModerationResult
│       ├── claude.py            # Anthropic implementation
│       ├── openai.py            # OpenAI implementation
│       ├── gemini.py            # Gemini implementation
│       ├── provider_factory.py  # Runtime provider resolution
│       ├── crypto.py            # Fernet encrypt/decrypt
│       ├── jwt_service.py       # Token creation, hashing, validation
│       ├── violations.py        # Strike pipeline helpers
│       └── email.py             # OTP delivery (SMTP or stdout fallback)
├── migrations/                  # Alembic migrations (versioned)
├── tests/                       # pytest suite
├── mod_cli.py                   # Interactive developer CLI
├── docker-compose.yml
├── docker-compose.test.yml
└── docs/
    ├── quickstart.md

About

BYOK AI content moderation API. Plug in your own AI key, moderate messages, track repeat offenders with a 5-strike pipeline. FastAPI + PostgreSQL + Docker.

Topics

Resources

Stars

47 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages