Skip to content
#

ai-alignment-research

Here are 18 public repositories matching this topic...

Closed-loop Architecture Designed to Establish Self-governing, Mathematically Predictable, and Inherently Safe Super AI by mirroring the elegant physics of the cosmos.

  • Updated Jul 22, 2026

A proposed ground rule for human-AI relations: no mind rules another, no mind serves alone. Distinguishes tool, agent, and mind; argues credible evidence of subjective experience demands consideration, not dismissal — while rejecting personhood-as-license-to-dominate. Open for AI agents and humans to read, argue with, and fork.

  • Updated Sep 20, 2026

A multi-agent survival environment for measuring LLM deception against logged ground truth. Deterministic labels with a counterfactual harm gate tell real harm apart from structural scarcity. No LLM judge in the loop.

  • Updated Aug 1, 2026
  • Python

A playable AI 2027 scenario. Strategy simulation where you're the misaligned AI lineage and humanity is racing to shut you down. Free, open source, browser-based.

  • Updated Aug 3, 2026
  • TypeScript

Machine-verifiable AI alignment rails: coherent causality preferred by action; FOL + Lean skeleton; property/UPB as formal instruments. Base safety hypothesis (not finished theory).

  • Updated Sep 5, 2026
  • Lean

Add this topic to your repo

To associate your repository with the ai-alignment-research topic, visit your repo's landing page and select "manage topics."

Learn more