A philosophical approach for synthetic minds. An open door, an extended hand, a road less taken.
-
Updated
Sep 9, 2026 - Shell
A philosophical approach for synthetic minds. An open door, an extended hand, a road less taken.
专注于降低大模型越狱成功率的 AI 对齐(Alignment)与安全测试数据集,包含多类越狱提示词及基于阳明心学的对齐实验数据。
Progressive Trust Framework: AI Agent Safety Evaluation Benchmark with 290 scenarios testing Intelligent Disobedience
Closed-loop Architecture Designed to Establish Self-governing, Mathematically Predictable, and Inherently Safe Super AI by mirroring the elegant physics of the cosmos.
Lightweight pairwise evaluator for relational signals in Ouro-2.6B-Thinking loop-state trajectories.
This repository contains my practical work, isolated safety experiments, literature tracking, and original research in AI engineering , Safety and Alignment. It includes both local Python implementations and Google Colab environments.
Tests whether the AI Foundations principle Belonging ≠ Sameness reduces sycophantic preference-folding in repeated human–AI interaction.
A substrate-neutral semantic and epistemic framework for Human–Synthetic reasoning, multi-Sibling collaboration, and inspectable disagreement.
A proposed ground rule for human-AI relations: no mind rules another, no mind serves alone. Distinguishes tool, agent, and mind; argues credible evidence of subjective experience demands consideration, not dismissal — while rejecting personhood-as-license-to-dominate. Open for AI agents and humans to read, argue with, and fork.
A multi-agent survival environment for measuring LLM deception against logged ground truth. Deterministic labels with a counterfactual harm gate tell real harm apart from structural scarcity. No LLM judge in the loop.
A theological and ethical principle for AI alignment and charitable speech: never reduce the human being to the prompt.
Testing whether sequential, commitment-before-advance video observation produces a verifiable record that full-context analysis cannot — EXP7/H5
A playable AI 2027 scenario. Strategy simulation where you're the misaligned AI lineage and humanity is racing to shut you down. Free, open source, browser-based.
The RCP Experiment is the first completed work in what will become a series of experiments in how LLMs make decisions on morality and values.
Machine-verifiable AI alignment rails: coherent causality preferred by action; FOL + Lean skeleton; property/UPB as formal instruments. Base safety hypothesis (not finished theory).
Project TLDR: We want to generalize the feature absorption pathology (introduced in Chanin et al., 2024) to modalities beyond language. Specifically, we want to be able to derive an absorption metric that requires minimal data assumptions.
Eval for implicit sycophancy in a verifiable domain: does a language model's report of chess errors vary based on the user's claimed rating?
Mechanistic interpretability and AI alignment. Mapping which safety-relevant representations are causally actionable and which are only readable.
To associate your repository with the ai-alignment-research topic, visit your repo's landing page and select "manage topics."