One Personal Evolving Memory System
English | 简体中文
Live Sites showcase · Owner-only deployment
OPEM, short for One Personal Evolving Memory System, aggregates conversations, durable observations, and tool outcomes from Codex instances across multiple machines, consolidates them into long-term knowledge, supports contextual recall, and evolves as its owner continues to use it.
MVP scope: trusted LAN HTTP, one shared server, SQLite or PostgreSQL, no authentication or TLS.
- Cross-machine collection — multiple Codex CLI, Desktop, or Remote hosts write to one Memory Server.
- MCP-first integration — submit, recall, report verified outcomes, request route advice, record controlled path ablations, archive complete chat, and check health.
- Raw history plus durable Memory — the complete transcript remains available while reusable knowledge is compressed separately.
- Hybrid recall — exact phrase, BM25, fuzzy, metadata, and optional vector similarity are fused with weighted RRF and deduplicated with MMR.
- Session Trace — group each Session into user-led request paths and show Codex/tool nodes with measured or estimated timing.
- Tool Call Groups — build usage profiles from inputs/outputs, language and keywords, then cluster functionally similar tools with explainable similarity and faceted filters.
- Tool archive and FAQ — successful tools become recallable context; failed tools become evidence-backed FAQ entries.
- Daily summaries — browse decisions, learnings, unresolved problems, and Memory highlights on a calendar.
- Small deployment footprint — FastAPI, SQLAlchemy, one Worker, server-rendered pages, and SQLite or PostgreSQL/pgvector.
The editable Mermaid source is available at docs/architecture.mmd.
The server keeps raw ChatMessage and Observation records independent from generated Memory. Every generated Memory retains source Observation and Session links.
Requirements: Python 3.12 or newer.
git clone git@github.com:dawncc/OPEM.git
cd OPEM
python -m venv .venvActivate the environment and install the project:
# Linux / macOS
source .venv/bin/activate
pip install -e ".[dev]"
# Windows PowerShell
.\.venv\Scripts\Activate.ps1
pip install -e ".[dev]"Initialize SQLite:
# Linux / macOS
export DATABASE_URL="sqlite:///./demo.db"
python -c "from memory_server.db import Base,engine; import memory_server.models; Base.metadata.create_all(engine)"# Windows PowerShell
$env:DATABASE_URL="sqlite:///./demo.db"
.\.venv\Scripts\python -c "from memory_server.db import Base,engine; import memory_server.models; Base.metadata.create_all(engine)"Run these in two terminals with the same DATABASE_URL:
# Linux / macOS
codex-memory-server
codex-memory-worker# Windows PowerShell
.\.venv\Scripts\codex-memory-server
.\.venv\Scripts\codex-memory-workerOptionally load demonstration data:
.\.venv\Scripts\python scripts\seed_demo.py
.\.venv\Scripts\python scripts\seed_conversation_demo.pyOpen http://127.0.0.1:8000.
The Compose deployment starts PostgreSQL/pgvector, the API/UI server, and the Worker.
cp .env.example .env
cd deploy
docker compose up --build -dOpen http://<server-lan-ip>:8000. PostgreSQL is only exposed to the internal Compose network.
Install the package on every Codex machine:
pipx install .Add the MCP server to $CODEX_HOME/config.toml:
[mcp_servers.codex_memory]
command = "codex-memory-mcp"
args = ["--server", "http://192.168.1.100:8000"]
startup_timeout_sec = 30Copy skill/codex-memory to $CODEX_HOME/skills/codex-memory, restart Codex, and call memory_health.
The Skill instructs Codex to:
- recall relevant project history before substantial work;
- save durable decisions, learnings, problems, and solutions;
- archive the complete conversation and all available tool results;
- include
duration_mswhen exact timing is available; - submit a structured summary before session end or compaction.
- report whether recalled memories were helpful, irrelevant, or harmful.
| Tool | Purpose |
|---|---|
memory_submit |
Save a durable observation, decision, learning, problem, solution, or session summary. |
memory_recall |
Return structured matches and a prompt-ready <chat-memory> context block. |
memory_feedback |
Record recall utility and calibrate future ranking confidence. |
memory_task_outcome |
Report evidence-backed task results without treating silence as success. |
memory_route_recommend |
Return a versioned observe, shadow, canary, or active route recommendation. |
memory_path_intervention |
Record paired-replay or randomized path-ablation evidence. |
memory_session_end |
Submit completed work, decisions, unresolved items, and files. |
memory_chat_submit |
Store the complete chronological user/assistant/system/tool transcript. |
memory_health |
Check the server/database and flush the local pending queue. |
The MVP keeps the business API intentionally small:
| Method | Path | Purpose |
|---|---|---|
POST |
/api/v1/observations |
Receive durable observations and session summaries. |
POST |
/api/v1/chat/messages |
Receive complete chat batches and tool metadata. |
POST |
/api/v1/recall |
Run hybrid Memory retrieval. |
POST |
/api/v1/memories/feedback |
Record auditable recall feedback and evolve confidence. |
POST |
/api/v1/tasks/outcomes |
Attach strong, medium, or weak outcome evidence to a TaskRun. |
POST |
/api/v1/routes/recommend |
Assign a versioned execution strategy; observe/shadow never change runtime behavior. |
POST |
/api/v1/paths/interventions |
Record causal evidence from a controlled path ablation. |
GET |
/api/v1/evolution/report |
Compare quality confidence, failures, cost, and latency by route and task bucket. |
GET |
/api/v1/health |
Check database health and pending jobs. |
Main UI routes:
/— overview/search— Memory search/calendar— daily Memory calendar/traces— Session request paths and timing/tools— tool execution archive/faq— failed-tool FAQ/sessions/<id>— complete conversation
- Capture — save the raw Observation or complete chat first.
- Compress — use an optional OpenAI-compatible model, or deterministic rule compression when no LLM is configured.
- Consolidate — merge only compatible Memory types and keep every source link.
- Retrieve — combine exact phrase, BM25, fuzzy, metadata, and optional embedding results.
Successful tools become tool_success Memory. Failed tools become separate tool_failure_faq Memory so similar success and failure output can never be consolidated together. Unknown outcomes are archived without promotion.
Recall outcomes form a bounded evolution loop: explicit feedback calibrates ranking confidence without changing application code or model weights. See the self-evolution design for the safety boundary and roadmap.
Pass top-level duration_ms on a memory_chat_submit message when exact timing is known. Trace renders exact timing in blue. Older messages fall back to the adjacent message timestamp and are explicitly marked as estimates. Active duration excludes idle time between separate user requests.
If a remote host cannot route directly to the LAN server, create a reverse SSH tunnel from the Memory Server machine:
ssh -N -R 18000:127.0.0.1:8000 <remote-host>Configure the remote MCP bridge with --server http://127.0.0.1:18000. The included scripts/test_remote_codex.sh validates health, chat submission, successful/failed tool archival, FAQ generation, and recall from a real remote Codex process.
A directly routable LAN, VPN, or private network address is preferred for permanent deployment; the reverse tunnel only exists while its SSH process is running.
| Variable | Default / example | Description |
|---|---|---|
DATABASE_URL |
postgresql+psycopg://memory:memory@postgres/memory |
SQLAlchemy database URL; use sqlite:///./demo.db locally. |
MEMORY_SERVER_HOST |
0.0.0.0 |
API bind address. |
MEMORY_SERVER_PORT |
8000 |
API/UI port. |
MEMORY_LLM_ENABLED |
false |
Enable OpenAI-compatible compression. |
MEMORY_LLM_BASE_URL |
http://localhost:11434/v1 |
OpenAI-compatible endpoint. |
MEMORY_LLM_MODEL |
qwen3 |
Compression model name. |
MEMORY_EMBEDDING_ENABLED |
false |
Enable local embeddings. |
MEMORY_EMBEDDING_MODEL |
BAAI/bge-small-zh-v1.5 |
Sentence-transformers model. |
MEMORY_RECALL_CANDIDATE_LIMIT |
0 |
0 searches every Memory; use a limit for larger datasets. |
MEMORY_MODEL_PRICES_JSON |
{} |
Optional model price overrides as input/cached/output rates per million tokens. |
MEMORY_STRATEGY_CANARY_ENABLED |
false |
Enable bounded low-risk canary assignment after evidence gates pass. |
MEMORY_STRATEGY_MIN_SHADOW_CASES |
20 |
Strong-evidence tasks required before proposing a shadow route. |
MEMORY_STRATEGY_MIN_CANARY_CASES |
50 |
Strong-evidence tasks required before canary eligibility. |
MEMORY_DAILY_REFRESH_SECONDS |
60 |
Daily-summary refresh interval. |
MEMORY_TIMEZONE |
Asia/Shanghai |
Daily-summary grouping timezone. |
MCP/Skill-only capture depends on Codex choosing to call a tool and can miss ordinary requests. Install deterministic lifecycle hooks on every Codex host:
python scripts/install_codex_hooks.py --server http://192.168.1.100:8000The installer preserves existing hooks and adds UserPromptSubmit, PostToolUse, Stop, and PreCompact capture. User prompts, tool results, and final assistant messages become visible immediately; Stop also creates an Observation for asynchronous Memory consolidation. The standard-library-only client queues failed deliveries under $CODEX_HOME/codex-memory/pending.jsonl and retries them on the next lifecycle event.
After installation, open /hooks in Codex, review and trust the new definitions, then start a new task. The home page refreshes its live-session list every five seconds. hooks/session-end.py remains only as a legacy compatibility entry point.
Preview completed turns stored under $CODEX_HOME/sessions, then synchronize them:
codex-memory-sync --server http://192.168.1.100:8000 --since 2026-07-01 --dry-run
codex-memory-sync --server http://192.168.1.100:8000 --since 2026-07-01Only turns with a task_complete event are imported. The backfill uses the same turn, tool-call, and event identifiers as the real-time hooks, so reruns and hook/backfill overlap are idempotent. Legacy rows without event identifiers use role/content occurrence matching as a compatibility fallback. Rollout JSONL remains a best-effort local import source; real-time hooks are the supported primary capture path.
- Cost control and context compression for multi-turn tasks (Chinese)
- Zero-sample execution-path evolution (Chinese)
packages/memory-common/ Shared schemas, settings, pending queue
packages/memory-server/ FastAPI, SQLAlchemy, recall, Trace, Jinja2 UI
packages/memory-worker/ compression, consolidation, embeddings, daily summaries
packages/memory-mcp/ local stdio MCP-to-HTTP bridge
skill/codex-memory/ Codex workflow Skill
hooks/ Stop / PreCompact fallback hook
migrations/ Alembic migrations
deploy/ Docker Compose deployment
scripts/ seed and remote validation scripts
tests/ unit and integration tests
docs/images/ README screenshots
pip install -e ".[dev]"
pytestCurrent test coverage includes chat preservation, Chinese/English retrieval, recall isolation, pending retries, tool outcome classification, TaskRun/DAG derivation, evidence aggregation, causal-path guards, strategy gates, FAQ construction, Worker consolidation boundaries, Trace timing, and HTML escaping.
This MVP assumes a trusted private network. It does not implement TLS, authentication, device tokens, tenant isolation, or automatic redaction. Add those controls before exposing the service beyond a trusted LAN or VPN.
Copyright (c) 2026 dawncc.
This project is open-source software licensed under the MIT License. You may use, copy, modify, merge, publish, distribute, sublicense, and sell copies subject to the license terms.