Tablex is the current working name for a tabular-first agentic data science workbench. It runs alongside Codex or another AgentRunner: the runner reasons, analyzes, models, writes reports, and authors notebooks; Tablex supplies data access, artifacts, lineage, evaluation contracts, safety boundaries, tools, memory, and the human interface.
The product name may still change. Keep code, package names, API paths, and database tables neutral where practical.
Read docs/agent_interface_spec.md before changing Full Auto, Raw, Chat, Research Plan, or notebook behavior.
The most important constraints are:
- Full Auto is one continuing main agent session. Support jobs may train models, ingest artifacts, render notebooks, or build split manifests, but the main Codex reasoning thread must not become a chain of small tickets.
- Raw is the runner transcript. Harness records may be interleaved as clearly labeled sidecar events, but Raw must preserve the real Codex CLI JSONL-style execution record.
- Chat is human accountability, not filtered Raw. Progress, errors, and next actions should be understandable to users. Internal implementation vocabulary and maker notes do not belong in Chat.
- Research Plan is agent-operated. Plan structure and current work come from Codex-authored plan artifacts or schema-validated ResearchPlan requests. UI/process presence must not infer plan progress.
- marimo source is the notebook artifact. Native marimo Python source is first-class evidence. Static HTML snapshots are not notebook artifacts and must not hide runtime failures.
- The harness validates fixed formats, not natural language. JSON schemas, artifact IDs, metric IDs, credentials boundaries, EvaluationSpec, and SplitManifest consistency are harness responsibilities. Objective reasoning, analysis narrative, hypotheses, and modeling judgment belong to the agent.
- Backend: FastAPI, SQLite metadata, local filesystem artifact store, DuckDB-backed profiling, local worker/supervisor CLIs.
- Frontend: React/Vite workbench with Home, Data, Insights/Reports, Leaderboard, Notebooks, Assets, Jobs, Lineage, Agent Chat, and Raw transcript surfaces.
- Agent execution: persistent
AgentSessionfor Full Auto, Codex CLI integration, transcript persistence, workspace materialization under.tablex/, inbox/ack request protocols, and supervisor recovery. - Artifacts and lineage: datasets, semantic/relational catalogs, evidence, assumptions, reports, notebooks, model results, prediction pipelines, pilot runs, and generated assets are stored with hashes and lineage.
- Evaluation and modeling: EvaluationSpec, SplitManifest, ExperimentRun, ModelVersion, and model diagnostics. Every promoted model candidate keeps a validated, versioned prediction runtime for drag-and-drop scoring from the Leaderboard. Complete local training/prediction bundles are built on demand for selected models, so export packaging does not block hypothesis iteration or reduce the visible candidate set.
- Research and skills: project and cross-project Skills, controlled research requests, source-backed findings, and research reports. External claims should be stored as Evidence or source-backed artifacts, not left only in logs.
- Pilot workflow: registered prediction pipelines can be used for prediction batches, outcome ingestion, pilot scoring, and agent-authored validation audits that feed the same continuing Full Auto loop.
- Codex-managed prediction: Predict sends the selected model, actual inputs, evaluation/OOF lineage, purpose, and available evidence to the same continuing Codex session. Codex decides what to inspect, invokes the canonical pipeline, reviews the real output, and can investigate, repair, or rerun before releasing a reasoned result. Tablex supplies timely context, execution, artifacts, lineage, progress, and the user interface without replacing that judgment with fixed drift rules.
- Verified experiment evidence: Validated candidates can link hypothesis, parent run, exact change set, fold results, OOF predictions, learning, and decision in
experiment_evidence.v1. Tablex replays supported metrics from OOF rows joined to frozen DatasetSnapshot labels while leaving analytical judgment to Codex. - Environment-aware compute: Tablex observes CPU, memory, visible NVIDIA devices, compute capability, driver support, and real library GPU probes. Codex receives those facts and chooses CPU or GPU for each useful experiment; Tablex records the requested, selected, and actually reported device with lineage instead of assuming that every GPU-capable host or library is usable.
- Create a project.
- Upload one or more data files. A primary table and objective may be deferred when the task requires data understanding, derived tables, aggregation, clustering, anomaly detection, multi-target setup, or another non-standard objective.
- Start Full Auto or use explicit controls.
- Codex receives the project context, data manifest, artifacts, equipped Skills, runtime facts, evaluation boundaries, and validated tool protocols.
- Codex updates Research Plan, writes progress reports, creates source-backed research findings, authors native marimo notebooks, registers model results, and submits pipeline or pilot artifacts through schema-validated requests.
- Tablex validates fixed-format submissions, stores artifacts and lineage, updates Chat/Activity/Leaderboard/Notebooks/Data/Assets links, and returns structured errors to Codex when a request must be corrected.
Full Auto workspaces are created under:
data/artifacts/agent_sessions/<project_id>/<agent_session_id>/
Common workspace paths:
.tablex/context.json # project context given to Codex
.tablex/GOAL.md # current project goal
.tablex/data/ # stable workspace-readable dataset paths
.tablex/requests/ # schema-checked requests from Codex to Tablex
.tablex/acks/ # structured success/error responses back to Codex
reports/chat_update.md # Codex-authored human progress update
notebooks/*.py # native marimo notebooks
outputs/ # structured plan/results outputs
artifacts/ # additional generated artifacts
The production artifact root defaults to:
data/artifacts/
The metadata database defaults to:
data/metadata/app.db
Backend requires Python 3.10 or newer.
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"Frontend requires Node.js 20 or newer.
cd apps/frontend
npm installRun the backend:
source .venv/bin/activate
uvicorn tabular_harness.main:app --app-dir apps/backend --reload --port 8000Run the frontend in another shell:
cd apps/frontend
npm run devOpen:
http://localhost:5173
Health check:
curl http://localhost:8000/healthThe FastAPI app starts a lightweight local worker by default for concrete sidecar jobs. For heavier debugging or split process operation:
tablex-worker --once
tablex-worker --interval 2 --worker-id local-workerFor manual host-side debugging, run a dedicated Full Auto supervisor from the
managed runtime created by scripts/tablex setup:
.tablex-runtime/venv/bin/tablex-agent-supervisor --interval 15 --owner-id local-agent-supervisorIf the supervisor is split out, start the API with TABLEX_API_AGENT_SESSION_SUPERVISOR_ENABLED=false and run worker processes with --no-agent-session-supervisor where appropriate.
scripts/tablex is the supported complete local launcher. It keeps the API, UI, and non-Codex jobs in ordinary Docker containers, while the Full Auto supervisor and Codex-required jobs run as host companions using the user's installed and authenticated latest Codex CLI. This avoids privileged containers and nested-sandbox failures. Project data is shared through the ignored data/ directory; credentials are never copied into project data or runner workspaces.
From the repository root:
codex login --device-auth # only when Codex is not already authenticated
scripts/tablex upThe device-auth command prints a URL and one-time code. scripts/tablex up reuses that host authentication, creates its managed Python runtime on first use, checks Codex authentication and local sandboxing without a model call, then starts Tablex. On Linux with a running systemd user session, the launcher also registers non-root user services for the Full Auto supervisor and Codex worker. They restart after a process failure and after the next user login following an OS reboot, then continue active projects with their persisted AgentSession, workspace, transcript, and Codex thread ID. No system service or root-owned Tablex process is installed. If the runtime check fails, Tablex stops before starting a misleading UI-only deployment. On Linux, install the official Codex bubblewrap/AppArmor prerequisites; do not disable the host security restriction globally.
Tablex does not pin Full Auto to an older fallback model. The agent model remains the authenticated Codex default unless the user explicitly selects another model.
The project power button is an execution boundary. Turning it off stops the project's Codex process tree and child compute, cancels queued token-consuming work, and prevents automatic recovery from restarting that work until power is explicitly turned on again. Project deletion applies the same boundary and reports success only after execution, project metadata, and project artifacts have been removed and verified.
The launcher requires Docker and the Codex CLI, but its runtime bootstrap does not require host pip, python3-venv, or sudo. It bootstraps a pinned, official uv binary from its digest-locked container image and creates the companion runtime under .tablex-runtime/. Ubuntu may require a one-time administrator installation of its distribution-provided bubblewrap AppArmor profile; use the exact commands in the troubleshooting guide.
When an NVIDIA GPU and the Docker NVIDIA runtime are available, scripts/tablex up builds the optional GPU image and runs real XGBoost, LightGBM, CatBoost, and PyTorch capability probes. Only libraries that pass are reported to Codex as GPU-ready. If detection, image construction, or the probe fails, Tablex starts the CPU runtime instead. CPU-only hosts need no GPU setup. Agent-authored compute runs in a private executor with no Codex authentication mount, metadata database, external network, or published port. Its root filesystem is read-only, while the trusted worker registers outputs and device/resource evidence with lineage.
Open http://localhost:8080 and verify the API:
curl http://localhost:8080/health
scripts/tablex statusAuthentication and project state survive normal restarts. On a systemd-based Linux desktop, the host companions restart automatically after login and resume active Full Auto work. The commands below remain available to refresh the runtime or manage autostart explicitly:
scripts/tablex up
scripts/tablex enable-autostart
scripts/tablex disable-autostartscripts/tablex up enables user-session autostart by default when supported; set TABLEX_AUTOSTART=0 for an intentionally manual session. On hosts without a running systemd user session, the launcher starts ordinary detached companions and clearly reports that scripts/tablex up must be run after an OS reboot. Use scripts/tablex down to stop Tablex and disable its user services. scripts/tablex status shows Docker, companion, and automatic-recovery state, while scripts/tablex logs follows either the user-service journal or detached companion log. Project data remains under data/; host Codex authentication remains in the user's normal CODEX_HOME. Never put Codex tokens, API keys, or authentication contents in Compose environment variables, images, logs, or repository files.
Password auth is optional and controlled by environment configuration. When enabled, Tablex stores user preferences such as locale, avatar, and model preferences server-side. Authentication credentials, password hashes, cookies, OAuth tokens, and connector secrets must never be passed to Codex workspaces, prompts, artifacts, logs, or reports.
Backend tests:
source .venv/bin/activate
pytestBackend type/lint checks:
source .venv/bin/activate
ruff check apps/backend
mypy apps/backendFrontend build:
npm --prefix apps/frontend run buildBrowser golden-slice smoke:
node apps/frontend/e2e/golden_slice_smoke.mjsLive Full Auto research/pilot audit smoke:
node apps/frontend/e2e/live_full_auto_research_pilot_smoke.mjsThis runs a real Codex session with live web research in an isolated temporary data directory and writes audit evidence under output/live/.
Frontend lint:
npm --prefix apps/frontend run lintFor README-only changes, at minimum run:
git diff --check README.md- Public documentation site source
- Development guide
- Agent interface spec
- Benchmark catalog and workflows
- Execution plans
- E2E and audit evidence
The public user documentation is a Docusaurus site under apps/docs/. English is the canonical source; Japanese, Simplified Chinese, and Korean pages live under matching i18n/ paths.
Local docs development:
npm --prefix apps/docs install
npm --prefix apps/docs run start
npm --prefix apps/docs run buildThe GitHub Pages workflow publishes the site to https://nejumi.github.io/Tablex/ when Pages is configured to use GitHub Actions. On first setup, open GitHub repository settings and set Settings → Pages → Build and deployment → Source to GitHub Actions. If the workflow fails at actions/configure-pages with Get Pages site failed, this repository setting is the missing step.
- Keep Codex autonomy intact. Do not add harness-side heuristics that infer targets, task shape, modeling intent, or analytical conclusions from column names or natural-language keywords.
- Prefer artifact-backed outputs and schema-validated request protocols over UI inference or background guessing.
- Preserve EvaluationSpec and SplitManifest integrity for comparable experiments.
- Keep user-facing text practical: what is happening, what changed, what needs attention, and how to continue.
- When a runtime or notebook fails, surface the failure and let the agent repair the source. Do not hide it behind a static fallback.