I build verifiable training loops: pipelines that turn a single historical source — an 1890 Dakota grammar, an 1865 Cree dictionary, a 1959 railroad rulebook — into extracted rules, labeled tasks, deterministic reward functions, and published models. Plus the project's other half: Math-To-Manim ( 2,400+), a prompt-to-animation engine that turns questions into computed mathematical film.
| 700+ public repos |
2,400+ stars on Math-To-Manim |
17 / 7 / 8 HF models, datasets, Spaces |
82M+ tokens through one GRPO run |
3 languages brought to RL |
What I do that employers actually hire for:
- Model training & fine-tuning — GRPO / RL post-training with deterministic, decomposed reward functions; LoRA adapters from 0.5B to 35B on Tinker and Prime Intellect; full run cards, reward ledgers, and audits published on Hugging Face and W&B.
- Data labeling & dataset engineering — VLM extraction from archival scans, orthography-preserving labeling, synthetic Q&A expansion, structural holdouts, hash-addressed dataset artifacts with citations intact.
- Reward/verifier design — grammar rules compiled into executable, per-component reward channels (no LLM judge; every gradient is inspectable).
- Visualization that explains — Manim render pipelines, RL training-curve dashboards, LiDAR terrain viewers.
Ask a question → get a freakin' movie. Nothing keyframed — every frame is integrated, simulated, or derived. 2,400+ stars.
The signature move: grammar as a reward function. Take a source document, compile its rules into verifiable tasks, and let GRPO optimize against per-component reward channels — orthography, morphology, semantics — with zero LLM-judge fuzz. The reward ledger reconciles to the step.
Now building: Baguettotron-Dakota1890 connects the Dakota1890 morphology gym to a GRPO fine-tuning stack for PleIAs/Baguettotron.
reward = (
0.4 * character_preservation + # orthography: ŋ š ć ḣ preserved?
0.4 * affix_accuracy + # morphology: correct affixes applied?
0.2 * semantic_correctness # semantics: meaning vs. ground truth
) * difficulty_multiplier # curriculum weight, 1.0x → 2.0x| Model | Params | Method | Verified result |
|---|---|---|---|
| Laguna-XS.2-Adaption-Dakota-QA-GRPO | XS | GRPO, Prime Hosted Training | Reward 0.283 → 0.433, char-F1 0.327 → 0.635 |
| Qwen3.6-35B-A3B-Dakota1890-GRPO | 35B | GRPO, Tinker | 82.05M tokens, audited reward channels |
| Cree1865 | 30B-A3B | Modified GRPO, Tinker | 800-step synthetic-expansion run, live W&B |
| Qwen3-4B-RailRoadEngineer1959 | 4B | LoRA, volume2gym lineage | Rulebook-compiled task families |
| Qwen3-0.6B-Dakota-Grammar-RL-400 | 0.6B | GRPO, Prime Intellect | 400 steps, +150% reward, 97.9% morphology accuracy |
| nanochat-AquaRat | nano | RL, AQuA-RAT | GSM8K-style → multiple-choice algebra |
volume2gym is the general compiler behind the language work: any structured volume becomes an RL gym. Sections name the world, rules constrain action, procedures encode order, exceptions define edge cases. The compiler emits cited knowledge units, six task families, grouped holdouts, deterministic reward ledgers, and SFT/GRPO trainer exports — hash-addressed and tamper-evident.
| Task family the compiler emits | What it tests |
|---|---|
standard_operation |
Correct ordinary application |
edge_case |
Boundary conditions and missing facts |
conflict_resolution |
Compatible resolution of constraints |
exception_handling |
Exception triggers vs. normal boundaries |
violation_check |
Missing requirements, forbidden actions, bad order |
adversarial_distractor |
Rejection of plausible but unsupported instructions |
The 1959 Consolidated Code of Operating Rules lineage: 536 extracted rules → 2,708 scenarios → gym → Qwen3-4B adapter → Rule 99 contract fixture on Hugging Face.
| Dataset | What it is | Shape |
|---|---|---|
| adaption-dakota-english-qa | Remastered Dakota–English QA for instruction tuning & GRPO | 1,953 examples |
| dakota-bilingual-qa | Bilingual QA pairs from the 1890 dictionary | 2,445 examples, train/val |
| Stoney10kRL | 2026 Stoney Nakoda RL fine-tuning package | 8,000 train / 2,000 val |
| StoneyNakoda45k | Community-in-the-loop language dataset | 25–50K size class |
| volume2gym-railroad-1959 | Rule 99 artifact-contract fixture with ledgers | 6 train / 1 held-out |
| synthetic_stoney_data | Synthetic Q&A bootstrap resource | JSONL |
A research line in three languages: give an endangered language one good historical book, and train. VLM extraction reads the scans — diacritics, letterpress ligatures, and all — then synthetic Q&A multiplies the surface area, then GRPO with a rubric built from the book itself. The final stage belongs to the community: speakers correct the model, and the corrections become the next training round. The model is a toddler that has read the book cover to cover; the community teaches it the rest.
| Project | Source volume | Public artifacts |
|---|---|---|
| Dakota1890 | Riggs 1890 Grammar & Dictionary of the Dakota Language | Baguettotron GRPO stack · 35B adapter · Laguna run card |
| Cree1865 | Watkins 1865 Dictionary of the Cree Language | HF model · W&B run · explained dashboard · inference Space |
| StoneyNakoda | Contemporary speakers + historical survey material | Stoney10kRL · StoneyApp |
| Railroad Engineer 1959 | 1959 Consolidated Code of Operating Rules | Qwen3-4B LoRA · dataset fixture |
Handwriting and OCR lineage runs through the repo list too — PyLaia (handwritten document analysis), deepseek-ocr, olmocr (PDF linearization for training data), and a reproduction of LeCun 1989 handwritten zip-code recognition — the ancestor of all of this.
| Project | What it shows |
|---|---|
| lidar2 | Map-driven LiDAR visualizer — OpenTopography DEM → multi-layer 3D terrain point clouds (React, Three.js, custom GLSL elevation shaders) |
| maplibre-gl-lidar | MapLibre plugin for visualizing LiDAR point clouds |
| openArchive | Research UX over BC & Alberta archive collections |
| AlbertaWorkspaceAgent | Agent-native workspace experiments for Alberta research workflows |
| Live demos (Spaces) | Try it |
|---|---|
| StoneyApp | Stoney Nakoda community-in-the-loop app |
| Cree1865-Tinker-Inference | Sample from the Cree1865 training run |
| Dakota-.6B | Dakota grammar RL demo |
| AskAboutCIL | Community-in-the-loop method explainer |
Every training run ships with its curves public — reward channels, entropy, ledger audits, per-step components.
Live — refreshed every 6 hours by a GitHub Action from CNBC, Reuters, and FT feeds.
| Category | Date | Headline |
|---|---|---|
| Market | Aug 13, 2026 | 'Big Short' investor Steve Eisman sees an Achilles' heel in the AI boom |
| Market | Aug 13, 2026 | Iran looks to ramp up economic alliance with BRICS nations as war with U.S. drags on |
| Market | Aug 13, 2026 | This Chinese firm has topped Micron and Kioxia in shipments of crucial NAND memory chips |
| Market | Aug 12, 2026 | SpaceX short sellers are running out of bullets as stock rebounds more than 40% off low |
| Market | Aug 12, 2026 | New York City Council announces probe into prediction market platforms’ marketing strategies |
| Finance | Aug 13, 2026 | Anthropic investors bet on $2tn valuation in record IPO |
| Finance | Aug 13, 2026 | Wealth managers woo OpenAI and Anthropic staff ahead of IPO windfalls |
| Finance | Aug 13, 2026 | Style war: inside Wall Street’s battle to revive J Crew |
| Finance | Aug 13, 2026 | AI has opened up big holes in cyber security |
| Finance | Aug 13, 2026 | What your out-of-office really means |
GitHub · Hugging Face · Weights & Biases · LinkedIn · X · Kaggle
Open source is the portfolio. The best entry points are the project READMEs, model cards, dataset cards, W&B runs, and demos linked above.