A Go 1.26 app for a collection of sutta quotes. cmd/extract distills the
quotes embedded in the essay dumps (dumps/*.txt) into one canonical format and
a SQLite seed. cmd/server is the web app: a single binary that serves the
quotes as editable, rune-sorted blocks persisted in SQLite in real time.
dumps/ source essays (input, hand-written)
discerning-truth-from-deception.txt prose only, no sutta quotes
sacredness-and-profanity.txt sutta quotes, inline-cited
stream-entry-for-lay-buddhists.txt sutta quotes, inline + header-cited
internal/quote/ parser, normalizer, renderer, near-duplicate detection, canonical-text import, seed emitter (+ tests)
internal/search/ pure full-text filter (Query/Parse/Match/Filter) over []store.Quote (+ tests)
internal/store/ SQLite store: CRUD + collections + categories + merge + JSON backup/restore, ordered by char_count
internal/store/storetest/ CloneFixture: a clone-of-main test database for integration tests
internal/seed/ EnsureSeeded: canonical seed (+ sample categories) on a fresh database
internal/server/ HTMX handlers + server-rendered templates (+ tests)
internal/coverbadge/ Go-cover parser + README badge renderer (+ tests)
cmd/extract/ CLI: reads dumps/ and writes database/ + exports/
cmd/server/ web server: opens + seeds the DB, serves the UI
cmd/coverage/ CLI: parses a cover profile, refreshes the README badge
cmd/screenshot/ CLI: serves the seeded app in-process, captures docs/home.png
cmd/fixture/ CLI: dumps the main DB into the clone-of-main test fixture
database/
seed.sql generated schema + inserts (committed, embedded)
quotes.db SQLite database (gitignored, created on run)
docs/
home.png README home screenshot (committed, regenerated by `make screenshot`)
exports/
shortest-first.md generated export, shortest-first (committed)
web/
templates/ layout, rail_left, rail_right, root_zone, collection_zone, check_zone, check_results, collection_list, quote_list, quote_block(_ro), quote_form, quote_import_form, quote_chips, quote_collection_chips, quote_category_editor
static/ app.css (typography/components) + layout.css (4-zone grid), app.js, htmx (vendored)
go.mod
readme.md
changelog.md
.gitignore
Run the server (CGO is required for the SQLite driver):
CGO_ENABLED=1 go run ./cmd/server # http://localhost:8080
CGO_ENABLED=1 go run ./cmd/server -addr :9000 # custom port
CGO_ENABLED=1 go run ./cmd/server -db /tmp/q.db # custom database pathOn first run the server creates database/quotes.db and loads the canonical
seed (109 quotes, three sample categories, one sample collection). After that
your edits persist; the seed is never reapplied, so deleting a quote is
permanent.
The UI is a dual-pane workspace. A left rail (Home, Categories, Duplicates, Backup) and a right rail (Collections) flank two text columns: the root corpus on the left and the active collection on the right. Each text column scrolls independently, so you can keep a different spot open in each. A thin header atop each column shows a name and a count, and the two columns' header rows are kept aligned across the workspace. Each root block:
- renders the quote in the canonical format, with passages in italics and the
sutta id bolded and linked to
https://suttacentral.net/<id-without-spaces>(e.g.MN 22tomn22), opening in a new tab; - shows its categories as chips, with an inline editor to tag it;
- shows the collections it belongs to as a second chip row;
- has Copy, Edit, and Delete actions;
- has a checkbox for bulk delete, with a select-all control in the toolbar.
Home is kept in shortest-first (rune-count) order. A newly added quote slots into its sorted place automatically; there is no drag on home, since home is the canonical corpus.
New opens a form (content, attribution, text ID) that also tags the quote with
categories: check any existing categories and/or type a new one, which is
created case-insensitively unique. An empty attribution defaults to "the
Buddha". Copy all copies every quote as one text joined by the dot
separator. Each rail also has a Copy ids button: the left rail copies the whole
corpus's text ids (/ids.txt), the right rail copies the active
collection's (/collections/{id}/ids.txt), each deduped and sorted, one per
line. Selecting a category or a collection swaps just that pane in place, and
the URL carries ?cat= or ?col= for deep linking.
Every mutation is written to SQLite before the UI updates, and the rails and column counts refresh live via out-of-band swaps. The Duplicates section, the category and collection counts, and the root "N blocks" header all stay current without a full reload.
Check one or more root quotes and insert gaps (+ markers) appear between every
pair of collection blocks. Clicking one inserts the selection at that 1-based
position, shifting later items down; duplicates are skipped. "+ from selection"
in the right rail creates a new collection from the selection and makes it
active. Collections are named (inline rename in the right rail) and default to
"Col {id}" until renamed.
A collection's blocks are copyable (copy-one, copy-all via
/collections/{id}/export.txt) and drag-to-reorder, saved to the collection's
own order. They are read-only for content: no New, edit, or delete, so home
stays the sole source of truth. Each collection has a Delete button.
Categories are named tags managed independently in the left rail. Create,
rename, or delete them inline; names are unique, case-insensitive. Each root
block shows its categories as chips, and the inline editor toggles any
combination and can create a new category on the spot; the New form offers the
same choice when a quote is first added. Clicking a category (in
the rail or on a chip) filters the root column to its quotes, with Copy all via
/categories/{id}/export.txt. Deleting a category untaggs its quotes; deleting
a quote clears its tags.
The left rail's Duplicates section surfaces near-duplicate quotes: clusters of
two or more passages whose word-level Jaccard similarity exceeds 0.8
(quote.GroupDuplicates, joined transitively through a disjoint set, so a group
is a connected component rather than a clique). Grouping is by content, not by
text id: only quotes whose bodies are near-identical are listed, and each row is
labelled with the group's representative (shortest) text id and a member count.
Clicking a group jumps to the representative in the root column, switching to
Home first if a category filter is active, and briefly highlights it. The
canonical seed already contains one such cluster: the MN 22 trio that differs
only in "Bhikkhus"/"Mendicants" and "sexual"/"sensual". Adding, editing, or
deleting a quote refreshes the section live. Each group's ‖ button merges
every member into the shortest representative in one transaction: the merged
quotes' collection and category memberships fold onto the keeper, then the
duplicates are deleted (store.MergeQuotes, POST /duplicates/{rep}/merge).
The root column and both rails refresh in place.
Each text column has a search box in its toolbar, scoped to the column's active
set: Home or the current category on the left, the active collection on the
right. Typing filters the column in place over htmx (debounced); matching is
case-insensitive against the quote body and citation, with hits wrapped in
<mark>. Bare words are AND-ed (all must appear); surround a multi-word phrase
with double quotes ("right view") to match it verbatim; prefix a term or
phrase with - (-suffering, -"the buddha") to exclude it. ?rq= and
?cq= deep-link the two columns independently. Switching category or collection
clears the search. While a collection search is active the insert gaps and drag
handles are hidden, so a filtered subset cannot be mis-reordered.
The Check ids button atop the right rail repurposes the right text column as a
workspace: paste a list of text ids (one per line) and Check reports each one's
membership in the corpus (GET /pane/check swaps in the workspace, POST /check
returns the results fragment). Each input is canonicalized via
quote.CanonicalSuttaID (so "the Buddha, MN 22" matches "MN 22") and matched
case-insensitively against the lowercased corpus id set; a non-canonical input
falls back to a literal match. Found ids show a check plus the number of
matching quotes; missing ones show a cross. Clicking the button again (or a
collection) leaves the workspace. The workspace reuses the #collection-zone
target but carries no
data-cid, so collection-mutation actions stay inert while it is shown.
The root toolbar's Import button swaps a textarea into the New-form slot. Paste
quotes in the canonical format (the same one exports/shortest-first.md uses:
italic passage lines and a - **citation** tail, blocks separated by the . . .
divider) and submit; quote.ParseCanonical recovers each block and
POST /quotes/import creates them, de-duplicating within the paste. The root
list re-renders in sorted order with the rails refreshed and the form cleared.
A block without a citation tail is skipped; a plain - citation tail (no bold
markers) is accepted as a fallback for lightly edited pastes. Editing a quote
re-sorts the list (an edit can change its rune count) and briefly flashes the
saved block at its new position.
The left rail's Backup section holds a portable, whole-database snapshot.
Download JSON (GET /backup.json) writes a versioned store.Dump (every quote
plus the named collections and categories with their ordered memberships) as a
quotes-backup.json attachment. Restore uploads one: POST /restore accepts the
same JSON and, after a confirm, replaces all five tables in a single transaction
(store.Import), preserving explicit ids so the canonical ranking, collection
order, and tags survive the round-trip. An unknown dump version is rejected with
store.ErrUnsupportedDump. Restore is destructive and replaces everything; for
adding quotes without overwriting, use Import instead.
Home is ordered by char_count (rune count), so the table matches the seed
schema:
CREATE TABLE quotes (
id INTEGER PRIMARY KEY, -- shortest-first rank (canonical)
sutta_id TEXT NOT NULL,
citation TEXT NOT NULL,
body_md TEXT NOT NULL, -- canonical italicized format
body_text TEXT NOT NULL,
line_count INTEGER NOT NULL,
char_count INTEGER NOT NULL, -- rune count; home is ordered by this
sources TEXT NOT NULL
);id and char_count carry the canonical shortest-first ranking, which List
orders by. Collection order is separate.
Collections are named (or autonumbered) subsets curated from home. Deleting a quote on home also removes it from every collection:
CREATE TABLE collections (
id INTEGER PRIMARY KEY,
name TEXT NOT NULL DEFAULT '' -- empty renders as "Col {id}"
);
CREATE TABLE collection_items (
collection_id INTEGER NOT NULL,
quote_id INTEGER NOT NULL,
position INTEGER NOT NULL, -- 1-based; insert-at-index shifts this
PRIMARY KEY (collection_id, quote_id)
);Categories are named tags. Deleting a quote clears its tags; deleting a category untaggs its quotes:
CREATE TABLE categories (
id INTEGER PRIMARY KEY,
name TEXT NOT NULL UNIQUE COLLATE NOCASE
);
CREATE TABLE category_items (
category_id INTEGER NOT NULL,
quote_id INTEGER NOT NULL,
PRIMARY KEY (category_id, quote_id)
);database/seed.sql and exports/shortest-first.md are generated. Regenerate
with go run ./cmd/extract. database/quotes.db is gitignored; populate it
from the seed with sqlite3 database/quotes.db < database/seed.sql. Never
hand-edit generated files.
Every extracted quote is normalized into one format and written to both the
database (body_md) and exports/shortest-first.md:
*"first passage*
*second passage*
*last passage"* - **the Buddha, MN 22**
- Each passage line is wrapped in italics (
*…*). - Lines 1..n-1 end with two spaces (a Markdown line break); there are no blank lines between passages of the same quote.
- The last line ends with
- **<citation>**, outside the italics. <citation>keeps the full attribution as found in the source (the Buddha, MN 22,the Buddha to layman Pessa, MN 51,layman Siha, AN 8.12). Any quote recorded without an attribution is attributed to the Buddha (e.g.the Buddha, AN 4.180,the Buddha, SN 55.1); suttacentral URLs in( … )are dropped.- Source curly quotes (
“ ”) and Pāli diacritics are preserved.
Consecutive quotes in exports/shortest-first.md are divided by:
.
.
.
Two blank lines before and after the divider; the first two dots carry two trailing spaces.
The dumps quote suttas in several formats; all are reduced to the canonical
form above (internal/quote).
- Inline-cited: a block whose last line ends with
- <citation>. Covers single-line quotes, multi-line dialog, and narrative-framed passages. - Header-cited: a lone
SUTTA:line (e.g.SN 55.1:,MN 13:); every following block becomes the quote's passages until the next header or a.divider. Such quotes include any framing narrative the essay placed between the header and the divider (e.g.MN 13). - Verse with stanza breaks: a quote may span several blank-separated blocks.
Leading blocks that open with
“but carry no citation are absorbed into the next cited block (e.g.SN 5.2).
Per-line cleanup: a leading (N) numbering marker and stray * / _ Markdown
artifacts are stripped; the - <citation> tail of the closing line is removed
(it is rendered separately). A citation with no attribution (just the sutta id,
as with all header-cited quotes) is normalized to the Buddha, <id>.
Sutta-ID forms recognized: (DN|MN|AN|SN) N[.N…][-N][#…], KN <sub> N[…],
pli-tv-…#…, and the abbreviated Vinaya Tv Vi Bu Pj1.
- De-dup by normalized passage text (whitespace collapsed). Source-file lists are merged; the first-seen citation and sutta id are kept.
- Order shortest-first by rune count of the concatenated passages (stable, with
deterministic tie-breakers on sutta id then body text). Row
idequals the shortest-first rank. - Five quotes recur across both essay files (e.g.
AN 8.53,MN 117,SN 20.7) and are collapsed to one row each.
CREATE TABLE quotes (
id INTEGER PRIMARY KEY, -- shortest-first rank
sutta_id TEXT NOT NULL, -- canonical id, e.g. "MN 22"
citation TEXT NOT NULL, -- full kept citation
body_md TEXT NOT NULL, -- canonical italicized format
body_text TEXT NOT NULL, -- plain passages joined by newlines
line_count INTEGER NOT NULL,
char_count INTEGER NOT NULL, -- rune count of passages (sort key)
sources TEXT NOT NULL -- ';'-joined dump files
);Indexes on char_count and sutta_id. Current seed: 109 quotes, char counts
from 51 to 5032.
CGO_ENABLED=1 go test ./... # full test suite (CGO for the SQLite driver)
CGO_ENABLED=1 go vet ./... # static checks
go run ./cmd/extract # writes database/seed.sql + exports/shortest-first.md
make coverage # recomputes Go test coverage, refreshes the README badge
make screenshot # serves the seeded app and captures docs/home.png
make fixture # dumps the main DB into the clone-of-main test fixture
CGO_ENABLED=1 go run ./cmd/server # run the web app (http://localhost:8080)To rebuild database/quotes.db from a changed seed.sql, delete the database
file and restart the server. It re-seeds only an empty or unseeded database:
rm database/quotes.db && CGO_ENABLED=1 go run ./cmd/serverdiscerning-truth-from-deception.txtis prose only. It mentions suttas inline (e.g.AN5.34#7.9) but contains no citation-terminated block quotes, so it contributes zero quotes.- The Vinaya "black snake" passage appears as
Tv Vi Bu Pj1in one essay andpli-tv-bu-vb-pj1#5.11.20in the other with identical text. De-duplication collapses them to one row, keeping the first-seenpli-tv-…id.