Kamibiki is a contextual search engine for agentic coding. It indexes git repositories locally using embeddings and provides fast semantic search from the command line or as an MCP server for coding agents.
- Rust — Install from https://rust-lang.org/tools/install
- Voyage AI API key — Sign up at https://www.voyageai.com/ and create an API key. Kamibiki uses Voyage AI for embeddings and reranking.
The quickest way is to install directly from the git repository with
Cargo, which builds the kb binary and places it in ~/.cargo/bin
(make sure that's on your PATH):
cargo install --git https://github.com/dominiccooney/kamibiki.gitAlternatively, clone and build from source:
git clone https://github.com/dominiccooney/kamibiki.git
cd kamibiki
cargo build --releaseThe binary is at target/release/kb. You can copy it somewhere on
your PATH, or run it directly.
kb initThis prompts for your Voyage AI API key and saves it to ~/.kb.conf.
kb add <name> <path/to/git/repo><name> is a shorthand alias you'll use in other commands. For
example:
kb add myproject ~/src/myprojectkb index <name>This chunks all tracked files, computes embeddings via the Voyage AI
API, and writes index files to a .kb folder at the repository root.
Add .kb to your .gitignore.
Indexing supports delta updates — subsequent runs only re-embed changed files. If interrupted, re-running resumes from where it left off.
You can index at a specific git revision with --commit:
kb index myproject --commit my-feature-branch
kb index myproject --commit abc1234
kb index myproject --commit HEAD~3When --commit is omitted, it indexes at HEAD.
Pass --compact to force a self-contained index at the target
commit. Kamibiki normally writes small delta files that point back
to an older "parent" index via a commit hash — searches walk that
chain to fill in unchanged chunks. A compact index is a single
standalone file (its parent hash is zero), so chain walks
terminate there. This is useful when you want to drop the tail of
a long delta chain, or stop a search from consulting older
history.
Compact rebuilds are embedding-preserving: every chunk whose
(path, text) matches a chunk in the existing .kbi or the next
most recent index is copied over verbatim — no API call — so
recompacting a repository you've already indexed typically only
embeds the handful of chunks whose bytes actually differ. The new
index is written to a side file, then atomically renamed over the
previous one; if interrupted, a re-run resumes the in-progress
compact build.
kb index myproject --compactkb search <name> <query>For example:
kb search myproject "how does authentication work"Use -n to control the number of results (default 10), and
--commit to search from a specific revision:
kb search myproject "error handling" -n 5 --commit my-branchUse . as the name to search the repository in the current directory:
kb search . "database connection pooling"Use --dir to focus the search on one or more directories, and
--exclude-dir to leave parts of the tree out. Both flags are
repeatable and accept absolute, cwd-relative, or repo-relative paths
(each is relativized against the repository root). Exclusions take
precedence over inclusions:
kb search . "request validation" --dir src/api --dir src/web --exclude-dir src/api/generatedBecause the paths are relativized rather than resolved on disk, you
can scope a search at an earlier --commit even if the directory no
longer exists in your working tree.
kb status
kb status <name>Shows indexed repositories, index file count and size, latest indexed commit, and embedding count.
| Command | Description |
|---|---|
kb init |
Set up your Voyage AI API key |
kb add <name> <path> |
Register a git repository for indexing |
kb index [names...] [-c commit] |
Update the index (delta-aware, restartable) |
kb status [name] |
Show indexing status |
kb search <name> <query> [-n top] [-c commit] [--dir d] [--exclude-dir d] |
Search a repository |
kb alias <name> <repos...> |
Create a shorthand for a group of repositories |
kb drop <name> |
Delete all index files for a repository |
kb gc [name] [--dry-run] |
Delete index files for commits no longer known to git |
kb start |
Launch the MCP server on stdio |
kb gc walks each repository's .kb directory and deletes any
.kbi file whose indexed commit no longer resolves in git. This
cleans up after branches that have been rewritten or rebased and
whose tip commits have been reaped by git's own garbage collector.
A dying commit's parent is often still live (you'd typically reindex
the tip of a rewritten branch, not its ancestors), so kb gc never
cascades deletions up or down the chain — it deletes only the files
whose own commit is gone. As a consistency check it warns if a
surviving index's parent_hash points at one of the files being
removed, which shouldn't normally happen.
Use --dry-run to preview what would be deleted without touching
the filesystem. Omit name to gc every registered repository.
kb start launches Kamibiki as an MCP server over stdio (JSON-RPC
2.0, newline-delimited). It exposes three tools:
kb_search — Search an indexed repository for code relevant to a query. Parameters:
name(required): repository name, or.for the current repoquery(required): natural language or code search querytop(optional): number of results, default 10commit(optional): git revision to search from (commit hash, branch name, tag,HEAD~1, etc.). Defaults to HEAD.dirs(optional): array of directories to focus the search on (absolute, cwd-relative, or repo-relative; relativized against the repository root). Omit for the whole repository.exclude_dirs(optional): array of directories to exclude. Same path forms asdirs. Exclusions take precedence overdirs.
kb_status — Show indexing status of registered repositories. Parameters:
name(optional): specific repository name; omit for all repos
kb_index — Update the search index for one or more repositories. Parameters:
names(optional): array of repository names; omit to index allcommit(optional): git revision to index at (commit hash, branch name, tag,HEAD~1, etc.). Defaults to HEAD.compact(optional): when true, write a self-contained root index (parent hash zero) at the target commit. Embeddings from the existing.kbiand the next most recent index are reused verbatim; only chunks whose text has no match get re-embedded. The new file is atomically renamed over the old one. Defaults to false.
Add Kamibiki to your MCP server configuration. For example, in Cline (VS Code or CLI), edit your MCP settings:
{
"mcpServers": {
"kamibiki": {
"command": "/path/to/kb",
"args": ["start"],
"disabled": false
}
}
}For Cline VS Code, this file is at
~/.config/Code/User/globalStorage/saoudrizwan.claude-dev/settings/cline_mcp_settings.json.
For Cline CLI, it's at ~/.config/cline-cli/cline_mcp_settings.json.
Once configured, your coding agent will see kb_search, kb_status,
and kb_index as available tools. Depending on your model, you may
need to prompt it to use the tool the first time — after that it
tends to use it regularly.
Kamibiki is implemented in Rust and uses gitoxide for high-bandwidth
read-only access to git repositories. Its indexes are flat files
stored in a .kb directory at the repository root. As your
repository changes, it writes overlay files. These are periodically
compacted to reclaim space and improve search speed.
Kamibiki skips binary files and extremely large files. Working with multimodal embedding models (e.g. to embed images) is something we might consider in future.
As an embeddings-based search engine, Kamibiki is sensitive to the chunking model, which breaks a file into chunks and possibly reformats those chunks to add context; and the embedding model, which projects those chunks into vector space. The set of supported chunking and embeddings models are hard-coded into the binary and written into index metadata.
Supported chunkers:
tsv1 — A tree-sitter based chunker which recursively divides parse trees for popular languages until a suitable chunk size is produced. For unsupported languages, it falls back to splitting on newlines and recursively merging chunks until a suitable size is reached. The exact chunk size depends on the capabilities of the embedding model.
Supported models:
1 = voyage-code-3@2048, binary quantization (~256 bytes/chunk)
The Kamibiki index file format is designed to be mmap'ed with the bulk of the data read-only. This lets us rely on the operating system's file system cache, page cache and virtual memory manager to be fast. Each index file is named after a git hash, and contains:
-
A version number, currently one byte = 1. These files always use the voyage-code-3@2048, binary quantization embedder. (The index is insensitive to the chunker used, as long as it produces non-overlapping chunks that start at the first byte of the file.)
-
The git hash this index was created at. 64 bytes to support repositories with SHA-256 hashes, but may just have 20 bytes/160 bit SHA-1 hash.
-
A parent git hash, if this index should be overlayed on an existing index. This way when a repository changes we can produce a small index containing just the differences.
-
A table of offsets. The offsets are offsets into the git index at the relevant revision. The offsets can be positive or negative s16 numbers. A positive number indicates the following n items from the git index, in order, are included in the index. A negative number indicates the following -n items from the git index, in order, are skipped. A 0 offset indicates the end of the table.
-
A table of chunk counts. For each file included (by virtue of being covered by a positive count in the index offsets table) a u16 number of chunks. Files with more than 65,534 chunks are not supported. (0 chunks is valid, meaning there are zero chunks. 65,535 is reserved.)
-
A table of lengths, per file and per chunk within each file, in order. The length is a u16 indicating the length of the chunk in bytes. For files with specific encodings, chunkers must break chunks on valid character boundaries.
-
Padding to align embedding data to a word boundary suitable for vector processing. The embedder component is responsible for describing this alignment.
-
The embedding data. This should be aligned to a suitable width for vector processing of the embeddings.
To serve a query, Kamibiki:
-
Embeds the query using the relevant embedding model.
-
Uses vector operations to compute cosine distance in embedding space. For binary quantized embeddings, this is XOR and POPCOUNT. It maintains a max-heap to keep the top N results with the smallest embedding distance. Each result is identified by its offset in the embeddings table. N is chosen to be close to the limits of the reranking model in step 5.
-
Sorts the top N results by offset in the embeddings table.
-
Scans with parallel pointers into the git index, index table, chunk count and offset tables. Finds the index entry and byte offsets of the top N chunks, and uses git to produce the chunk content.
-
Submits the query and top N chunks to a reranking model (not exceeding the context length limitations of the reranking model).
-
Returns the results from the reranker, in order.
When the index has a parent index, the results from steps 2–4 are refined iteratively walking backwards over the relevant indexes. Results from the newer indexes are preferred. If an item is deleted in the revision being searched (not in the index but in git itself), stale results from old indexes are suppressed before step 5.
Kamibiki does not search uncommitted changes. In future, we might do that with an ephemeral in-memory index for uncommitted changes.