MobileKV is a KV cache storage manager for edge/mobile LLM inference. It focuses on memory layout, allocation, addressing, and window retention. It does not run model kernels or quant/dequant math.
MobileKV currently provides:
- Multi-layer KV storage (
LayerStorage) with per-layer K/V template binding - Mixed precision per plane (
KandVcan use different scalar types) - Independent sequence capacity control per plane (
KandVcan use different initial/max seq) - Ring-buffer mode for sliding-window retention
- Plain and dim-block storage templates
- Opaque-scalar registry for dim-block template construction
- Raw-typed convenience accessors with runtime scalar-type checks
- Format metadata descriptor (
FormatDescriptor) attached to templates
Supported scalar types in core templates:
FP32,FP16,BF16,INT8,UINT8,INT16
Current implementation explicitly does not include:
- Attention kernel execution
- Quantization/dequantization compute
- Disk-backed KV runtime implementation
- Block-table/page-table scheduler APIs for large-scale continuous batching
StorageMode::Blocked, MemoryDomain::DiskMapped, and thread_safe flags exist in API surface for future extension, but are not fully implemented runtime features yet.
KVTemplate: Defines layout mapping from logical coordinates to physical bytes.KVPlane: Concrete storage instance (oneKor oneV) for a layer.LayerStorage: OwnsK/Vplanes for one layer.KVCacheStorage: Global container for all layers and templates.
Logical coordinate:
LogicalCoord{layer, seq, head, dim}
Physical address:
PhysicalAddr{byte_offset, byte_size, ...}
When ring mode is enabled (max_seq_capacity > 0):
- Plane keeps only the latest window up to
max_seq_capacity. locate()interpretsseqas a window-local logical index in[0, seq_length).- Internal logic maps logical index to physical slot after wrap-around.
- Overflow append overwrites oldest tokens.
Strict validation:
- Build fails if
initial_seq_capacity > max_seq_capacity(no silent clamp).
KVCacheStorageConfig includes:
default_max_seq_capacity
Behavior:
- 4-arg
add_layer(layer, k_template, v_template, initial)inheritsconfig.default_max_seq_capacity. - 5-arg
add_layer(..., initial, max)explicitly overrides it. - 7-arg
add_layer(..., k_initial, v_initial, k_max, v_max)enables independent K/V capacities. - Set
default_max_seq_capacity=0to keep non-ring growth behavior.
When using 7-arg builder API:
builder.add_layer(
0, // layer_id
1, // k_template
2, // v_template
512, // k_initial
256, // v_initial
2048,// k_max
1024 // v_max
);K and V ring windows then evolve independently. Check each plane via:
layer.plane(PlaneKind::K).stats()layer.plane(PlaneKind::V).stats()
For attention runtime integration, use a consistent effective length policy
(for example effective_len = min(k_seq_length, v_seq_length)) when K/V windows differ.
#include "mobilekv/kv_cache.h"
using namespace mobilekv;
KVCacheStorageBuilder builder;
builder.config({64, false, 2048}); // default ring window for 4-arg add_layer
auto k_fp16 = std::make_shared<PlainKVTemplate<ScalarType::FP16>>(8, 128, 1, "k_fp16");
auto v_int8 = std::make_shared<PlainKVTemplate<ScalarType::INT8>>(8, 128, 2, "v_int8");
builder.add_template(k_fp16);
builder.add_template(v_int8);
// Inherits default_max_seq_capacity=2048
builder.add_layer(0, 1, 2, 1024);
auto storage = builder.build();
if (!storage) {
// Build can fail on invalid config or missing templates.
return;
}
auto& layer0 = storage->layer(0);
auto& k_plane = layer0.plane(PlaneKind::K);
k_plane.append_seq(1); // ring append if max_seq_capacity > 0
auto addr = k_plane.locate(LogicalCoord(0, 0, 0, 0));Prefill typically writes a prompt chunk (N tokens) into every layer's K/V planes, then commits length.
constexpr uint32_t NUM_LAYERS = 4;
constexpr uint32_t NUM_HEADS = 8;
constexpr uint32_t HEAD_DIM = 128;
constexpr uint32_t PROMPT_LEN = 16;
constexpr size_t TOKEN_ELEMS = static_cast<size_t>(NUM_HEADS) * HEAD_DIM;
auto storage = create_fp32_storage(NUM_LAYERS, NUM_HEADS, HEAD_DIM, 2048);
if (!storage) return;
// Reserve once for prefill.
storage->reserve_all(PROMPT_LEN);
// Your runtime output buffers for one token.
float src_k[TOKEN_ELEMS];
float src_v[TOKEN_ELEMS];
for (uint32_t layer = 0; layer < NUM_LAYERS; ++layer) {
auto& layer_storage = storage->layer(layer);
auto& k_plane = layer_storage.plane(PlaneKind::K);
auto& v_plane = layer_storage.plane(PlaneKind::V);
float* k_ptr = static_cast<float*>(k_plane.data());
float* v_ptr = static_cast<float*>(v_plane.data());
// Write prompt token-by-token (or bulk memcpy from your runtime buffer).
for (uint32_t t = 0; t < PROMPT_LEN; ++t) {
size_t token_offset = static_cast<size_t>(t) * TOKEN_ELEMS;
// Fill src_k/src_v with model output for token t before memcpy.
std::memcpy(k_ptr + token_offset, src_k, TOKEN_ELEMS * sizeof(float));
std::memcpy(v_ptr + token_offset, src_v, TOKEN_ELEMS * sizeof(float));
}
// Commit visible length for this layer.
k_plane.resize_seq(PROMPT_LEN);
v_plane.resize_seq(PROMPT_LEN);
}Decode typically appends one token per step, writes only the newest token slot, then reads a window for attention.
constexpr uint32_t DECODE_STEPS = 32;
constexpr uint32_t TOKEN_STRIDE = NUM_HEADS * HEAD_DIM;
float step_k[TOKEN_STRIDE];
float step_v[TOKEN_STRIDE];
for (uint32_t step = 0; step < DECODE_STEPS; ++step) {
// 1) Grow all layers by one token.
storage->append_all(1);
// 2) Write newly generated K/V for the newest logical token.
for (uint32_t layer = 0; layer < NUM_LAYERS; ++layer) {
auto& layer_storage = storage->layer(layer);
auto& k_plane = layer_storage.plane(PlaneKind::K);
auto& v_plane = layer_storage.plane(PlaneKind::V);
uint32_t newest = k_plane.stats().seq_length - 1; // logical index in current window
auto k_addr = k_plane.locate(LogicalCoord(layer, newest, 0, 0));
auto v_addr = v_plane.locate(LogicalCoord(layer, newest, 0, 0));
float* k_base = reinterpret_cast<float*>(static_cast<Byte*>(k_plane.data()) + k_addr.byte_offset);
float* v_base = reinterpret_cast<float*>(static_cast<Byte*>(v_plane.data()) + v_addr.byte_offset);
// Fill step_k/step_v with model output for this decode step before memcpy.
std::memcpy(k_base, step_k, TOKEN_STRIDE * sizeof(float));
std::memcpy(v_base, step_v, TOKEN_STRIDE * sizeof(float));
}
// 3) Acquire attention window view (window-local logical indices).
auto& layer0_k = storage->layer(0).plane(PlaneKind::K);
uint32_t len = layer0_k.stats().seq_length;
uint32_t win = std::min<uint32_t>(len, 512);
uint32_t begin = len - win;
AccessView view = layer0_k.acquire_seq_view(begin, win, AccessMode::ReadOnly);
layer0_k.release_seq_view(view);
}Notes:
- In ring mode,
seqindices are always window-local[0, seq_length). locate()already resolves logical-to-physical mapping after wrap-around.- If
acquire_seq_view(...)returnscontiguous=false, treat it as non-contiguous window access in your runtime path.
include/mobilekv/kv_cache_convenience.h provides:
create_simple_storage(...)create_complex_storage(...)create_fp32_storage(...),create_fp16_storage(...),create_int8_storage(...)create_storage_from_init_config(...)load_storage_init_config_from_file(...)create_storage_from_config_file(...)KVAccessor<T>with runtime scalar-type validation
load_storage_init_config_from_file(...) supports a simple line-based format:
model num_heads=32 head_dim=128
storage default_alignment=64 thread_safe=false default_max_seq_capacity=0
defaults k_type=fp16 v_type=fp16 initial=512 max=2048
group 0-31 k_type=int8 max_k=1024
override 7 v_type=uint8 initial_v=256 max_v=1024k_type/v_type/type currently support:
- Built-in:
fp32,fp16,bf16,int8,uint8,int16 - Registered named custom types (for example
int8_pack4)
Named custom types in cfg:
- Generic token
customis rejected. - Register custom type in code first, then reference by name in cfg.
- Custom types in cfg path are materialized as dim-block templates.
head_dimmust be divisible by registereddim_pack_factor.
Register + create example:
ConfigTypeRegistry registry;
std::string error;
registry.register_type({"int8_pack4", 4, 4, 4}, &error);
auto storage = create_storage_from_config_file("mobilekv.cfg", ®istry, &error);
if (!storage) {
// handle parse/build error
}For built-in-only cfg (no named custom types), you can build directly:
std::string error;
auto storage = create_storage_from_config_file("mobilekv.cfg", &error);
if (!storage) {
// handle parse/build error
}Rule precedence (low to high):
defaultsgroupoverride(and legacylayeralias)
Layer resolution:
- Layers are materialized from
group/override/layerselectors. - If a plane max is still
0, it inheritsstorage.default_max_seq_capacitywhen that value is non-zero.
Accessor compatibility:
KVAccessor<float>->FP32KVAccessor<uint16_t>->FP16/BF16raw storage bitsKVAccessor<int8_t>->INT8KVAccessor<uint8_t>->UINT8KVAccessor<int16_t>->INT16
Template config contains FormatDescriptor metadata to describe storage format for downstream runtime/kernel decisions.
See:
This is metadata only; MobileKV does not perform quant/dequant compute.
For pre-packed D-last storage (for example pack4 in [H, S, D/4] block space), use builder-level opaque scalar registration.
This is the recommended end-to-end path:
KVCacheStorageBuilder builder;
builder.config({64, false, 2048});
OpaqueScalarId pack4 = builder.register_opaque_scalar({"int8_pack4", 4, 4});
if (pack4 == 0) {
return; // invalid descriptor
}
auto k_t = builder.make_dim_block_template(32, 128 / 4, pack4, 1, "k_pack4");
auto v_t = builder.make_dim_block_template(32, 128 / 4, pack4, 2, "v_pack4");
if (!k_t || !v_t) {
return; // unknown scalar id
}
builder.add_template(k_t);
builder.add_template(v_t);
builder.add_layer(0, 1, 2, 1024, 2048);
auto storage = builder.build();
if (!storage) {
return; // invalid layer/template wiring
}Notes:
register_opaque_scalar(...)returns0on invalid descriptor (bytes==0oralignment==0).make_dim_block_template(...)returnsnullptrfor unknown scalar IDs.dim_blocksis already block-space (D/4), not rawD.
Configure and build:
cmake -S . -B build_mobilekv
cmake --build build_mobilekv -jRun tests:
./build_mobilekv/kv_cache_testRun examples:
./build_mobilekv/mobilekv_demo
./build_mobilekv/mobilekv_mixed_precision_demo
./build_mobilekv/mobilekv_dim_block_demo
./build_mobilekv/mobilekv_convenience_demo
./build_mobilekv/mobilekv_cfg_basic_demo ./example/configs/basic.cfg
./build_mobilekv/mobilekv_cfg_multi_demo ./example/configs/low.cfg ./example/configs/high.cfg
./build_mobilekv/mobilekv_cfg_custom_demo ./example/configs/custom_type.cfgRun benchmark:
./build_mobilekv/mobilekv_benchmarkBenchmark includes both:
- Ring steady-state cases
- Growth-stress cases (non-ring expansion pressure)
example/fp32_prefill_decode_example.cppexample/mixed_precision_example.cppexample/dim_block_example.cppexample/convenience_api_example.cppexample/config_file_basic_example.cppexample/config_file_multi_cfg_example.cppexample/config_file_custom_type_example.cppexample/configs/basic.cfgexample/configs/low.cfgexample/configs/high.cfgexample/configs/custom_type.cfg
Unit tests currently validate:
- Template addressing and byte sizing
- Mixed precision K/V assignment
- Dim-block layout locate/view correctness
- Ring-buffer logical-to-physical mapping and wrap behavior
- Default ring configuration inheritance and explicit override
- Strict build failure for invalid
initial > max - Accessor scalar-type mismatch rejection
For simple single-request or light batching workloads, current ring-window storage is usable. For large-scale continuous batching runtimes, add a separate scheduling/index indirection layer above MobileKV (request-token mapping, gather plans, block tables) to avoid coupling runtime policy with storage internals.