Impact: High · Priority: P6 · Go client · Fix size: Small + larger
Where. The inline completion is at oxia/batch/batcher.go:77, and the linger-0 path is at :93. The RPC blocks in oxia/internal/batch/write_batch.go:101 and oxia/internal/write_stream.go:63. The queue size is set in oxia/batch/batcher_factory.go:26, and oxia/sync_client_impl.go:41 forces linger to 0.
Mechanism. The per-shard batcher goroutine calls batch.Complete() inline, which sends the request and waits for the response. While it waits, no new batch forms. Once the call queue fills, callers of Put, Delete and Get block inside Add, and that queue holds only GOMAXPROCS calls. SyncClient sets linger to 0, so every call becomes its own single-operation batch. The result is one operation in flight per shard, no matter how many goroutines call the client.
When it bites. Every Go client is affected, and SyncClient under concurrency is the worst case. Its per-shard throughput is roughly one operation per round trip, and that round trip includes the leader's fsync and replication.
Fix, cheap. When linger is 0, drain the calls already waiting in callC without blocking, then complete the batch. This batches concurrent callers naturally and adds no latency.
Fix, larger. Separate sending from completion. Send in order from the batcher goroutine and complete batches asynchronously, with a bound on in-flight batches per shard. Per-client write order is a documented guarantee since the op_index work (#834), so sends must stay ordered on the stream. A stream failure must fail or resend the pending batches in order, not let each batch retry on its own.
From the Oxia performance review of main at cf2a644. Line numbers refer to that commit.
Impact: High · Priority: P6 · Go client · Fix size: Small + larger
Where. The inline completion is at
oxia/batch/batcher.go:77, and the linger-0 path is at:93. The RPC blocks inoxia/internal/batch/write_batch.go:101andoxia/internal/write_stream.go:63. The queue size is set inoxia/batch/batcher_factory.go:26, andoxia/sync_client_impl.go:41forces linger to 0.Mechanism. The per-shard batcher goroutine calls
batch.Complete()inline, which sends the request and waits for the response. While it waits, no new batch forms. Once the call queue fills, callers of Put, Delete and Get block insideAdd, and that queue holds onlyGOMAXPROCScalls.SyncClientsets linger to 0, so every call becomes its own single-operation batch. The result is one operation in flight per shard, no matter how many goroutines call the client.When it bites. Every Go client is affected, and
SyncClientunder concurrency is the worst case. Its per-shard throughput is roughly one operation per round trip, and that round trip includes the leader's fsync and replication.Fix, cheap. When linger is 0, drain the calls already waiting in
callCwithout blocking, then complete the batch. This batches concurrent callers naturally and adds no latency.Fix, larger. Separate sending from completion. Send in order from the batcher goroutine and complete batches asynchronously, with a bound on in-flight batches per shard. Per-client write order is a documented guarantee since the op_index work (#834), so sends must stay ordered on the stream. A stream failure must fail or resend the pending batches in order, not let each batch retry on its own.
From the Oxia performance review of main at cf2a644. Line numbers refer to that commit.