(h/t to my teammate @tanaykm for finding this issue)
Summary
On a single node, a LogQL query that matches a stream with multiple log lines
returns only one line for that stream instead of all of them. The
regression was introduced by the label-resolution join change in
fix(logql): drop log entries whose fingerprint has no labels yet (commit
2edfa85c71ee834fab50a42ee4f6d8f7ae13b3b6, released in v5.4.3).
Both /loki/api/v1/query and /loki/api/v1/query_range are affected, because
they share the same label-resolution planner.
Affected version
- v5.4.3 (and any build containing commit
2edfa85).
- Earlier versions are not affected.
Root cause
LabelsJoinPlanner.Process
(reader/logql/logql_transpiler/clickhouse_planner/planner_labels_joiner.go)
resolves stream labels by joining the samples subquery (main) to the
time-series/labels subquery (_time_series) on fingerprint.
The commit changed the single-node join from ANY LEFT JOIN to
ANY INNER JOIN to drop samples whose fingerprint has no labels row yet:
SELECT main.fingerprint, main.timestamp_ns, _time_series.labels,
main.string, main.value
FROM main
ANY INNER JOIN _time_series
ON main.fingerprint = _time_series.fingerprint
The relationship is one-to-many: main has one row per log line (many rows
per fingerprint), while _time_series has one row per fingerprint. In
ClickHouse, the ANY modifier keeps at most one matched pair per join key, so
ANY INNER JOIN collapses all the sample rows sharing a fingerprint down to a
single row. Every multi-line stream is reduced to one line.
The previous ANY LEFT JOIN preserved every left-hand (sample) row, so all
lines of a stream survived. Switching the join type conflated two different
goals — "drop rows whose fingerprint has no labels" and "keep only one row per
key" — and only the first was intended.
Note the cluster-mode branch was fixed differently, with an explicit
WHERE length(labels) > 0 filter, and does not exhibit this collapse. Only
the single-node path regressed.
End-to-end reproduction (single-node gigapipe)
- Run a single-node gigapipe (
IsCluster = false) backed by ClickHouse.
- Push three log lines that share the same label set (one stream), e.g. via
POST /loki/api/v1/push with three entries under a single stream such as
{k8s_node_name="test_node1"}, plus one line under a second stream
{k8s_node_name="test_node2"}.
- Query them back within the query window:
GET /loki/api/v1/query?query={k8s_node_name=~"test_node[12]"}
(a range query over the same selector shows the same result).
Expected: the test_node1 stream returns three entries; test_node2
returns one.
Actual: each stream returns exactly one entry. test_node1 is missing two
of its three lines.
Impact
Any single-node deployment silently drops log lines whenever a stream has more
than one line in the query window. This affects log browsing and any consumer
that reads raw stream entries, including downstream tooling that expects the
full set of lines for a selector.
Suggested fix
Resolve labels without collapsing the many-side. Keep the samples-preserving
ANY LEFT JOIN and drop unresolved rows with an explicit filter, mirroring the
cluster branch:
FROM main
ANY LEFT JOIN _time_series
ON main.fingerprint = _time_series.fingerprint
WHERE length(_time_series.labels) > 0
This still discards samples whose fingerprint has no labels row yet (the
original intent of the commit) while preserving every line of a resolved
stream. Verified against ClickHouse: the WHERE length(...) > 0 filter drops
only the genuinely unresolved rows and returns all three lines of the
multi-line stream.
(h/t to my teammate @tanaykm for finding this issue)
Summary
On a single node, a LogQL query that matches a stream with multiple log lines
returns only one line for that stream instead of all of them. The
regression was introduced by the label-resolution join change in
fix(logql): drop log entries whose fingerprint has no labels yet(commit2edfa85c71ee834fab50a42ee4f6d8f7ae13b3b6, released in v5.4.3).Both
/loki/api/v1/queryand/loki/api/v1/query_rangeare affected, becausethey share the same label-resolution planner.
Affected version
2edfa85).Root cause
LabelsJoinPlanner.Process(
reader/logql/logql_transpiler/clickhouse_planner/planner_labels_joiner.go)resolves stream labels by joining the samples subquery (
main) to thetime-series/labels subquery (
_time_series) onfingerprint.The commit changed the single-node join from
ANY LEFT JOINtoANY INNER JOINto drop samples whose fingerprint has no labels row yet:The relationship is one-to-many:
mainhas one row per log line (many rowsper fingerprint), while
_time_serieshas one row per fingerprint. InClickHouse, the
ANYmodifier keeps at most one matched pair per join key, soANY INNER JOINcollapses all the sample rows sharing a fingerprint down to asingle row. Every multi-line stream is reduced to one line.
The previous
ANY LEFT JOINpreserved every left-hand (sample) row, so alllines of a stream survived. Switching the join type conflated two different
goals — "drop rows whose fingerprint has no labels" and "keep only one row per
key" — and only the first was intended.
Note the cluster-mode branch was fixed differently, with an explicit
WHERE length(labels) > 0filter, and does not exhibit this collapse. Onlythe single-node path regressed.
End-to-end reproduction (single-node gigapipe)
IsCluster = false) backed by ClickHouse.POST /loki/api/v1/pushwith three entries under a single stream such as{k8s_node_name="test_node1"}, plus one line under a second stream{k8s_node_name="test_node2"}.GET /loki/api/v1/query?query={k8s_node_name=~"test_node[12]"}(a range query over the same selector shows the same result).
Expected: the
test_node1stream returns three entries;test_node2returns one.
Actual: each stream returns exactly one entry.
test_node1is missing twoof its three lines.
Impact
Any single-node deployment silently drops log lines whenever a stream has more
than one line in the query window. This affects log browsing and any consumer
that reads raw stream entries, including downstream tooling that expects the
full set of lines for a selector.
Suggested fix
Resolve labels without collapsing the many-side. Keep the samples-preserving
ANY LEFT JOINand drop unresolved rows with an explicit filter, mirroring thecluster branch:
This still discards samples whose fingerprint has no labels row yet (the
original intent of the commit) while preserving every line of a resolved
stream. Verified against ClickHouse: the
WHERE length(...) > 0filter dropsonly the genuinely unresolved rows and returns all three lines of the
multi-line stream.