Skip to content
Permalink

Comparing changes

Choose two branches to see what’s changed or to start a new pull request. If you need to, you can also or learn more about diff comparisons.

Open a pull request

Create a new pull request by comparing changes across two branches. If you need to, you can also . Learn more about diff comparisons here.
base repository: metrico/gigapipe
Failed to load repositories. Confirm that selected base ref is valid, then try again.
Loading
base: v5.5.1
Choose a base ref
...
head repository: metrico/gigapipe
Failed to load repositories. Confirm that selected head ref is valid, then try again.
Loading
compare: v5.5.2
Choose a head ref
  • 11 commits
  • 39 files changed
  • 4 contributors

Commits on Sep 21, 2026

  1. Change runner from self-hosted to ubuntu-latest (#1001)

    Signed-off-by: rita7lopes <105070795+rita7lopes@users.noreply.github.com>
    rita7lopes authored Sep 21, 2026
    Configuration menu
    Copy the full SHA
    7e30c49 View commit details
    Browse the repository at this point in the history
  2. fix(promql): one definition of the range-function bucket cap (#983)

    A function that measures a change across samples (rate, irate, deriv,
    delta, idelta, resets, increase, changes) cannot answer from a single
    bucket, so its bucket has to be sized against its range rather than
    against the query's step. After #981 that rule existed in three places,
    with three different lists of function names and two different floors,
    and they did not agree.
    
    deriv is where the disagreement was reachable: it has no ClickHouse
    pushdown, so it goes through DownsampleHintsPlanner, and it was also in
    rateFunctions, so adjustHintsForRate had already capped its step by the
    time it got there. Below a 30s range the request layer floored the cap at
    15s -- metrics_15s rows are stamped on a 15s grid, so a finer bucket
    cannot hold a second row -- and the planner then overrode that back down
    to range/2, asking ClickHouse for buckets narrower than the table's own
    resolution. The two also disagreed on the threshold: the request layer
    capped once step passed range/2, the planner only once step reached the
    full range, so a step between the two was capped in one place and not the
    other.
    
    Consolidate both decisions into planner.NeedsDistinctSamples and
    planner.BucketResolution. BucketResolution is idempotent, so the request
    layer and the planners can both apply it to the same query without
    fighting over the answer; a test asserts the two layers agree on the
    width they settle on.
    
    Rename rateFunctions to prolongFunctions. That list is consulted by
    isProlong, which decides whether the raw iterator carries a series
    forward between steps -- a different question from which functions need a
    capped bucket, and the two must not be folded back together.
    
    Restructure DownsampleHintsPlanner into an explicit three-way switch on
    what the function needs from its window. The step > range branch was
    unreachable for change functions and read as if it still applied; it is
    the shape a reducer takes, reading only the trailing range of each step
    and snapping it forward, and it is now stated as such. That branch had no
    test at all, so its output is captured verbatim first.
    
    Behavior changes, all of them narrow:
    
      - irate and idelta are now capped at the request layer as well. Their
        SQL is unchanged -- the planner already capped them to the same width
        -- but hints.Step now reflects it.
      - The 15s floor is gone. A cap that drops the step under the 15s grid
        now trips the existing useRawData check, so a range too small for
        metrics_15s to serve is read from raw samples at real timestamps
        instead of from a bucket finer than the table's own resolution.
        Affects change functions with a range under 30s.
      - The counter path now caps at min(step, range/2) rather than leaving
        any step below the full range alone, matching what the request layer
        has always done and what #981 documents itself as doing. A step
        between range/2 and range previously kept only two grid slots inside
        (t-range, t], the outer one on its first millisecond, so a series
        that had not reported into that one bucket collapsed to a single
        sample and was dropped.
    rita7lopes authored Sep 21, 2026
    Configuration menu
    Copy the full SHA
    fd04909 View commit details
    Browse the repository at this point in the history
  3. docs: add the compression dimension to samples_v3 sort key guidance (#…

    …985)
    
    The sort-key section added in #982 framed the choice purely as scan
    selectivity: which of fingerprint or timestamp is the more selective
    predicate for a given workload. That is only half the tradeoff, and it
    closed with "gigapipe ships no benchmark comparing these layouts", which
    reads as though nothing is known about it.
    
    The storage half has in fact been measured. A codec only helps when
    adjacent rows are similar, and the sort key is what decides which rows are
    adjacent. Under the default timestamp_ns ordering consecutive rows belong
    to different series, so fingerprint is interleaved and stays close to
    incompressible whatever codec it carries. #976 reported that the column
    only compresses once ADVANCED_SAMPLES_ORDERING groups by it, and #977
    picked the shipped per-column codecs on that basis - DoubleDelta for
    timestamps in tables sorted by time, Delta for the ones that are not.
    
    Add that dimension to the guidance, and narrow the closing caveat to what
    is actually unmeasured here: scan performance across the two layouts.
    
    Also note in the rebuild procedure that SHOW CREATE TABLE is how the #977
    codecs get carried over, since hand-writing the column list is where they
    would quietly be lost.
    
    Docs only, no behaviour change.
    rita7lopes authored Sep 21, 2026
    Configuration menu
    Copy the full SHA
    aa8342b View commit details
    Browse the repository at this point in the history
  4. fix(promql): evaluate range functions over (t-range, t], not a shifte…

    …d window (#984)
    
    * fix(promql): reject a zero fill resolution instead of rendering STEP 0
    
    FillGapsPlanner.Resolution is the grid the fill densifies onto, and it has
    no safe zero value: Go zeroes it by default, so a planner built from a
    context whose Step was never set renders ORDER BY ... WITH FILL STEP 0, an
    interval no server can advance by. A negative one renders STEP -1000.
    
    Not hypothetical. TestTranspilerV2 built its PlannerContext without a Step
    and had been printing exactly that query, passing all the while because it
    asserts nothing about what it prints. The guard is what surfaced it; the
    context now carries the 15s step a real request would.
    
    The request path itself cannot reach this today -- adjustHintsForRate
    gives hints.Step a value before any planner sees it -- so this is a
    guardrail for the next caller, and the fixture it already caught.
    
    * fix(promql): evaluate range functions over (t-range, t], not a shifted window
    
    Every accelerated range function reads its samples from buckets that a
    window frame selects by key. A bucket was keyed by the floor of its
    samples' timestamps, so it held [key, key+width) -- the interval STARTING
    at its key -- and the frame at t therefore took in samples up to t+width.
    The window was shifted a whole bucket into the future, reporting data the
    caller could not yet have seen, where prometheus evaluates (t-range, t]
    and nothing after t.
    
    Key buckets by the ceiling instead. A bucket then holds (key-width, key],
    the frame ends exactly at t, and the buckets tile the range backwards from
    there. This also settles a latent one in the counter path: last_ts could
    previously exceed the row's own timestamp, driving c_fwd_edge negative.
    
    Tiling only works if a whole number of buckets spans the range, so
    BucketResolution now returns a width that divides it rather than whatever
    the query's step happened to be. A step that already divides the range is
    still used as-is; one that does not rounds down to the next width that
    does. At a 5m range a 149s step used to leave the frame reaching three
    buckets back to t-447s -- 147s of over-inclusion on a 300s window -- where
    100s tiles it exactly.
    
    OverTimePlanner now takes that width too, which is the other half of the
    same bug. It bucketed at ctx.Step however coarse, on the grounds that a
    reducer needs only one bucket in the window to have an answer. It does,
    but the answer is then computed over that whole bucket rather than over
    the range: sum_over_time(x[5m]) at a 30m step summed thirty minutes and
    labelled it five. TestOverTimeKeepsQueryStepEvenWhenCoarserThanRange
    asserted that behaviour and is replaced by one asserting the window.
    
    * test(promql): pin the keying that keeps c_fwd_edge non-negative
    
    c_fwd_edge is how far the range's last real sample sits before the
    timestamp being evaluated, and it feeds c_reach, which carries the
    observed change out to the edges of the range. last_ts is an argMax over a
    frame ending at the current row, so for a row carrying data it is that
    row's own newest sample. A bucket keyed by the floor of its samples holds
    the interval starting at its key, so that sample is routinely newer than
    the key and c_fwd_edge goes negative.
    
    Measured against the e2e fixture at a 150s bucket over 15s samples: 61 of
    62 buckets negative, worst case -135s, which is bucket minus sample
    interval exactly. The same data keyed by the ceiling gives 0 of 61.
    
    This is the fingerprint of the shifted window rather than a defect of its
    own, and it costs the counter functions nothing: the reach terms telescope,
    
        c_span + c_back + c_fwd = (last-first) + (first-(t-R)) + (t-last) = R
    
    so c_reach is R/c_span and the reported rate is c_change/c_span, the true
    slope between the two samples, whatever the shift. What is wrong is which
    samples those are -- the window ran to t+bucket rather than to t -- which
    a constant-rate fixture cannot show and count_over_time can.
    
    The guard is therefore on the keying, which is what actually constrains
    the sign, not on a value that happens to come out right.
    rita7lopes authored Sep 21, 2026
    Configuration menu
    Copy the full SHA
    719e72d View commit details
    Browse the repository at this point in the history
  5. build(deps): bump go.opentelemetry.io/collector/pdata/pprofile (#995)

    Bumps [go.opentelemetry.io/collector/pdata/pprofile](https://github.com/open-telemetry/opentelemetry-collector) from 0.160.0 to 0.161.0.
    - [Release notes](https://github.com/open-telemetry/opentelemetry-collector/releases)
    - [Changelog](https://github.com/open-telemetry/opentelemetry-collector/blob/main/CHANGELOG-API.md)
    - [Commits](open-telemetry/opentelemetry-collector@v0.160.0...v0.161.0)
    
    ---
    updated-dependencies:
    - dependency-name: go.opentelemetry.io/collector/pdata/pprofile
      dependency-version: 0.161.0
      dependency-type: direct:production
      update-type: version-update:semver-minor
    ...
    
    Signed-off-by: dependabot[bot] <support@github.com>
    Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
    dependabot[bot] authored Sep 21, 2026
    Configuration menu
    Copy the full SHA
    c9e36a9 View commit details
    Browse the repository at this point in the history

Commits on Sep 22, 2026

  1. fix(promql): key selector buckets by the ceiling, like every other re…

    …ad (#998)
    
    * fix(promql): key selector buckets by the ceiling, like every other read
    
    #984 corrected the bucket keying a BucketProducer renders, so a bucket
    holds (key-width, key] and the buckets tile backwards from the timestamp
    being evaluated. The other read of metrics_15s was left behind.
    
    A query carrying no range function is planned by DownsampleHintsPlanner's
    default branch rather than by a BucketProducer, and that branch still keyed
    by the floor. Its bucket therefore held [key, key+Step), and the value
    column over it is argMaxMerge(last) -- the newest sample in the bucket --
    so the value reported at t was the one belonging to t+Step.
    
    Measured against a reference Prometheus fed the same samples by
    remote_write, over demo_memory_usage_bytes at a 15s step: every point of
    the series was one step early, gigapipe's value[i] equal to Prometheus's
    value[i+1] for 20 of 20 consecutive pairs, and 0 of 21 matching at the same
    index. With this change the same comparison is 19 of 19 at the same index
    and 0 shifted.
    
    The keying is now one definition, bucketTimestampCol, applied at both
    sites. The width stays the caller's -- a range function sizes its bucket
    against its range, a bare selector against the query step -- but the rule
    that a bucket ends at its key is one decision, and duplicating it is what
    let #984 correct one path and leave this one behind.
    
    The default branch also serves the reducers the engine evaluates itself
    (sum_over_time and friends, when they are not accelerated); they were
    shifted by the same construction and are fixed by the same change, though
    the end-to-end measurement above is of the bare selector.
    TestDownsampleHintsLeavesPlainAggregatesAtTheQueryStep asserted the
    rendered floor expression while meaning to assert that the bucket width is
    the uncapped step; it now asserts that through the same helper, so it
    survives a change of keying.
    
    Not included: DownsampleValuesPlanner patches timestamp_ms and then
    DownsampleHintsPlanner patches the same alias over it, so the first is
    dead for this path. Left alone rather than changed silently.
    
    * fix(promql): key change-function buckets by the ceiling too
    
    bucketTimestampCol exists so the keying is one rule, but the
    NeedsDistinctSamples branch still rendered the floor form by hand, so
    deriv, irate and idelta -- the change functions with no accelerated
    planner, which reach DownsampleHintsPlanner instead -- bucketed at
    [key, key+width). Their rows are handed to the engine as if they were
    the raw series, so argMaxMerge(samples.last) reported at t the newest
    sample of [t, t+width): a window shifted a whole bucket into the future.
    It is also the interval BucketResolution already documents its width to
    be chosen against.
    
    TestDownsampleHintsCapsChangeFunctionBucket pinned the floor literal
    while asserting only the width; it now renders through the same helper.
    TestBucketCapAgreesAcrossLayers reads the width back out of the column
    instead of matching its whole text, so a keying change no longer fails
    a test that has no opinion about keying.
    
    * docs(promql): record what keeps DownsampleValuesPlanner's zero step unreachable
    
    Its timestamp_ms column is floor-keyed and divides by ctx.Step, and both
    are only harmless because DownsampleHintsPlanner patches the column away
    whenever Hints.Step is non-zero -- and because adjustHintsForRate, in
    another package, rewrites a zero step before any planner sees one. A
    reader of this file could not tell that from this file. Say it here, so
    a future caller that reaches this planner by another route knows what it
    has to uphold.
    rita7lopes authored Sep 22, 2026
    Configuration menu
    Copy the full SHA
    7316b7b View commit details
    Browse the repository at this point in the history

Commits on Sep 23, 2026

  1. deps: pin grpc to the v1.84.0 release (#1002)

    v1.85.0-dev is a pre-release tag on grpc-go master and lacks the
    :authority check from GO-2026-6443. Bump x/crypto for GO-2026-6354/6355.
    janeblower authored Sep 23, 2026
    Configuration menu
    Copy the full SHA
    2989966 View commit details
    Browse the repository at this point in the history

Commits on Sep 24, 2026

  1. chore(e2e): point at gigapipe-tests and de-duplicate its path (#992)

    * chore(e2e): point at gigapipe-tests and de-duplicate its path
    
    The e2e suite repository was renamed from qryn-test to gigapipe-tests. The
    old URL still resolves through GitHub's redirect, so nothing was broken,
    but the name was stale in both the Makefile and the CI workflow.
    
    The path was also defined twice: the workflow checked the suite out to
    ./deps/qryn-test with actions/checkout while the Makefile cloned and
    mounted the same path independently, so renaming in one place alone would
    have left CI mounting an empty directory.
    
    Hoist the repository and directory into E2E_TESTS_REPO / E2E_TESTS_DIR in
    the Makefile, and have the workflow call `make e2e-deps` instead of
    repeating the checkout. The Makefile is now the only place either value
    appears.
    
    * chore: ignore the built binary and local tooling directories
    
    `make build` and the CI build step both write a `gigapipe` binary to the
    repository root, where it sits untracked at ~70MB. Ignore it.
    
    Also carries the previously uncommitted entries for the local tooling
    directories `/.claude/` and `/.superpowers/`, which are working-tree
    scratch and should never be committed.
    rita7lopes authored Sep 24, 2026
    Configuration menu
    Copy the full SHA
    b5d7660 View commit details
    Browse the repository at this point in the history
  2. chore: use the Gigapipe name where the old names are branding (#993)

    * docs: use the Gigapipe name in prose, badges and asset links
    
    The project is Gigapipe; qryn is the former name. Update the places where
    the old name appears as branding rather than as an identifier: the README's
    badge, asset and stars URLs (metrico/qryn redirects to metrico/gigapipe, and
    all four targets were verified to return 200 under the new path), one prose
    mention, a stale comment in the version helper, and the K6 workflow's
    display name.
    
    Deliberately unchanged: the "formerly known as qryn" note, which is
    intentional history; the #qryn:matrix.org room address, which is a real
    address; and every name that code or data depends on.
    
    * chore: use the Gigapipe name in internal keys and user-facing strings
    
    Follows the prose pass with the remaining places the old name appears
    without anything depending on it:
    
    - the ctrl project registry key, a one-entry package-level map whose key is
      used only for the lookup; it is never passed to init/upgrade/rotate, is
      not operator-supplied, and never reaches ClickHouse (migration bookkeeping
      keys on `ver.k`, an int64 script identifier)
    - the CI CLUSTER workflow's display name; master's branch protection lists
      no required status checks, so no rule matches the old name
    - the collector resource attribute in Tempo trace-by-id responses
    - the writer's settings log lines
    
    Unchanged: the "cloki" default database name, which existing deployments
    that never set CLICKHOUSE_DB depend on; and the database-error hint naming
    "/cloki-writer", which points at the same binary as
    writer/config.NAME_APPLICATION — read as a file path by
    writer/metric/reload.go:26 — so the two must stay consistent.
    rita7lopes authored Sep 24, 2026
    Configuration menu
    Copy the full SHA
    3d0316d View commit details
    Browse the repository at this point in the history
  3. feat(config): accept GIGAPIPE_ environment variables (#994)

    New deployments should configure gigapipe with GIGAPIPE_-prefixed variables.
    The former QRYN_ and CLOKI_ names keep working, so existing deployments are
    unaffected.
    
    Two consumers read the old names, which is why this rewrites the environment
    rather than changing each call site. cloki-config binds every config-file key
    to an environment variable through viper using Setting.EnvPrefix, which
    defaults to "QRYN" and accepts exactly one prefix, so the new names cannot be
    added alongside the old ones there. Separately, a few call sites in this repo
    read specific names directly: the basic-auth credentials in cmd, the ruler
    settings in ruler/router, and the writer's log-file overrides.
    
    shared/envalias.Apply runs before any of that, copying each GIGAPIPE_ variable
    onto its legacy equivalent. It overwrites, so when both are set the GIGAPIPE_
    value wins. Credentials feed both legacy prefixes because cmd reads QRYN_LOGIN
    first and then lets CLOKI_LOGIN override it, so writing only one would let a
    stale value take precedence. A deployment that sets no GIGAPIPE_ variable is
    left untouched, and one that sets only a legacy name gets a deprecation
    warning naming its replacement.
    
    The e2e stack now configures gigapipe with GIGAPIPE_LOGIN, GIGAPIPE_PASSWORD
    and GIGAPIPE_RULER_ENABLED, so a green e2e run exercises the mapping against a
    real deployment rather than only in unit tests.
    rita7lopes authored Sep 24, 2026
    Configuration menu
    Copy the full SHA
    f6bb314 View commit details
    Browse the repository at this point in the history
  4. fix(logql): return the empty label set for a grouping-less aggregation (

    #1000)
    
    An aggregation with no by/without clause keeps no labels, so it yields a
    single series whose label set is empty. Since #961 that series came back
    as an empty matrix instead: HTTP 200 with result: [], reading as "no
    data" rather than an error.
    
    Two independent causes, both on that path.
    
    planAgg synthesises `by ()` for a grouping-less aggregation, which
    projects the labels to an empty map and re-fingerprints the series as a
    hash of that empty set. No time_series row carries that fingerprint, so
    the labels join appended afterwards resolves nothing and #961's
    length(labels) > 0 filter drops every row. #961 already exempted
    by/without from the drop, since an empty label set is a legitimate
    projection there; the exemption just never covered the grouping-less
    case, where the by/without is synthetic and the join is attached after
    the aggregation. Track the collapse and skip the join, which can only
    ever shed rows here: the labels are already resolved to the empty map.
    
    That alone still returns nothing. A bare map() is Map(Nothing, Nothing),
    which the driver hands back as map[*interface{}]*interface{} and fails
    to Scan into the map[string]string the reader uses. The scan error lands
    on an entry with Value == 0 and ZeroEaterPlanner drops it, so a hard
    failure surfaced as an empty result. Cast the projection to
    Map(String, String). This predates #961 -- the labels join used to
    overwrite the untyped column with a typed one, hiding it.
    
    by/without with labels, and `without ()`, are unaffected: the collapse
    is only tracked for `by` with no labels.
    
    Fixes #997
    
    Co-authored-by: 44YHC <161518859+44YHC@users.noreply.github.com>
    rita7lopes and 44YHC authored Sep 24, 2026
    Configuration menu
    Copy the full SHA
    0b50d9a View commit details
    Browse the repository at this point in the history
Loading