Skip to content

fix(o11y): move the Loki index to an index_ prefix so the compactor can run (DEV-3144) - #402

Merged
demtario merged 2 commits into
masterfrom
fix/DEV-3144-loki-index-prefix
Oct 1, 2026
Merged

demtario merged 2 commits into
masterfrom
fix/DEV-3144-loki-index-prefix

Conversation

@demtario

@demtario demtario commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Context

Part of DEV-3144. Loki's compactor never compacts because index.prefix: "index/" contains a slash: reproduced on the filesystem backend with grafana/loki:3.3.2, where the compactor logs skipping compaction since we can't find schema for table every cycle, while the same setup with index_ compacts and removes the source files. This PR appends a second schema period with prefix index_ (same store, schema and 24h period) and keeps the old period so existing data stays readable. supervisor/shutdown.sh now lists both index/index/<day>/ and index/index_<day>/ per day, still fail-closed on either listing, because records up to 7 days old land in the old-prefix table after the switch. ADR-0041 §B.4 documents the periods, the time budget (worst case about 630 s inside the documented 900 s stop grace) and when the old period can be dropped (not before 90 days after from). The exact S3 error string (invalid prefix index/<n>/browser) was not reproduced because only the filesystem path was run; the failure class is confirmed.

Deploy constraint: the new period starts at from: 2026-10-03 (00:00 UTC) and must be live in the deployed image before then. If this deploys after from, index entries already written under index/index/<day> for days on or after from become unreachable. If it will slip, change from in both Loki configs to at least one day after the real deploy date before merging. Compaction only starts working for new-prefix tables after from; old tables stay uncompacted until they age out.

Types of changes

  • New example
  • Update to an existing example
  • README / documentation change
  • Demo runner (runner/) change
  • CI / tooling change

How was this verified?

node --test on o11y-shutdown-snapshot.test.mjs and o11y-box-config.test.mjs is 46 tests, 46 pass, 0 fail, 0 skipped; with the old shutdown.sh and Loki configs restored, 8 of them fail (new schema-period tests, new dual-prefix tests, and existing tests whose call counts changed). loki -verify-config reports the config valid for both Loki configs. The full stop-roundtrip.mjs docker compose run exited 0 with all checks passing (it seeds uploader-named keys under both prefixes). Not run: the rest of the pipeline suite, and anything against R2 or MinIO's S3 path for the compactor. Because from is in the future, local runs and production write the old prefix until then, so the new prefix is covered by the seeded keys and unit tests rather than real traffic. One reasoned, not staged, risk: a compaction in flight could remove the freshly uploaded final file before the post-exit listing, which would refuse the stop as unclean (the safe direction, a replay that dedupes).

Checklist

  • New/renamed example: added to runner/config/frameworks.json (see CONTRIBUTING.md); otherwise it won't appear on demos.handsontable.com
  • New example: added a row to the tables in README.md
  • Ran pnpm build (and pnpm dev) in the affected example/server-example locally

None of the example checklist items apply; this PR changes only runner/.

Related issue(s):

  1. DEV-3144 (https://app.clickup.com/t/123kvxebzp5)

Note

Medium Risk
Changes SIGTERM clean-shutdown index confirmation and introduces a time-bound schema cutover; mis-timed deploy past from could make new index entries unreachable, though tests and fail-closed listing reduce marker false positives.

Overview
Adds a second Loki schema period (from: 2026-10-03) with index.prefix: "index_" (no slash) in both S3 and filesystem configs, while keeping the original index/ period so existing data stays readable. A slash in the old prefix prevented the compactor from mapping tables, so compaction never ran.

Shutdown / clean-stop checks now list both table layouts per day (index/index/<day>/ and index/index_<day>/), still fail-closed if any listing fails. ADR-0041 documents the migration, deploy-before-from requirement, and the ~240s worst-case listing budget per snapshot.

Tests and local harnesses follow: config pins for two periods, snapshot tests with doubled curl call counts, stub curl keys under the requested prefix, and stop-roundtrip.mjs seeds pre-existing index objects under both prefixes for the C1 negative control.

Reviewed by Cursor Bugbot for commit 682ffb1. Bugbot is set up for automated code reviews on this repo. Configure here.

demtario and others added 2 commits October 1, 2026 10:00
…ctor can run

index.prefix "index/" contains a slash, so the compactor cannot map a table back to
its schema period and skips it ("can't find schema for table"), leaving the index
uncompacted. A second schema period from 2026-10-03 uses "index_"; the first stays so
existing data remains readable. The stop check lists both prefixes per day because
records up to 7 days old still land in the old-prefix table after that date.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…time budget

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@demtario demtario self-assigned this Oct 1, 2026
@demtario
demtario merged commit 41e9433 into master Oct 1, 2026
10 checks passed
@demtario
demtario deleted the fix/DEV-3144-loki-index-prefix branch October 1, 2026 08:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant