Skip to content

feat: L2 Cache with Valkey Client Side Hash Ring - #4033

Merged
szuecs merged 90 commits into
zalando:masterfrom
larry-dalmeida:feat/cache-client-side-valkey-hash-ring
Sep 2, 2026
Merged

szuecs merged 90 commits into
zalando:masterfrom
larry-dalmeida:feat/cache-client-side-valkey-hash-ring

Conversation

@larry-dalmeida

@larry-dalmeida larry-dalmeida commented May 26, 2026 •

Copy link
Copy Markdown
Collaborator

Related Issue

Follow up of #3991

Description

Extends the cache() filter with an optional Valkey-backed L2 cache using a client-side consistent hash ring. When Valkey is configured (via --swarm-valkey-urls, --kubernetes-valkey-service-name, or --swarm-valkey-endpoints-remote-url with --enable-swarm), responses are stored in Valkey (L2) with the in-process LRU (L1) serving as a look-aside cache in front of Valkey, warmed lazily on Valkey hits and successful stores.

On Valkey write errors, the filter falls back transparently to L1. On Valkey read errors, the request is treated as a cache miss and fetched from origin. Valkey delete errors are best-effort - the local L1 entry is always removed, but other Skipper instances in the fleet retain their own L1 copies until each entry's warmed TTL (bounded by --cache-l1-ttl) expires naturally.

Why? Valkey is a network call - it can fail due to timeouts, connection drops, or shard unavailability. The filter is designed to degrade gracefully rather than return errors to the requester. Without Valkey configuration, the filter operates as a pure in-process LRU cache.

Storage architecture

flowchart TD
    A[Incoming Request] --> B["Skipper cache() filter"]
    B --> C{L1 LRUStorage}
    C -->|l1_hit — fresh hit| S[Entry Served]
    C -->|miss| D{ValkeyStorage L2}
    C -->|stale hit — no l1_hit| D
    D -->|l2_hit — key found| S
    D -->|valkey_miss — key absent| E[Origin Fetch\nContentful CDN / Pegasus]
    D -->|valkey_get_error — error / timeout| E
    E -->|write back: Set warms Valkey + L1<br/>min TTL: cache-l1-ttl vs entry.TTL| S
Loading

Write path: successful Valkey Set warms L1 with min(--cache-l1-ttl, entry.TTL) (default 60s). Valkey Get hit warms L1 with min(--cache-l1-ttl, remaining freshness). L1 is also populated on Valkey Set errors (fallback). Set --cache-l1-ttl=0 for write-around.

Read path: L1 checked first. Fresh L1 hit → returns immediately (increments l1_hit), no Valkey call. Stale L1 hit (past TTL but within the stale retention window) → falls through to Valkey without incrementing l1_hit. L1 miss → Valkey. Valkey hit → increments l2_hit, warms L1 with min(--cache-l1-ttl, remaining freshness) when --cache-l1-ttl > 0, returns entry. Valkey miss (nil) → cold miss (increments valkey_miss). Valkey error → increments valkey_get_error, treated as a cold miss.

Caching semantics

Each stored entry has three time zones relative to its CreatedAt timestamp:

Zone Range Behaviour
Fresh [0, TTL) Served directly from cache
Stale-while-revalidate [TTL, TTL + StaleWhileRevalidate) Served from cache; background revalidation against origin enqueued. Some directives (must-revalidate, proxy-revalidate, no-cache, s-maxage) bypass this and force a synchronous origin fetch.
Stale-if-error [TTL + SWR, TTL + max(SIE, SWR)) Only exists when StaleIfError > StaleWhileRevalidate. Origin is always contacted (coalesce); the stale entry is served as a fallback only if origin returns 5xx.

Valkey expiry: entries are stored in Valkey for TTL + max(StaleIfError, StaleWhileRevalidate), ensuring the key outlives both stale windows.

L1 TTL: capped to min(--cache-l1-ttl, entry.TTL) on the write path, and min(--cache-l1-ttl, remaining freshness) on the read-promotion path. In both cases L1 never holds an entry beyond Valkey's authoritative freshness window for the fresh portion. L1 warming is skipped for TTL=0 entries (no-cache / conditional-revalidation-only entries) to avoid polluting L1 with entries that must not be served directly.

StaleIfError scope: StaleIfError activates on the coalesce path — cold misses and SIE-zone re-hits both go through coalesce, which snapshots any eligible stale entry before the origin fetch and serves it on 5xx. StaleIfError does not activate on the background revalidation path (doRevalidate): if the background fetch fails with a transport error, the error is logged and the stored entry is left unchanged. If the background fetch returns a 5xx HTTP response, it may be stored (subject to errorTTL) — StaleIfError does not apply as a fallback on this path. Valkey errors are treated as cache misses — StaleIfError does not protect against Valkey unavailability.

L1 staleness and cross-instance consistency: ValkeyStorage.Get only returns an L1 entry if it is still fresh (!IsStale). Entries in the stale-while-revalidate window fall through to Valkey, so a fresher copy written by another instance is not bypassed. Entries in the SIE-only zone are returned from L1 (IsStale is false outside the SWR window), but filter.go calls coalesce for them regardless, which re-reads from Valkey when snapshotting the SIE candidate.

Valkey ring topology

flowchart LR
    subgraph Pods
        P1[Pod A]
        P2[Pod B]
        P3[Pod C]
    end

    subgraph Ring["Valkey Hash Ring (client-side)"]
        direction LR
        V1[(Shard 0)]
        V2[(Shard 1)]
        V3[(Shard 2)]
    end

    P1 -- "key → consistent hash → same shard" --> V1
    P2 -- "same key → same shard" --> V1
    P3 --> V2
    P1 --> V3
Loading

All pods share the same ring, so a response stored by pod A lands in the same shard that pod B would read from - cross-pod cache sharing is guaranteed by consistent hashing. However, there is no cross-pod thundering herd protection: if N pods miss simultaneously before any one of them completes the fetch and writes to Valkey, all N will forward to origin. This is most likely at cold start, TTL expiry on high-traffic keys, or during Valkey shard failures.

Observability

Counter Meaning
hit Fresh entry served from cache (filter layer; incremented regardless of whether L1 or L2 was the source)
stale Stale entry served; background revalidation enqueued
miss Cache miss; origin fetched
l1_hit L1 returned a warm entry; Valkey not consulted (Valkey mode only; a fresh L1 hit increments both l1_hit and hit)
l2_hit Valkey returned the entry; L1 warmed as a side-effect when --cache-l1-ttl > 0
valkey_miss Key absent in Valkey (clean miss)
valkey_get_error Valkey error on Get; treated as a cache miss
valkey_set_fallback Valkey error on Set; L1 written instead
lru_oversized Entry exceeded shard capacity and was not stored
lru_eviction Entry evicted from LRU (capacity pressure)
storage_error Storage Set/Delete failed (Valkey or L1); entry may not have been persisted
coalesce_error Singleflight fetch failed or returned nil; request not served from cache
reval_error Background revalidation fetch failed (transport error or unreadable body)
reval_dropped Background revalidation job dropped; revalidation queue full or filter shutting down

lru_bytes gauge is updated by a background scraper every 10s instead of only on eviction, so it stays accurate when capacity is not exceeded.

Storage Set and Delete errors are now logged at Warn instead of being silently discarded.

@larry-dalmeida
larry-dalmeida force-pushed the feat/cache-client-side-valkey-hash-ring branch 2 times, most recently from 7047126 to 573ebdf Compare May 27, 2026 04:40
@larry-dalmeida larry-dalmeida changed the title Add L2 Cache with Valkey Client Side Hash Ring feat: Add L2 Cache with Valkey Client Side Hash Ring May 27, 2026
@larry-dalmeida larry-dalmeida changed the title feat: Add L2 Cache with Valkey Client Side Hash Ring feat: L2 Cache with Valkey Client Side Hash Ring May 27, 2026
@larry-dalmeida
larry-dalmeida force-pushed the feat/cache-client-side-valkey-hash-ring branch from 55ba444 to c73eb72 Compare May 27, 2026 09:55
@larry-dalmeida
larry-dalmeida marked this pull request as ready for review May 28, 2026 07:56
@szuecs szuecs added the architectural all changes in the hot path, big changes in the control plane, control flow changes in filters label May 29, 2026
Comment thread skipper.go Outdated
@szuecs

szuecs commented Jun 1, 2026

Copy link
Copy Markdown
Member

👍

@szuecs
szuecs requested review from MustafaSaber and a4180p June 1, 2026 20:02
@larry-dalmeida

Copy link
Copy Markdown
Collaborator Author

👍

@a4180p

a4180p commented Jun 2, 2026

Copy link
Copy Markdown
Member

Write path: successful Valkey Set does not warm L1 (write-around). L1 is only populated on Valkey errors.
Read path: Valkey miss → nil (clean miss, no L1 consulted). Valkey error → L1 consulted as fallback.

That actually implies that l1 cache will be empty in case if Set was successfull and only Get fails.
Also, this model means that cache will be slower - there is no fast l1 layer, positive path always reads and writes from/to network storage.
Is that desired behaviour?
I would expect a small write/read-through L1 cache to serve as a fast local layer for the common case, but I may be missing some context around the intended use case.

@szuecs

szuecs commented Jun 2, 2026

Copy link
Copy Markdown
Member

Write path: successful Valkey Set does not warm L1 (write-around). L1 is only populated on Valkey errors.
Read path: Valkey miss → nil (clean miss, no L1 consulted). Valkey error → L1 consulted as fallback.

That actually implies that l1 cache will be empty in case if Set was successful and only Get fails. Also, this model means that cache will be slower - there is no fast l1 layer, positive path always reads and writes from/to network storage. Is that desired behaviour? I would expect a small write/read-through L1 cache to serve as a fast local layer for the common case, but I may be missing some context around the intended use case.

Good point, I wondered about this and missed to ask the question.
Normally (for example CPU) you would check L1 cache and if you get a cache miss, you check L2 cache and if you have a hit you will save in L1 and respond.

Comment thread skipper.go Outdated
@larry-dalmeida

larry-dalmeida commented Jun 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

@a4180p @szuecs @MustafaSaber

Intended behavior: L1 is a degraded-mode fallback only, not hot cache layer path. Happy path always reads from and writes to L2. L1 is only populated when a valkey operation fails.

CPI L1/L2 cache hierarchy

works because 1) L2 is on-chip with 5-30ns latency. L1 saves nanoseconds, not milliseconds. 2) Cache coherence is handled at hardware level (MESI protocol) - when core writes to L1, hardware broadcasts invalidation to other core’s L1 automatically. This does not transfer here:

L1 is an LRU per Skipper pod - there can be n pods, each with own private LRU. No hardware coherence protocol

Valkey round trip is ~0.5-1ms over local cluster network - latency is higher than L1 cache but problem being solved is not sub-ms latency but rather cross-pod cache sharing.

Primary goal: cross-pod consistency

Without L2, every pod maintains an independent LRU. A cold miss on pod A fetches from origin, pod B's LRU is empty and fetches independently. Under load (e.g. a popular campaign going live), this produces an N-way thundering herd - one upstream fetch per pod, not one per cluster.

Valkey solves this because the consistent-hash ring maps a given cache key to the same shard regardless of which pod is making the request. A cold-miss coalesced by pod A writes to Valkey shard S. Pod B's next request for the same key hits shard S directly - no upstream fetch.

Why warming L1 on write undermines this goal

If L1 is warmed on a successful Valkey Set (write-through):

  1. A pod serves future requests from its local LRU, bypassing Valkey.
  2. Valkey Delete or TTL-based expiry does not reach L1 - the pod continues serving stale content until L1's own TTL expires.
  3. There is no invalidation channel. Adding one (e.g. Valkey pub/sub) would require every pod to subscribe, handle reconnects, and accept a bounded staleness window on subscriber lag.

The tradeoff is: faster reads on the hot path vs. stale content served after invalidation, with non-trivial invalidation infrastructure.

Why Valkey-miss does not consult L1

A Valkey miss (nil, no error) means the key is genuinely absent - no pod has fetched and stored it yet, or the TTL expired. Consulting L1 on a clean miss would serve stale content beyond the intended TTL. The filter must go to origin.

L1 is only consulted when Valkey returns an error, because in that case we have no authoritative answer. Serving a potentially-stale L1 entry is preferable to a 5xx.

Considered alternatives

Write-through with TTL-bounded staleness

Warm L1 on every successful Set, accept that L1 entries can be served up to their TTL after a Valkey Delete. For content that is never explicitly invalidated (only TTL-expired), this is semantically equivalent to the current design - and would reduce Valkey read load.

Not chosen because explicit Delete is on the roadmap (cache invalidation on content publish events). Once that lands, write-through without an invalidation channel produces observable stale responses.

Write-through + Valkey pub/sub invalidation

Warm L1, subscribe each pod to a Valkey pub/sub channel for invalidation events. This matches the CPU L1/L2 mental model most closely.

Not chosen for this PR. The complexity cost is high (subscribe lifecycle, reconnect handling, message delivery guarantees, lag-bounded staleness), and the latency benefit does not yet justify it. This is the natural next step if Valkey read latency becomes a bottleneck.

@szuecs

szuecs commented Jun 2, 2026

Copy link
Copy Markdown
Member

Considered alternatives

Write-through with TTL-bounded staleness

Warm L1 on every successful Set, accept that L1 entries can be served up to their TTL after a Valkey Delete. For content that is never explicitly invalidated (only TTL-expired), this is semantically equivalent to the current design - and would reduce Valkey read load.

Not chosen because explicit Delete is on the roadmap (cache invalidation on content publish events). Once that lands, write-through without an invalidation channel produces observable stale responses.

We could have L1 entry TTL of a fixed acceptable amount of time.
We have for example token result caching in tokeninfo filters set to 60s and it offloaded 50% of total CPU time and latency dropped from 10-20ms to <1ms.
Given our experience I would be in favour of having a fixed TTL of 60s for L1 cache in case there is an L2 cache setup configured.

@larry-dalmeida

larry-dalmeida commented Jun 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

Given our experience I would be in favour of having a fixed TTL of 60s for L1 cache in case there is an L2 cache setup configured.

Good point. tokeninfo data is a strong precedent.

The write-around choice was conservative: L1 and Valkey TTLs are independent, and if a Valkey entry expires or gets evicted, an L1 entry with a longer TTL would silently serve stale content with no signal. The intent was to keep Valkey authoritative for the lifetime of every entry.

That said, your proposal sidesteps the problem cleanly.
A fixed short TTL (e.g. 60 s) on L1 should result in:

  • staleness is bounded and operator-visible - it's a deliberate configuration choice, not a silent race
  • the Delete path already calls l1.Delete unconditionally, so explicit invalidations would propagate correctly
  • no invalidation channel needed

Trade-off is: cache filter TTL must be meaningfully longer than the L1 TTL for the L1 layer to be useful.
For short-lived entries (< 60 s) the L1 TTL would need to be proportionally smaller, so it is likely to be configurable rather than hard-coded.

For now I will proceed with setting 60s as fixed TTL as a start.

@larry-dalmeida
larry-dalmeida marked this pull request as draft June 2, 2026 14:21
@a4180p

a4180p commented Jun 2, 2026

Copy link
Copy Markdown
Member

@larry-dalmeida

The intent was to keep Valkey authoritative for the lifetime of every entry.

After every successfull read from L1 you can call EXPIRE to valkey comand to:

  • prolongate TTL of the record if it exists
  • get a signal if it not exists and cleanup the record in L1

You can also consider using GETEX instead of GET if you wan TTLs to be updated on read ops.

@szuecs

szuecs commented Jun 2, 2026 •

Copy link
Copy Markdown
Member

@larry-dalmeida

The intent was to keep Valkey authoritative for the lifetime of every entry.

After every successfull read from L1 you can call EXPIRE to valkey comand to:

This you can't really do because l2 cache is shared by all skipper instances
In the rate limit cases we set Expiry time on create entry. I think it's the way to go for expiration by time.

@larry-dalmeida

larry-dalmeida commented Jun 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

@szuecs @a4180p @MustafaSaber thanks a lot for the feedback and patience 🧡 I will review it thoroughly and get back to you with concrete proposal. Currently I am busy with a business critical reliability related topic but I will resume this on Monday 15th June.

@larry-dalmeida

Copy link
Copy Markdown
Collaborator Author

Thanks for your patience folks, our 🔥 are out, resuming this today.

@larry-dalmeida
larry-dalmeida force-pushed the feat/cache-client-side-valkey-hash-ring branch 2 times, most recently from 7d1af73 to c1fb3cd Compare June 18, 2026 07:19
@larry-dalmeida

Copy link
Copy Markdown
Collaborator Author

Update: fixing final failing tests. PR will be ready for review shortly.
@MustafaSaber @a4180p @szuecs

@larry-dalmeida
larry-dalmeida force-pushed the feat/cache-client-side-valkey-hash-ring branch from 1071146 to 8760752 Compare June 22, 2026 06:52
@larry-dalmeida
larry-dalmeida marked this pull request as ready for review June 22, 2026 11:55
@larry-dalmeida

Copy link
Copy Markdown
Collaborator Author

@szuecs ran all tests locally - all pass locally. Let me know if you have concerns.

…onsistency

Signed-off-by: Larry D Almeida <hello@larrydalmeida.com>
@larry-dalmeida

Copy link
Copy Markdown
Collaborator Author

In the description

valkey_miss 	Key absent in Valkey (clean miss)
valkey_get_error 	Valkey error on Get; treated as a cache miss
valkey_set_fallback 	Valkey error on Set; L1 written instead

please rename these to l2_* instead of valkey_ there is already one metric starting with l2_

Also this line makes me think that it's not write-through cache (i.e. write to L2 always means a write to L1). Is that true?

Valkey error on Set; L1 written instead

Fixed - renamed valkey_miss, valkey_get_error, valkey_set_fallback to l2_miss, l2_get_error, l2_set_fallback for consistency. On the write-through question: yes, a successful Set writes to both L2 and L1. l2_set_fallback fires only when the L2 write fails - in that case L1 is written instead (not L2). Updated the description to make this explicit.

Comment thread docs/operation/operation.md
Comment thread filters/cache/valkey_storage.go Outdated
Comment thread filters/cache/valkey_storage.go Outdated
Comment thread filters/cache/valkey_storage.go Outdated
Comment thread skipper.go Outdated
…cache other than having the expose interface to make sure it has all its needs. Like this everyone can pass their own implementation

Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
…ent if not set

Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
…ting the real proxy with the filter in a route

Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
Signed-off-by: Sandor Szücs <sandor.szuecs@zalando.de>
@szuecs

szuecs commented Sep 1, 2026

Copy link
Copy Markdown
Member

I refactored a bit the code as I saw in my review. I did a bit more refactoring to drop out Valkey from the ./filter/cache package. Valkey is only used in tests, but not required so library users could use whatever they want as L2 storage as long as they implement cache.L2Client interface. This will make it easy to support also Redis for example or whatever people can think of.
I instructed claude to read coverage.out to increase coverage and used the skipper test skill to write some more advanced tests.
Last I added storage architecture docs from the nice PR description and updated the entries with cache. prefix and tried to be consistent in the use of L2 vs Valkey, so we focus not on the implementation but on the idea.

From my side we are good to go.

@szuecs

szuecs commented Sep 1, 2026

Copy link
Copy Markdown
Member

👍

1 similar comment
@larry-dalmeida

Copy link
Copy Markdown
Collaborator Author

👍

@szuecs
szuecs merged commit 0d0ac17 into zalando:master Sep 2, 2026
19 checks passed
@szuecs

szuecs commented Sep 2, 2026

Copy link
Copy Markdown
Member

Thanks for this contribution!


// testMetrics is a minimal metrics.Metrics stub for testing.
// Only IncCounter does real work; all other methods are no-ops.
type testMetrics struct {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Shouldn't this use mockMetrics?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

architectural all changes in the hot path, big changes in the control plane, control flow changes in filters documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants