Section 05

HiCache

Reuse beyond GPU memory

The handbook assistant is popular enough to serve hundreds of different documents. Some are used every minute; others return after an hour. Keeping all of their KV state in GPU memory would leave too little space for active users. Discarding every document as soon as it leaves the GPU would repeat expensive prefills.

HiCacheHiCacheSGLang’s hierarchical KV cache spanning GPU memory, instance-local host memory, and a storage backend that can be shared when configured.See in glossary → extends prefix reuse across a memory hierarchy. Its terminology calls GPU memory L1, host memory L2, and a storage backend L3. These are software cache tiers, not the GPU’s physical hardware L1 and L2 caches. The hierarchy changes where computed state can survive and how it can become useful again. HiCache system design

Capacity and speed pull in different directions

A GPU can consume KV efficiently when it is already in accessible device memory. Host DRAM generally provides more capacity per accelerator, but moving data back consumes a device interconnect. A storage backend can retain still more data and may allow reuse by other instances, at the cost of lookup, transfer, and coordination.

Consider a 128 MiB prefix. At an effective host-to-GPU bandwidth of 25 GB/s, moving that payload takes about 5.37 milliseconds. A storage read at 4 GB/s takes about 33.55 milliseconds, before a separate host-to-GPU copy or other overhead. If recomputing the prefix takes 20 milliseconds, a host hit looks attractive while that serialized storage path does not.

These numbers are chosen examples. They expose the decision: a cache hit saves computation only if retrieving the relevant state is preferable to computing it again under the current conditions. A hit in a directory is not the same as a low-latency hit in GPU memory.

There can still be a system-wide reason to fetch a prefix when its isolated request latency barely improves. Avoiding a compute-heavy prefill might release GPU capacity for other work. Conversely, fetching might congest a link needed by several other requests. We should examine both the request’s critical path and the shared resources it occupies.

What is private, and what can be shared?

In the documented HiCache architecture, GPU and host caches belong to an inference instance. A second instance does not automatically read the first one’s host cache, even when both run on the same node. Cross-instance reuse depends on an appropriately configured storage tier. A local file backend is not automatically a cluster-wide cache. HiCache tier-sharing scope

This matters for our document service. Suppose a request warms handbook A on worker 1. Routing the next question to worker 2 does not make worker 1’s GPU state magically local to worker 2. There must be an accessible copy and a supported transfer path. If a shared backend has not received the data yet, its apparent promise of global reuse is irrelevant to that request.

“Shared storage” also describes an arrangement, not just a product name. Workers need compatible namespaces, cache identity, model configuration, and access to the same stored objects. Two processes pointing at different local directories have two private stores even if both enable the same file-backend option.

A tree can describe several residencies

The hierarchical radix metadata tracks token spans and their local residency. A prefix may exist in device memory, host memory, or both. Storage lookup extends the search beyond those local tiers. The design does not require synchronizing every storage object’s exact address into every instance’s radix tree. HiCache design at v0.5.20

Imagine a request whose first 4,000 tokens are on the GPU, the next 4,000 have a host copy, and the remaining 200 are new. The system needs a coherent prefix covering the old 8,000 before it can extend that history with the suffix. Restoring a random later segment is not enough if earlier required state is missing.

While a transfer is in flight, its destination is not yet a usable cache hit. Memory must remain protected until the copy is complete, and execution must wait for whatever dependencies it actually needs. Moving bytes asynchronously permits overlap; it does not remove those ordering requirements.

When to write a backup

There are several reasonable times to move computed state to a slower tier. A write-through policy backs it up promptly. A selective policy waits for evidence of reuse. A write-back policy defers the move until faster-tier eviction needs it. SGLang documents these policy families and associated tuning controls. HiCache best practices

For the handbook assistant, write-through spends bandwidth early so another reader is more likely to find the document later. That is valuable for a popular shared handbook. It is wasteful for a one-off upload that nobody will revisit. A selective policy can favor frequently accessed prefixes, but the first repeat may arrive before a backup has been made.

Write-back delays the expense and can avoid unnecessary copies. Its downside is that eviction now has a transfer obligation. Under a burst, the system may need to reclaim GPU space exactly when the copy path is busiest. An asynchronous implementation must coordinate both the transfer and the eventual reuse of the source allocation.

No single policy wins independently of the workload. Measure the fraction of written bytes subsequently read, the latency of restores, and the effect on foreground transfers. A large cache that constantly writes cold data can consume bandwidth while contributing little useful reuse.

Fetch, prefetch, or recompute

A demand fetch begins when a request needs the data. Prefetching tries to begin earlier, while the request is waiting or while other work proceeds. If the data arrives before the GPU needs it, much of the fetch is hidden. If it arrives too late, waiting for it can delay useful execution.

SGLang exposes prefetch policies and timing controls rather than treating every storage hit as something to wait for indefinitely. The design and best-practices documents explain that this is a tradeoff between hit completion and responsiveness. A time budget prevents an optional reuse opportunity from becoming an unbounded stall. HiCache design

Our widget makes the tradeoff visible with a deliberately simpler policy. It knows a fixed recompute time and chooses the cheaper of fetching and recomputing. Real scheduling has to estimate costs, respect budgets, and account for overlap; the widget is not a reproduction of HiCache’s admission logic.

Where should this prefix come from?

Each prefix occupies 128 MiB. Host→GPU bandwidth is 25 GB/s. Storage loads stage through host memory; transfer times add. This model uses exclusive GPU/host residency, LRU replacement, instantaneous storage backup, and a known recompute time. Real HiCache can keep multiple copies and has asynchronous policies.

Request A, B, C, then A again. Reduce capacity to move a prefix further down the hierarchy.

GPU · private

empty

Host · private

empty

Storage · shared if configured

empty

GPU and host list most recently used first. Changing a capacity resets the sequence.

Illustrative model, not a SGLang benchmark. Controls update immediately; no automatic animation.

With two GPU slots, request A, B, then C. A leaves the GPU. Request A again and observe the host hit. Now use one GPU slot and zero host slots, reset, and request A, B, A. The final A must come from storage or be recomputed. Raise storage bandwidth until fetching becomes the cheaper option.

The model keeps one local copy per prefix, with GPU and host exclusive, and assumes instantaneous storage backup. Real HiCache can retain multiple copies and must pay for writes. Those simplifications isolate residency and fetch cost without pretending to simulate asynchronous transfer scheduling.

A capacity benefit can become a bandwidth problem

Suppose ten workers simultaneously restore a 1 GiB prefix through a shared path that sustains 25 GB/s in aggregate. The ten copies contain roughly 10.74 GB. Even before overhead, the shared path needs about 430 milliseconds of aggregate service time. Calculating 43 milliseconds for each request in isolation and assuming all ten finish together would create bandwidth out of nothing.

A hierarchical cache therefore moves the resource question. You can admit workloads whose reusable state exceeds GPU capacity, but you must provision and schedule the paths that restore that state. Warm local hits, host hits, and shared-storage hits deserve separate measurements.

The same distinction becomes crucial in disaggregated inference. HiCache moves state to preserve it for possible reuse. A prefill-to-decode transfer moves state because the next stage of a live request runs elsewhere. They can use related infrastructure, but one is an optional optimization and the other is a dependency of the chosen execution path.

Sources and further reading

  • HiCache design: architecture and residency semantics; online documentation checked September 24, 2026.
  • Pinned design document: release-specific reference.
  • HiCache best practices: policy, layout, and deployment considerations.
  • Mooncake: a related KV-centric serving architecture; its complete scheduler should not be equated with a HiCache storage backend.