Introducing ContextCube

Reuse Context
Across Your GPU Cluster

One shared, persistent KV cache pool.
Less repeated
computation. Help your inference GPUs generate more tokens.

Built around the shared Tier 3.5 architecture.
Powered by NVIDIA BlueField DPUs.

120GB/s
KV bandwidth per appliance¹
Shared
one pool across inference nodes
Simple
storage software on DPUs. Simpler hardware architecture.
Open
choice of inference engines and KV cache software
THE CHALLENGE

Growing context.
Isolated caches.

As local cache fills up and inference clusters grow, KV cache
needs more capacity and a shared home across nodes.
Without it, GPUs must process the same context again.

Memory fills up

Longer prompts and more concurrent requests exceed local memory and storage capacity. Valuable KV cache gets evicted.

Context stays isolated

Cache on one server may be unavailable to a request on another. The same context gets stored—or computed, again.

Repeated work adds up

Rebuilding context consumes GPU cycles and delays the first token, reducing token throughput and the number of concurrent requests the cluster can serve.

THE SHARED TIER 3.5

One pool.
Connected to every node.

ContextCube adds a purpose-built KV cache appliance alongside your existing memory and storage tiers.
Participating inference nodes can retrieve matching cached prefixes over RDMA, reducing repeated prefill work.

TuringData ContextCube shared Tier 3.5 architectureTuringData ContextCube shared Tier 3.5 architecture
FEATURES & BENEFITS

Make your inference
infrastructure go further.

Share more context. Reduce repeated computation. Match shared KV cache capacity to the scale of your models and inference cluster.

01 / SHARING

Reuse across GPU nodes

A KV pool makes cached prefixes accessible to participating nodes, reducing the need to maintain separate copies on each server.

Less redundant prefill. More capacity for requests.
02 / LARGE CAPACITY

More room for reusable context

Go beyond a single node's capacity with hundreds of terabytes of shared KV cache. Keep reusable context resident for hours to weeks, depending on configuration, workload, and cache policy.

Cache more KV data to support higher cache hit rates.
03 / DATA MOVEMENT

Move KV with DPUs

Storage software runs on NVIDIA BlueField-3 DPUs, simplifying the appliance architecture. Inference nodes read and write KV over RDMA, with 120 GB/s of bandwidth per appliance.

Simpler hardware. Fast KV access with less host CPU involvement.
04 / OPEN ECOSYSTEM

Compatibility across your stack

Works with inference engines such as vLLM and SGLang, and connects to KV cache software including TuringData Cache Fabric, LMCache, and Mooncake.

More freedom to choose the software combination that fits your deployment.
BUILT FOR CONTEXT REUSE

When your AI returns
to familiar ground

The strongest opportunities are workloads that repeatedly use matching context, and can retrieve it faster than recomputing it.

  • Reuse matching instructions, code prefixes, and prior context across repeated requests.
  • Retain reusable conversation prefixes as sessions grow or move between serving nodes.
  • Reuse matching document and prompt prefixes when requests revisit the same material.
  • For large inference clusters serving many concurrent requests, share and reuse KV cache to reduce redundant GPU computation and produce more tokens from the same compute resources.
Inside ContextCube

Purpose-built,
ready to share.

A dedicated appliance for your cluster's
shared KV cache tier.

Discuss your configuration →
COMPONENTPER APPLIANCE
KV bandwidth120 GB/s ¹
NVMe configuration26 × U.2 NVMe drives
Data processing4 × NVIDIA BlueField-3 DPUs
Connectivity4 x dual-port 200Gb/s per appliance
Starting deploymentOne appliance

¹ Reference performance. Actual performance depends on configuration and workload.

Discuss your configuration →

Put your cached
context to work

Explore how ContextCube fits your models, context lengths, and concurrency targets.
Reach out to our team of experts for personalized support, inquiries, and solutions tailored to your specific requirements.