Historical past is affected by the corpses of applied sciences that have been forward of their time, and Intel’s Optane storage and reminiscence merchandise are definitely amongst them.
Costly, badly misunderstood, and awkwardly priced, the expertise struggled to discover a place available in the market, surviving simply 5 years earlier than Chipzilla pulled the plug.
Issues may need been completely different if it have been launched in the present day. The shift from AI coaching to inference has pushed super demand for DRAM and NAND flash. The truth is, the identical properties that made Optane a distinct segment product just a few years in the past would have made it ideally suited to those write intensive AI workloads.
The rise and fall of Optane
Intel and Micron co-developed Optane – or extra particularly, 3D XPoint reminiscence – and unveiled it in 2015. It promised to bridge the hole between standard NAND flash utilized in SSDs and DRAM utilized in DDR4 and (later) DDR5.
Like NAND, 3D XPoint was non-volatile, which suggests information persists when unpowered, making it acceptable for storage functions.
However not like NAND, 3D XPoint didn’t retailer information utilizing trapped electrons and as a substitute relied on a section change materials that was each extraordinarily quick, reaching sub-10-microsecond latencies (even decrease for later PMem modules), and absurdly write-endurant.
Intel’s penultimate Optane SSDs boasted endurance of 100 drive writes a day and a imply time between failures of two million hours — specs that stay unmatched in the present day.
These figures are in fact extrapolated, primarily based on what we find out about how 3D XPoint reads and writes information. But when something, we suspect Intel was in all probability being conservative with its claims.
This one-two punch of endurance and latency, significantly for random reads and writes meant that it might be used as a second tier of system reminiscence. Optane SSDs are actually, actually low latency for storage-class reminiscence, however they’re nonetheless orders of magnitude much less responsive than DRAM which might hit 100 nanoseconds or decrease. We should always observe Intel’s PMem DIMMs have been able to latencies of roughly 350 nanoseconds.
Intel’s Optane persistent reminiscence additionally had the advantage of capacities as much as 512 GB per DIMM, at a time when probably the most you can anticipate out of DDR4 was about 128 GB – and provided that you had deep pockets.
These persistent reminiscence DIMMs might be made to behave like primary reminiscence, with normal DDR4 functioning as a large L4 cache when enabled; as a storage pool; or in an application-aware reminiscence mode, which uncovered the Optane reminiscence straight to pick functions.
Regardless of Optane’s many strengths, it was moderately awkwardly priced. For a similar reminiscence density, 3D XPoint was constantly costlier than NAND and, whereas cheaper than DRAM, considerably much less performant.
Because of this, there was hardly ever a state of affairs during which customers have been higher off with Optane than they’d be simply shopping for extra NAND or fewer, larger DIMMs.
To make issues worse, because the boffins at TechInsights noted on the time, whereas 3D XPoint had its deserves, it wasn’t enhancing shortly sufficient to maintain up with NAND flash on bit density.
By 2021, Micron had had sufficient. The reminiscence vendor introduced it might finish growth of 3D XPoint merchandise. With no supply of latest silicon for its SSD and chronic reminiscence merchandise, the writing was on the wall for Optane.
Intel’s then-CEO Pat Gelsinger formally pulled the plug in 2022, ending growth of latest merchandise beneath the Optane banner. The P5810X and P5811X, the ultimate Optane merchandise to roll off the meeting line, launched in late 2022.
Forward of its time
The identical yr Intel pulled the plug on Optane, a brand new workload was rising that might utterly flip the datacenter on its head, and may very nicely have been the killer app to justify 3D XPoint’s continued growth.
In November 2022, OpenAI, then a comparatively obscure AI startup, lifted the curtain on ChatGPT, kicking off an AI arms race and sparking a large inflow of cash into the tech sector.
Massive language fashions are among the many most resource-intensive workloads on the planet. However till 2025, most of that compute was devoted to coaching larger, smarter, and fewer hallucination-prone fashions.
The arrival of DeepSeek early in that yr signaled the shift towards inference, a workload that wants loads of compute, reminiscence, and storage.
Trendy inference engines are tuned to maximise effectivity. A technique is by caching the important thing worth pairs used to trace mannequin state.
Chatbots and brokers are inherently iterative. For multi-turn classes, each immediate comprises not solely new data, but additionally each request and response that got here earlier than it. Recomputing all of that each time a brand new request is made is wasteful, so as a substitute that data is computed, cached, and reused.
The issue is at giant context sizes and excessive concurrency, these KV caches can get moderately giant. A single 64,000-token sequence in a mannequin like DeepSeek R1 can chew up 4 gigabytes of GPU reminiscence. Multiply that throughout a whole lot or hundreds of customers and it provides up shortly.
Trendy GPUs don’t have a lot reminiscence, and what they do have is often tied up internet hosting mannequin weights. Many inference engines due to this fact assist KV cache offloading, both natively or by way of plugins. As soon as GPU reminiscence is exhausted or chat classes stale, older chats are ejected to system reminiscence.
However DRAM is dear, briefly provide, and will not even be sufficient. The truth is, Nvidia is reportedly reducing the quantity of LPDDR5X shipped as a part of its Vera Rubin platform.
Due to this, KV caches the place, for instance, a consumer has walked away and the session has been idle for an hour or extra, could also be offloaded to flash storage arrays for longer-term storage.
The issue is that KV caching is a write-intensive workload and NAND has a finite variety of writes it could possibly make earlier than it’s shot. There’s quite a lot of work being completed to mitigate this, however the truth stays that if you happen to write sufficient information to NVMe storage, finally it wears out.
Optane’s otherworldly write endurance and low latency, significantly when regarding small random writes which dominate KV-caching, would have made it an ideal selection for this workload.
Optane’s successors
One of many nails in Optane’s coffin was that Compute Express Link (CXL) was already on the horizon. If the first motive for adopting Optane was to get your arms on a big pool of moderately quick DRAM-like reminiscence, then CXL provided all the advantages and not one of the compromises. In impact, early CXL implementations have been basically a distant reminiscence controller to which you can connect your selection of reminiscence, DDR4 or DDR .
Want extra reminiscence than your CPU helps? Simply plug a CXL reminiscence expander right into a free PCIe slot and also you have been off to the races. Later revisions of the reminiscence coherent protocol added assist for pooling, and later sharing.
Nonetheless, CXL hasn’t modified the truth that DDR5 is dear at one of the best of occasions, and we’re at the moment dwelling within the worst of occasions for DRAM costs. (With that stated, whereas your CPU might solely assist DDR5, the CXL controller may nonetheless have the ability to use DDR4. That is precisely what Meta is doing to boost the memory capacity of its techniques on a budget.)
Within the absence of Optane, reminiscence distributors have tried to fill the void with write-optimized NAND of their very own.
Kioxia’s XL-Flash is one such instance of storage class reminiscence (SCM) that guarantees most of the similar advantages as Optane, together with marketed write latencies beneath 10 microseconds, and endurance as much as 60 drive writes a day, by way of the usage of SLC (1-bit per cell) reminiscence.
There’s additionally a crossover with CXL right here. XL-Flash might be uncovered as a CXL system to simplify reminiscence tiering.
Samsung additionally tried to copy a lot of Optane’s qualities with its Z-NAND tech, however the tech nonetheless falls nicely wanting 3D XPoint. Samsung is reportedly revamping the tech and aiming for a 15x efficiency enhance over standard NAND.
Alas, neither fairly lives as much as Optane’s legacy. The tech was just too far forward of its time. Had Intel and Micron waited just some years longer, it may need ridden the AI wave all the way in which to victory. ®
Source link

