Breaking Storage Bottlenecks in AI Large-Scale Inference: Solidigm TLC+QLC Multi-Tier Storage Eases DRAM Pressure
As AI large-scale inference enters a new development era, artificial intelligence operations have evolved far beyond simplistic Token input and output mechanisms. A complete end-to-end AI workflow encompasses request processing, gateway policy orchestration, task and tool scheduling, intelligent retrieval, context assembly, KV cache population, decoding computation, tool loop iteration, response delivery, and full-link evaluation. Every stage involves intensive and high-density data read and write operations.
This industrial transformation indicates that the rollout of commercial AI applications no longer relies solely on computing power expansion. Instead, it faces unprecedented storage pressure driven by explosive data growth. Surging data volumes, diversified and complex read/write scenarios, and rising storage costs have become core bottlenecks restricting efficient AI deployment — precisely the pain points addressed by Solidigm’s multi-tier storage solutions.
Addressing AI Data Challenges: Multi-Tier Storage as an Optimal Solution
The 2026 Flash Memory World Summit (FMW2026) gathered upstream and downstream players across the flash storage industry to discuss emerging opportunities and core challenges. As a global leading provider of flash storage solutions, Solidigm released in-depth industry insights and tailored solutions for AI storage bottlenecks at the summit.
Currently, mainstream 8-card AI computing nodes are equipped with 1.5TB to 2TB of system memory by default. Such high-speed DRAM primarily stores mission-critical hot data, including agent running status, business context, tool output content, task queues, and retrieval indexes. With the continuous iteration of AI agents, enrichment of application scenarios, and growing model complexity, industry demand for larger memory capacity keeps rising with expanding market gaps.
Contrary to the booming memory demand, hardware costs have skyrocketed. Memory module prices have surged by over 400% since last year, accompanied by persistent supply shortages. This situation has severely hindered the deployment and expansion of AI businesses for enterprises and data centers worldwide, creating widespread cost and supply chain challenges across the industry.
To balance AI data storage requirements and hardware cost pressures, the industry has pioneered a multi-tier SSD storage architecture. By offloading DRAM workloads via hierarchical storage, the solution delivers balanced performance, capacity, and cost without compromising overall AI system efficiency. Beyond traditional HBM and system memory, a brand-new three-tier storage architecture tailored for AI scenarios has been established to match data of varying popularity and types:
G3 High-Performance Storage Tier: Built on high-performance SSDs, dedicated to storing model files, task checkpointing data, and high-popularity KV cache data, requiring ultra-high bandwidth and ultra-low latency;
G3.5 Dedicated Context Cache Tier: A customized storage layer exclusive to AI context caching, enabling fast reading and writing of contextual data for rapid invocation;
G4 High-Capacity Storage Tier: Optimized for massive cold data storage, accommodating low-temperature KV data, AI training data lakes, and enterprise document resources with a focus on superior storage density and rack space utilization.
The SSD-centric multi-tier storage system effectively alleviates massive data storage pressure on DRAM. It perfectly adapts to the data growth brought by complex AI agents, large-model inference, and large-scale AI training, breaking enterprise data flow bottlenecks and achieving optimal balance between operational efficiency and total cost.
Ni Jinfeng, Vice President of Sales for APAC at Solidigm, stated at the summit that Solidigm’s product portfolio covers full-scenario AI applications including training, inference, and edge computing. The lineup ranges from high-performance PCIe 5.0 TLC SSDs (D7-PS1010 and PS1030) to ultra-high-capacity 122TB QLC SSDs (D5-P5336), fully catering to the hierarchical requirements of AI multi-tier storage architectures. Moving forward, Solidigm will continue to deepen collaboration with local Chinese ecosystem partners to help industries eliminate AI data flow bottlenecks and seize opportunities in the large-model era.
Dual TLC+QLC Media Strategy Unlocks 65% DRAM Offloading Capability
The three-tier AI storage architecture features clear hierarchical positioning and differentiated performance requirements. The core G3 and G3.5 tiers demand high IO performance, high throughput, and ultra-low latency to guarantee smooth AI operations. In contrast, the G4 tier prioritizes storage density, single-device capacity, and cost efficiency with flexible tolerance for instantaneous write performance. Equipped with a comprehensive lineup of TLC and QLC products, Solidigm delivers customized storage solutions that precisely match the differentiated needs of each tier.
Solidigm D7-PS1010 TLC SSD: Solid Foundation for High-Frequency AI Workloads
Engineered to meet the rigorous demands of the G3 high-performance tier, the Solidigm D7-PS1010 TLC SSD delivers industry-leading performance for high-frequency AI workloads. Powered by premium TLC flash particles, it achieves a throughput of 14.5GB/s and 3.3M IOPS, with random read/write latency as low as 60μs / 7μs and a stable write endurance of 1DWPD.
Featuring high bandwidth, ultra-low latency, and exceptional reliability, the D7-PS1010 fully supports core G3-tier operations. It provides robust storage support for high-popularity KV cache access, rapid model reloading, and high-frequency data reading and writing, ensuring low-latency execution of large-scale inference and intelligent agent scheduling tasks.

Solidigm D5-P5336 QLC SSD: Maximizing Capacity and Optimizing AI Storage Costs
Built for massive cold data workloads in the G4 tier, the Solidigm D5-P5336 QLC SSD delivers industry-leading capacity density and space efficiency with a maximum single-drive capacity of 122TB. It significantly optimizes data center rack space utilization and reduces hardware deployment costs. The drive achieves sequential read/write latency of 8μs / 21μs, fully meeting the routine storage and access demands of AI training data lakes, low-temperature KV data, and massive enterprise documents.

To further enhance QLC SSD performance in commercial scenarios, Solidigm has open-sourced the CSAL software layer. It intelligently aggregates scattered small-block random writes into efficient large-block sequential writes, effectively boosting QLC SSD performance and extending service life in data center and cloud AI workloads.
For the specialized G3.5 context cache tier, Solidigm supports flexible combinations of multiple SSD models to build high-performance dedicated cache systems, enabling ultra-fast reading and invocation of AI context and retrieval data.
Real-world business tests verify that the TLC+QLC multi-tier storage architecture can offload 65% of data from AI computing node system memory. This core capability drastically cuts expensive DRAM procurement costs without sacrificing system performance or user experience, while providing sufficient data storage support for deploying more complex, diverse AI agents and iterating large-scale AI applications.
Technological Innovation and Ecosystem Collaboration: Scenario-Matched AI Storage Solutions
As a universal core hardware component, SSDs support diverse application scenarios, with hierarchical storage serving as a core application for AI workloads. Catering to the green and high-efficiency heat dissipation trends of modern data centers, Solidigm has launched the industry’s first single-sided cold-plate liquid-cooled variant of the D7-PS1010 E1.S SSD. Compatible with mainstream liquid cooling architectures, it fully leverages existing liquid cooling resources to achieve efficient heat dissipation for high-performance storage hardware, improving data center energy efficiency and lowering overall storage TCO (Total Cost of Ownership).
Solidigm’s full-range storage products and customized solutions have been widely validated in real-world Chinese industry scenarios with proven adaptability and stability. For instance, the RC-Vision2000 intelligent quality inspection solution developed by Ruike Lianchuang adopts the D5-P5336 SSD to deliver stable, high-capacity, and high-efficiency underlying storage for manufacturing AI visual inspection workflows, supporting the storage, retrieval, and analysis of massive industrial image data.
Ni Jinfeng commented that the industry’s storage evaluation logic has undergone fundamental changes. The focus has shifted from standalone hardware performance indicators to the overall effectiveness of complete solutions in real AI scenarios. To achieve precise matching between products, technologies, solutions, and industry demands, Solidigm launched its global AI Central Laboratory by the end of 2025, alongside a dedicated Shanghai AI Central Laboratory serving Chinese users and partners.
Equipped with the latest Solidigm storage products and high-end GPU devices, the laboratory simulates diverse real-world AI workloads to continuously test, optimize, and iterate storage solutions, verifying product stability and adaptability under complex operating conditions. Through sustained technological R&D, scenario verification, and ecosystem cooperation, Solidigm effectively addresses core data storage pain points in the AI era. Its hierarchical, high-efficiency, and cost-effective storage system helps enterprises maximize AI business value and accelerate industrial intelligent upgrading.










