As AI transitions from simple chatbots to complex, multi-turn agents, the demand for massive context windows has pushed traditional DRAM and on-chip memory to their physical limits. Huawei’s OceanStor M900 addresses this by creating a global, multi-tier KV cache that pools resources across compute, network, and storage layers. Utilizing a high-speed UnifiedBus interconnect, the system allows clusters to scale up to 64 PB of capacity, effectively shifting the bottleneck away from memory size for large-scale deployments.
In section Releases
Huawei Debuts OceanStor M900 to Solve AI Memory Bottlenecks
At HUAWEI CONNECT 2026, Huawei unveiled the OceanStor M900 Context Memory Storage, a hardware solution designed to overcome memory limitations in hyperscale AI inference. By enabling PB-scale shared memory capacity, the system aims to support the next generation of 10-trillion-parameter models and autonomous AI agents.

The architecture integrates CPU, network, and NAND controllers to enable one-hop connections between NPUs and SSDs. This design slashes access latency to 60 microseconds—a 90% reduction compared to standard protocols—and pushes aggregate bandwidth to 40 TB/s. Beyond performance, the M900 utilizes KV-aware adaptive storage to intelligently manage data lifecycles. By extending SSD endurance 16-fold, the system lowers long-term operational costs, facilitating the transition toward more sustainable and efficient large-scale AI infrastructure.
Comments (0)
No comments yet. Be the first!