PCIe 5.0 NVMe SSDs can increase a server's local storage throughput, but they are not a one-size-fits-all performance upgrade. The actual benefit is realized only when there is a proven I/O bottleneck and a consistently compatible Gen5 platform. For OLTP, virtualization, or AI, link rate, queue behavior, latency, SSD endurance, and PCIe topology are therefore all important factors. Manufacturer specifications provide valuable comparative data when the test profile and operating conditions are transparently documented.
PCIe 5.0, NVMe, and SSD: A Breakdown
To ensure a reliable assessment, three levels must be distinguished: PCIe The transport bus is the link between the host and the device; NVMe is the underlying protocol for non-volatile storage; and the SSD is the actual drive, consisting of a controller, NAND flash, and firmware. A fast PCIe connection therefore initially describes only the potential connection; it defines neither the internal processing of the SSD nor the performance of an application.
PCIe 5.0 operates at 32.0 GT/s per lane, while PCIe 4.0 reaches 16.0 GT/s per lane. This doubles the signaling rate per lane and, consequently, the theoretically available link bandwidth compared to Gen4. However, this statement refers to the transmission path: it is not a guarantee that database queries, virtual machines, or AI jobs will also run twice as fast.
GT/s
These values describe the signaling rate per lane. They are not a data rate or an application benchmark.
Data Table for the Chart
| Entry | GT/s |
|---|---|
| PCIe 4.0 | 16 |
| PCIe 5.0 | 32 |
In a typical x4 connection The order of magnitude can be derived from the PCI-SIG specifications. According to PCI-SIG, a PCIe 5.0 x32 link provides 128 GB/s of effective raw bandwidth in each direction. Four lanes are one-eighth of 32 lanes, so PCIe 5.0 x4 theoretically provides about 16 GB/s in each direction. This is a link specification and not an achievable SSD throughput; other protocol overheads, as well as the controller, NAND, and firmware, reduce or limit the usable performance.
In addition, the SSD controller, the number and configuration of the NAND chips, firmware algorithms, and the specific read and write patterns all play a role. Low queue depths or small, mixed accesses can slow down an application, even though the PCIe link theoretically still has capacity to spare. Conversely, a parallel sequential data stream can make good use of a faster host connection, though this does not constitute a general rule of performance.
For the procurement of Server Storage So the key question isn’t simply „Gen4 or Gen5?“ Rather, it is necessary to determine whether the existing PCIe path, the SSD itself, and the actual I/O workload all constitute the same bottleneck. Only when these factors align can the additional bandwidth of PCIe 5.0 become practically relevant for the application.
Why Enterprise SSDs Are Rated Differently
A Enterprise SSD is not selected based solely on sequential MB/s. In a data center, in addition to performance, the most important factors are predictable continuous operation, operational monitoring, and integration into server platforms. Consumer models may be technically suitable for individual tasks, but they do not automatically meet the requirements for availability, maintenance, protection mechanisms, and qualified firmware paths in a production environment.
There are significant differences in form factor: Enterprise drives are available in formats such as U.2, E1.S, or E3.S, making them suitable for hot-swap bays, dense server layouts, and dedicated cooling systems. The Solidigm D7 PS1010 series, for example, is available in U.2, E1.S, and E3.S form factors. The form factor therefore plays a role in determining the backplane, airflow, capacity density, and future replaceability.
The Endurance Class must match the write profile. High transfer rates are no substitute for adequate write endurance: Logging, streaming, and database writes place different demands on the drive than analytics workloads that are predominantly read-intensive. Solidigm positions the D7-PS1010 as a standard-endurance model and the D7-PS1030 as a mid-endurance model. Capacity, performance, and permissible write load are therefore selection criteria that must be evaluated separately.
Also important are Power loss protection and telemetry. Power failure protection can help manage write operations that have already been initiated in a controlled manner; telemetry data supports the monitoring of status, temperature, and wear. Depending on the environment, procurement requirements may also include encryption features, secure erasure procedures, and traceable firmware support. However, their specific availability must always be verified for each model and configuration.
Another enterprise-class feature may be dual-port operation. Samsung’s PM1743 is positioned as a PCIe 5.0 enterprise SSD with this feature, which may be relevant for certain high-availability architectures. This does not mean that every Gen 5 SSD is dual-port-capable or that this feature is required for every server. Rather, the example illustrates why data sheets and platform requirements must be read in conjunction with one another.
Even product names are no substitute for technical testing. The Micron 9550 series documents, among other things, high performance metrics and specified test profiles; however, this does not imply general suitability for every rack. A sound selection process combines the characteristics of the specific drive with endurance goals, service concepts, form factor, security requirements, and the actual server architecture in place.
The platform determines the possible link rate
A PCIe 5.0 SSD can only achieve its maximum link rate if the entire path supports Gen5. This includes the CPU platform and root port, the motherboard, riser or cable, the backplane, the PCIe switch (if applicable), the drive bay, and the firmware of the components involved. A single Gen5-capable component is not enough, because the connection between the involved components operates at a speed and bandwidth supported by all of them collectively.
PCIe 5.0 is backward compatible with earlier PCIe generations. If a Gen5 drive is connected to a path that is Gen4-capable only, it can generally still function, but only at the link rate supported by that configuration. The SSD will not have a Gen5 connection in this setup; however, it may still be reusable if you switch to a different platform later, provided the form factor, firmware, and server certifications are also compatible.
The following requires special attention: PCIe Topology. Multiple drives can be connected to the processor via a switch or a shared uplink connection. In that case, the SSD’s individual link is not necessarily the limiting factor: under parallel load, the shared uplink may reach capacity first. The same applies when additional accelerators, network cards, or other expansion cards are powered from the same lane budget.
The available Lane Budget It is therefore a planning factor, not just a specification in the datasheet. Four x4 SSDs require a total of 16 lanes on their device links, but their combined bandwidth depends on how these links are connected to root ports and switch uplinks. For high overall throughput, an architecture with multiple parallel drives may be more practical than replacing a single device.
Before making a purchase, you should check the server manual and compatibility lists against your specific configuration. Key factors include the supported generation and width per slot, limitations when fully populated, backplane and riser variants, and available firmware versions. Only this check can distinguish a theoretically fast enterprise SSD from a configuration that can actually deliver its Gen5 connectivity in a production server.
How to Read and Compare Manufacturer Specifications Correctly
Specifications describe clearly defined test scenarios, not necessarily the performance of a server workload. For the Micron 9550, the manufacturer specifies up to 14,000 MB/s sequential read speed, 10,000 MB/s write speed, and up to 3.3 million 4-KiB read IOPS. These metrics answer different questions: MB/s measures large, continuous transfers; IOPS describes the number of small input and output operations per unit of time.
The measurement parameters are crucial. The sequential values listed are based on 128-KiB transfers at a queue depth of 32, while the random 4-KiB values are based on a queue depth of 512 and a defined steady state. A database with few concurrent queries can therefore achieve significantly lower IOPS despite using a fast SSD. The read/write mix, capacity variant, and fill level also affect the results.
| Feature | PCIe 4.0 | PCIe 5.0 | Practical Classification |
|---|---|---|---|
| Signaling Rate per Lane | 16.0 GT/s | 32.0 GT/s | Gen5 doubles the signaling rate per lane. |
| Typical NVMe SSD Connection | x4 | x4 | The number of lanes for a single NVMe drive often remains the same. |
| Effective raw bandwidth at x4 | about 8 GB/s in each direction | about 16 GB/s in each direction | Mathematically, one-eighth of the values specified by PCI-SIG for x32; no SSD effective data rate. |
| Host Requirement | End-to-End Gen4 Path | End-to-End Gen5 Path | The connection operates only at the rate supported by the overall configuration. |
| Statement for Applications | No direct conclusion | No direct conclusion | Protocol shares, SSD configuration, topology, and workload limit the practical performance. |
PCI-SIG refers to the bandwidth specifications as the effective raw bandwidth of the link. To determine the usable SSD throughput, additional protocol overhead and implementation limitations must be taken into account. A comparison of link generations therefore shows the available transport headroom, but it does not replace either the data sheet for the specific drive or a measurement using the intended workload.
To ensure reliable product comparisons, manufacturers should specify the transfer size, queue depth, access pattern, test method, and operating condition. The SNIA Enterprise PTS It defines methods to make performance measurements of enterprise SSDs more transparent. While it does not replace your own load analysis, it helps ensure that peak values from different test profiles are not mistakenly treated as directly comparable.
Which Workloads Truly Benefit from Gen5
For AI data pipelines and HPC, Gen5 can be useful when large local datasets slow down the computing process. GPUDirect Storage enables direct DMA transfers between storage and GPU memory without a bounce buffer in CPU memory. NVIDIA describes fully GPU-offloaded compute pipelines as a favorable use case: The GPU handles the first and last accesses to the data moved between storage and the GPU. This describes a usage profile, not a universal, hard-and-fast functional requirement. Whether GDS helps depends on the actual data path, the APIs used, and the storage bottleneck.
In OLTP, the following are usually considered: tail latencies and write behavior under load more significantly than sequential peak rates. A high cache hit rate prevents many read accesses from even reaching the SSD; CPU-limited queries remain CPU-limited even after a drive replacement. Therefore, P95 and P99 latency, commit latency, I/O wait, and queue depth should be measured rather than just throughput.
| Condition | Implications for Potential Benefits | Variable to be tested |
|---|---|---|
| Storage I/O is on the critical path | A more direct data path may become important if storage bottlenecks the GPU pipeline. | Measure data load time, GPU utilization, and storage throughput together. |
| The GPU handles the first and last data accesses | Computational pipelines that are fully offloaded to the GPU are a suitable use case for GDS, not a universal functional requirement. | Check the first access after reading and the last access before writing, as well as the use of `cuFile` in the actual data path. |
| Large or coarse-grained streaming transfers | The documented GDS alignment is better suited for coarse-grained transfers than for fine-grained random accesses. | Record the application's data transfer sizes and access patterns. |
| Multiple Drives and Parallel Transfers | To fully utilize wide PCIe lanes, multiple x4 devices and simultaneous transfers may be required. | Check the number of devices, parallel processing, switch uplinks, and root port topology. |
Virtualization can benefit from a faster SSD when many VMs generate random I/O simultaneously and storage latencies limit consolidation. This assumption is known as Planning Heuristic ...should not be treated as a general Gen5 result. CPU readiness time, RAM pressure, shared PCIe uplinks, and a distributed storage system can limit the same workload more than the individual drive.
For analytics, ETL, backup, and archiving as well, the purchasing decision is based on the measured bottleneck. Large parallel local transfers can utilize higher link bandwidth, while the network, compression, or target system may limit the effectiveness of backup tasks. Therefore, compare the total job runtime and end-to-end throughput rather than drawing direct conclusions about application performance based solely on the sequential SSD peak value.
Write-intensive pipelines also require the appropriate endurance class. Solidigm positions the D7-PS1010 as a standard-endurance model and the D7-PS1030 as a mid-endurance model. This does not imply a universal product selection, but rather a rule for planning: performance, capacity, and permissible write load are separate criteria and must align with the actual write profile.
Server Storage Assessment Plan Prior to Procurement
Don't start the procurement process with a synthetic peak benchmark; instead, use baseline data from production-like operations. To do this, define a few representative load cases: such as database transactions, VM startup spikes, or an ETL run. Separate cold starts from Steady state; filled caches and prolonged write loads can lead to fundamentally different results. The article explains the same methodological principle How to Accurately Measure Performance After WordPress Updates using the example of cache-dependent hosting workloads.
Report the following metrics for each load case: P95/P99 latency, commit latency, I/O wait, queue utilization, and cache hit rate. A long queue with low SSD utilization may indicate an upstream bottleneck; conversely, a high cache hit rate explains why a faster SSD has little noticeable impact. Also include CPU, RAM, and network metrics to ensure that storage is not prematurely identified as the cause.
Before each performance evaluation, check which devices the system recognizes. On a Linux server with nvme-cli installed, use nvme list Inventory: According to the project documentation, the command scans the sysfs tree for NVMe devices and outputs their device nodes and associated information. Use this to first describe the detected inventory before comparing performance metrics.
A good starting point for PCIe diagnostics is the following examples from the older Red Hat guide for OpenStack Platform 10: LnkCap is set there for the maximum link speed and LnkSta Evaluated for the current link status. This is not a current Gen5 platform release. Replace the example address with your device's PCI address, and verify the syntax, full output, and meaning against the locally installed pciutils documentation and the server manual.
With regard to health data, the cited nvme-cli-master documentation states smart-log as an outdated alias for log smart; the alias is intended for older scripts, while new scripts should use the new format. However, "master" describes the development status and does not automatically reflect the command set of a stable distribution version. Therefore, first check the version of the package you have installed and its help or man page. Use only the command format documented there, and replace the example device path to match your system. No blanket version limit is derived from this.
A general destructive benchmark is unsuitable for production systems: It can overwrite data, affect caches, or produce an unrepresentative result due to an atypical queue depth. Instead, compare the same application use cases before and after a planned hardware upgrade. Document the firmware version, capacity, fill level, temperature, and topology so that any discrepancies can be traced later.
Planning Cooling, Endurance, and Sustained Performance
Peak values listed in data sheets describe defined test profiles, not necessarily the continuous operation of a loaded server. As the fill level increases, with mixed write operations, and during the Garbage Collection An SSD must reorganize data. NAND management and the specific write pattern can therefore affect throughput and, in particular, latency compared to short, empty test runs. Specifications become comparable only when the operating profile and steady state are taken into account.
Temperature is also a factor in performance planning. The D7-PS1010 family includes U.2, E1.S, and E3.S models; for the E1.S, Solidigm distinguishes between air-cooled versions and a variant for direct liquid cooling with a single-sided cold plate. Therefore, “E1.S” alone does not necessarily mean liquid cooling. Check the specific model and height you’ve ordered, as well as its compatibility with the server, slot, and, if applicable, cold plate. For air-cooled models, the platform evaluation should include approved power consumption, airflow, slot spacing, and inlet temperature.
The Endurance Class must be selected separately from the interface generation. For continuous log ingestion or write caches, a mid-endurance model may be more suitable than one designed for standard endurance, even if both read at similar speeds. Capacity not only affects cost and usable space but can also alter endurance and performance profiles depending on the model.
| Criterion | Issue to be clarified | Operational Consistency |
|---|---|---|
| Form Factor and Capacity | Will the U.2, E1.S, or E3.S fit in the bay, backplane, and riser; is there enough usable capacity? | Check for mechanical and electrical compatibility before placing an order. |
| Battery Life and Writing Profile | Does the class match the expected daily write volume and the read/write mix? | Design writing-intensive roles separately from reading-intensive roles. |
| Latency profile | Are the percentiles, load condition, and test profile documented? | Do not compare solely based on sequential peak values. |
| Protection and Telemetry | Are power-loss protection, safety features, and status data required? | Include maintenance, contingency planning, and compliance in the selection process. |
| Firmware and Platform | Are firmware support and server, backplane, and cooling certifications available? | Ensure compatibility before the rollout. |
Systematically Narrow Down Typical Gen5 Errors
Don't start troubleshooting by replacing the SSD; start with the negotiated link instead. A lower speed is, at first, a diagnostic finding and not yet a definitive cause. Compare the information explained in the measurement plan LnkSta and LnkCap Follow the instructions in the local tool documentation; the older OpenStack guide cited here demonstrates this basic check but does not replace the current server documentation. Then verify the slot assignment, riser, backplane, switch path, and firmware version against the platform specifications.
If the link is correct, then the following will appear: Queue Behavior of the application. Low IOPS with a shallow queue depth do not indicate an SSD failure; similarly, a full shared uplink can throttle multiple drives. Therefore, check I/O wait time, queue utilization, read/write mix, and latency percentiles simultaneously with CPU, memory, and network metrics. Only by correlating these metrics can you determine whether storage is on the critical path.
For GPU-close data paths, a theoretically direct PCIe path is not sufficient. GPUDirect Storage requires an appropriate hardware and software configuration; depending on the topology, ACS and IOMMU can route data through CPU root ports. Changes to these features affect isolation, security, and virtualization. Therefore, they must only be made after an architectural review and a documented risk assessment—not as a blanket performance toggle.
If performance remains unstable under load, check the temperature, health data, and application load. Select the SMART command based on the installed package version and its man page. The cited master documentation recommends nvme log smart and leads nvme smart-log as an outdated alias; it does not reflect the scope of every stable distribution. Compare cold start and steady state separately, and document the measurement conditions rather than concluding there is a defect based on a single outlier.
When to Buy Gen5 and When to Stick with Gen4
A PCIe 5.0 SSD is justified if benchmark data indicates a storage or host link bottleneck and the entire platform actually supports Gen5. Relevant indicators include consistently high queue utilization, insufficient throughput, or problematic latency under representative workloads. PCIe 5.0 doubles the supported bit rate per lane compared to PCIe 4.0; however, this does not automatically result in faster application performance.
It is appropriate to continue using Gen4 if existing drives are not operating at full capacity, the server platform only supports Gen4, or the CPU, memory, and network are limiting the service. Even with low concurrency or a high cache hit rate, the performance gap in the application may remain small. The budget saved can then be used more effectively, if necessary, for capacity, redundancy, memory, or eliminating the identified bottleneck.
Fewer Gen5 drives are not necessarily better than multiple Gen4 devices. For high overall transfer rates, what matters is the number of parallel paths, the Lane Budget and the uplinks from switches or root ports. Regarding GPU-proximal local storage paths, NVIDIA states that multiple x4 NVMe devices and parallel transfers may be required to fully utilize an x16 PCIe link. This is a statement regarding GDS topology and not a general rule for the number of devices in a server.
NVMe over Fabrics, or disaggregated storage, on the other hand, is an option when capacity needs to be flexibly shared among hosts. In this case, NICs, network switches, protocol paths, and the operating model also determine the achievable performance. Local NVMe SSDs prioritize host-near data paths; centralized systems can enable shared use and independent scaling. Both approaches must be weighed against requirements for latency, availability, and operation.
Before procurement, the workload and measurement window should be determined, the link rate and topology should be verified, and capacity, endurance, cooling, firmware support, and protection features should be thoroughly evaluated. Additionally, power supply, rack environment, and scaling options should be considered in the infrastructure decision. The article explains specific criteria for this. Criteria for Choosing the Right Data Center. The purchase decision is thus based on the proven bottleneck rather than on a single peak value.
Sources and Current State of Knowledge
Status of the research:
As of September 26, 2026. Information regarding interfaces, products, and nvme-cli syntax is based on the manufacturer and project documentation cited in this article; check compatibility and the firmware version before use.
https://pcisig.com/faq?field_category_value%5B%5D=pci_express_5.0
https://www.snia.org/solid-state-sss
https://nvmexpress.org/specification/nvme-over-pcie-transport-specification/
https://www.solidigm.com/products/data-center/d7/ps1010.html
https://www.solidigm.com/content/dam/solidigm/en/site/products/data-center/product-briefs/ps1010-ps1030/solidigm-d7-ps1010-d7-ps1030-product-brief.pdf
https://semiconductor.samsung.com/ssd/enterprise-ssd/pm1743/
https://www.micron.com/content/dam/micron/global/public/products/data-sheet/ssd/9550-nvme-ssd-tech-prod-spec.pdf
https://docs.nvidia.com/gpudirect-storage/overview-guide/
https://docs.nvidia.com/gpudirect-storage/pdf/design-guide.pdf
https://github.com/linux-nvme/nvme-cli/blob/master/Documentation/nvme-list.txt
https://docs.redhat.com/zh_hans/documentation/red_hat_openstack_platform/10/pdf/ovs-dpdk_end_to_end_troubleshooting_guide/red_hat_openstack_platform-10-ovs-dpdk_end_to_end_troubleshooting_guide-en-us.pdf
https://github.com/linux-nvme/nvme-cli/blob/master/Documentation/nvme-smart-log.txt
https://www.solidigm.com/archive-v2/products/data-center/d7/ps1010.html
https://docs.nvidia.com/gpudirect-storage/best-practices-guide/




