Nvidia published benchmark results for its Vera CPU on August 3, claiming up to 3.67 times faster throughput on CRC32C integrity checks than standard x86 processors. The chip, built into the BlueField-4 STX storage processor, targets AI-native storage operations like encryption, compression, and data recovery.

The performance gains vary sharply by workload. Compression speeds hit 3.29x faster, Reed-Solomon recovery clocked 3.26x, and multi-stage pipelines reached 3.21x. AES-128 encryption showed a smaller 1.43x improvement, where x86 chips already have dedicated hardware accelerators. The pattern is clear: Vera dominates tasks requiring raw parallel throughput, while specialized instruction sets narrow the gap.

What's inside Vera

The chip packs 88 custom Nvidia Olympus cores built on Armv9.2, delivering 176 threads total. Memory bandwidth reaches 1.2 TB/s through SOCAMM2 LPDDR5X, with 3.4 TB/s core-to-core bisection bandwidth via Nvidia's Scalable Coherency Fabric. Beyond storage acceleration, Vera doubles as the host CPU for Nvidia's Rubin GPU platforms, embedding it into what the company calls its unified AI factory stack. Commercial availability is scheduled for fall 2026.

The benchmarks measure real-world bottlenecks in modern AI storage: how fast you can verify data integrity, compress datasets, recover from failures, and encrypt at scale. But vendor benchmarks always come with caveats. Nvidia hasn't disclosed which specific x86 CPU served as the comparison point, leaving room for cherry-picked baselines. Real-world performance under production loads often diverges from controlled launch conditions.

This article is informational only and not investment or financial advice.