According to Nvidia's article, DeepInfra's benchmark results show that NVIDIA Vera CPUs are more than twice as fast as other CPUs and can support more concurrent AI agents. Cloud platform DeepInfra is an early experiential participant in Nvidia's open AI ecosystem, independently designed and ran benchmarks based on its production AI agent infrastructure. DeepInfra processes nearly 5 trillion tokens per week, of which around 30% are driven by proxy systems. Its cloud platform is built for high-throughput AI inference. Benchmarks showed that under the same service quality, up to 1.6 times more concurrent AI agents are supported, and the collaboration speed is as high as 2.2 times, while improving infrastructure utilization and cost efficiency. These results show that NVIDIA Vera CPUs can achieve the cost efficiency, low latency, and throughput required for production agent artificial intelligence. As AI agents take on more complex reasoning, planning, tool usage, and data flow, CPU performance becomes increasingly important in the coordination work around every model call. As part of Nvidia's extreme co-design strategy for AI factories, Vera CPUs are designed specifically for proxy workloads. DeepInfra's benchmark tests highlight how Nvidia Vera CPUs help cloud service providers improve infrastructure utilization, improve cost efficiency, and support more concurrent AI agents with the same service quality.

Zhitongcaijing · 2d ago
According to Nvidia's article, DeepInfra's benchmark results show that NVIDIA Vera CPUs are more than twice as fast as other CPUs and can support more concurrent AI agents. Cloud platform DeepInfra is an early experiential participant in Nvidia's open AI ecosystem, independently designed and ran benchmarks based on its production AI agent infrastructure. DeepInfra processes nearly 5 trillion tokens per week, of which around 30% are driven by proxy systems. Its cloud platform is built for high-throughput AI inference. Benchmarks showed that under the same service quality, up to 1.6 times more concurrent AI agents are supported, and the collaboration speed is as high as 2.2 times, while improving infrastructure utilization and cost efficiency. These results show that NVIDIA Vera CPUs can achieve the cost efficiency, low latency, and throughput required for production agent artificial intelligence. As AI agents take on more complex reasoning, planning, tool usage, and data flow, CPU performance becomes increasingly important in the coordination work around every model call. As part of Nvidia's extreme co-design strategy for AI factories, Vera CPUs are designed specifically for proxy workloads. DeepInfra's benchmark tests highlight how Nvidia Vera CPUs help cloud service providers improve infrastructure utilization, improve cost efficiency, and support more concurrent AI agents with the same service quality.