NVIDIA Expands Its Deep Learning Inference Capabilities for Hyperscale Datacenters

Company Unveils NVIDIA TensorRT 4, TensorFlow Integration, Kaldi Speech Acceleration and Expanded ONNX Support; GPU Inference Now up to 190x Faster Than CPUs

About NVIDIA
NVIDIA’s (NASDAQ:NVDA) invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots and self-driving cars that can perceive and understand the world. More information at http://nvidianews.nvidia.com/.

  1. Total cost of ownership based on a workload mix representative of a major cloud service provider: 60 percent neural collaborative filtering (NCF), 20 percent neural machine translation (NMT), 15 percent automatic speech recognition (ASR), 5 percent computer vision (CV) and per socket (Tesla V100 GPU vs CPU) workload speedups of: 10x NCF, 20x NMT, 15x ASR, 40x CV. CPU node configuration is two-socket Intel Skylake 6130. GPU recommended node configuration is the eight Volta HGX-1.
  2. Performance gains observed over a range of important workloads. Examples include ResNet50 v1 inference performance at a 7 ms latency is 190x faster with TensorRT on a Tesla V100 GPU than using TensorFlow on a single-socket Intel Skylake 6140 at minimum latency (batch = 1).

Certain statements in this press release including, but not limited to, statements as to: the benefits, impact, performance, uses and abilities of NVIDIA TensorRT 4 and its integration into the TensorFlow framework; GPU acceleration for deep learning inference gaining traction and its impact and benefits; the benefits and impact of NVIDIA working with Kaldi to optimize GPUs, NVIDIA collaborating with partners, TensorRT being used around the world, GPU acceleration for Kubernetes, NVIDIA Tesla GPU-accelerated servers and MATLAB software’s integration of TensorRT; NVIDIA’s technologies expanding its potential inference market; and the support for intelligent applications and frameworks improving the quality of deep learning and reducing costs of hyperscale servers are forward-looking statements that are subject to risks and uncertainties that could cause results to be materially different than expectations. Important factors that could cause actual results to differ materially include: global economic conditions; our reliance on third parties to manufacture, assemble, package and test our products; the impact of technological development and competition; development of new products and technologies or enhancements to our existing product and technologies; market acceptance of our products or our partners’ products; design, manufacturing or software defects; changes in consumer preferences or demands; changes in industry standards and interfaces; unexpected loss of performance of our products or technologies when integrated into systems; as well as other factors detailed from time to time in the reports NVIDIA files with the Securities and Exchange Commission, or SEC, including its Form 10- K for the fiscal period ended January 28, 2018. Copies of reports filed with the SEC are posted on the company’s website and are available from NVIDIA without charge. These forward-looking statements are not guarantees of future performance and speak only as of the date hereof, and, except as required by law, NVIDIA disclaims any obligation to update these forward-looking statements to reflect future events or circumstances.

© 2018 NVIDIA Corporation. All rights reserved. NVIDIA, the NVIDIA logo, NVIDIA DGX, NVIDIA DRIVE, Jetson and Tesla are trademarks and/or registered trademarks of NVIDIA Corporation in the U.S. and other countries. Other company and product names may be trademarks of the respective companies with which they are associated. Features, pricing, availability and specifications are subject to change without notice.

About NVIDIA
NVIDIA’s (NASDAQ:NVDA) invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots and self-driving cars that can perceive and understand the world. More information at http://nvidianews.nvidia.com/.

  1. Total cost of ownership based on a workload mix representative of a major cloud service provider: 60 percent neural collaborative filtering (NCF), 20 percent neural machine translation (NMT), 15 percent automatic speech recognition (ASR), 5 percent computer vision (CV) and per socket (Tesla V100 GPU vs CPU) workload speedups of: 10x NCF, 20x NMT, 15x ASR, 40x CV. CPU node configuration is two-socket Intel Skylake 6130. GPU recommended node configuration is the eight Volta HGX-1.
  2. Performance gains observed over a range of important workloads. Examples include ResNet50 v1 inference performance at a 7 ms latency is 190x faster with TensorRT on a Tesla V100 GPU than using TensorFlow on a single-socket Intel Skylake 6140 at minimum latency (batch = 1).

Certain statements in this press release including, but not limited to, statements as to: the benefits, impact, performance, uses and abilities of NVIDIA TensorRT 4 and its integration into the TensorFlow framework; GPU acceleration for deep learning inference gaining traction and its impact and benefits; the benefits and impact of NVIDIA working with Kaldi to optimize GPUs, NVIDIA collaborating with partners, TensorRT being used around the world, GPU acceleration for Kubernetes, NVIDIA Tesla GPU-accelerated servers and MATLAB software’s integration of TensorRT; NVIDIA’s technologies expanding its potential inference market; and the support for intelligent applications and frameworks improving the quality of deep learning and reducing costs of hyperscale servers are forward-looking statements that are subject to risks and uncertainties that could cause results to be materially different than expectations. Important factors that could cause actual results to differ materially include: global economic conditions; our reliance on third parties to manufacture, assemble, package and test our products; the impact of technological development and competition; development of new products and technologies or enhancements to our existing product and technologies; market acceptance of our products or our partners’ products; design, manufacturing or software defects; changes in consumer preferences or demands; changes in industry standards and interfaces; unexpected loss of performance of our products or technologies when integrated into systems; as well as other factors detailed from time to time in the reports NVIDIA files with the Securities and Exchange Commission, or SEC, including its Form 10- K for the fiscal period ended January 28, 2018. Copies of reports filed with the SEC are posted on the company’s website and are available from NVIDIA without charge. These forward-looking statements are not guarantees of future performance and speak only as of the date hereof, and, except as required by law, NVIDIA disclaims any obligation to update these forward-looking statements to reflect future events or circumstances.

© 2018 NVIDIA Corporation. All rights reserved. NVIDIA, the NVIDIA logo, NVIDIA DGX, NVIDIA DRIVE, Jetson and Tesla are trademarks and/or registered trademarks of NVIDIA Corporation in the U.S. and other countries. Other company and product names may be trademarks of the respective companies with which they are associated. Features, pricing, availability and specifications are subject to change without notice.

Media Contacts

Global contacts for media inquiries.

All Contacts

Stay Informed

Newsroom updates delivered to your inbox.

Subscribe