
In the latest MLPerf Inference V5.0 benchmarks, which reflect some of the most challenging inference scenarios, the NVIDIA Blackwell platform set records - and marked NVIDIA's first MLPerf submission using the NVIDIA GB200 NVL72 system, a rack-scale solution designed for AI reasoning.
Delivering on the promise of cutting-edge AI takes a new kind of compute infrastructure, called AI factories. Unlike traditional data centers, AI factories do more than store and process data - they manufacture intelligence at scale by transforming raw data into real-time insights. The goal for AI factories is simple: deliver accurate answers to queries quickly, at the lowest cost and to as many users as possible.
The complexity of pulling this off is significant and takes place behind the scenes. As AI models grow to billions and trillions of parameters to deliver smarter replies, the compute required to generate each token increases. This requirement reduces the number of tokens that an AI factory can generate and increases cost per token. Keeping inference throughput high and cost per token low requires rapid innovation across every layer of the technology stack, spanning silicon, network systems and software.
The latest updates to MLPerf Inference, a peer-reviewed industry benchmark of inference performance, include the addition of Llama 3.1 405B, one of the largest and most challenging-to-run open-weight models. The new Llama 2 70B Interactive benchmark features much stricter latency requirements compared with the original Llama 2 70B benchmark, better reflecting the constraints of production deployments in delivering the best possible user experiences.
In addition to the Blackwell platform, the NVIDIA Hopper platform demonstrated exceptional performance across the board, with performance increasing significantly over the last year on Llama 2 70B thanks to full-stack optimizations.
NVIDIA Blackwell Sets New Records The GB200 NVL72 system - connecting 72 NVIDIA Blackwell GPUs to act as a single, massive GPU - delivered up to 30x higher throughput on the Llama 3.1 405B benchmark over the NVIDIA H200 NVL8 submission this round. This feat was achieved through more than triple the performance per GPU and a 9x larger NVIDIA NVLink interconnect domain.
While many companies run MLPerf benchmarks on their hardware to gauge performance, only NVIDIA and its partners submitted and published results on the Llama 3.1 405B benchmark.
Production inference deployments often have latency constraints on two key metrics. The first is time to first token (TTFT), or how long it takes for a user to begin seeing a response to a query given to a large language model. The second is time per output token (TPOT), or how quickly tokens are delivered to the user.
The new Llama 2 70B Interactive benchmark has a 5x shorter TPOT and 4.4x lower TTFT - modeling a more responsive user experience. On this test, NVIDIA's submission using an NVIDIA DGX B200 system with eight Blackwell GPUs tripled performance over using eight NVIDIA H200 GPUs, setting a high bar for this more challenging version of the Llama 2 70B benchmark.
Combining the Blackwell architecture and its optimized software stack delivers new levels of inference performance, paving the way for AI factories to deliver higher intelligence, increased throughput and faster token rates.
NVIDIA Hopper AI Factory Value Continues Increasing The NVIDIA Hopper architecture, introduced in 2022, powers many of today's AI inference factories, and continues to power model training. Through ongoing software optimization, NVIDIA increases the throughput of Hopper-based AI factories, leading to greater value.
On the Llama 2 70B benchmark, first introduced a year ago in MLPerf Inference v4.0, H100 GPU throughput has increased by 1.5x. The H200 GPU, based on the same Hopper GPU architecture with larger and faster GPU memory, extends that increase to 1.6x.
Hopper also ran every benchmark, including the newly added Llama 3.1 405B, Llama 2 70B Interactive and graph neural network tests. This versatility means Hopper can run a wide range of workloads and keep pace as models and usage scenarios grow more challenging.
It Takes an Ecosystem This MLPerf round, 15 partners submitted stellar results on the NVIDIA platform, including ASUS, Cisco, CoreWeave, Dell Technologies, Fujitsu, Giga Computing, Google Cloud, Hewlett Packard Enterprise, Lambda, Lenovo, Oracle Cloud Infrastructure, Quanta Cloud Technology, Supermicro, Sustainable Metal Cloud and VMware.
The breadth of submissions reflects the reach of the NVIDIA platform, which is available across all cloud service providers and server makers worldwide.
MLCommons' work to continuously evolve the MLPerf Inference benchmark suite to keep pace with the latest AI developments and provide the ecosystem with rigorous, peer-reviewed performance data is vital to helping IT decision makers select optimal AI infrastructure.
Learn more about MLPerf.
Images and video taken at an Equinix data center in the Silicon Valley.
Most recent headlines
03/04/2025
PARIS BeNarative, an innovative video production platform, has announced a technical and commercial partnership with Haivision, a major provider of live video c...
03/04/2025
Innovations across Premiere Pro and After Effects deliver AI-powered upgrades and workflow improvements, said the company
By Matthew Corrigan
Published: Apri...
03/04/2025
The new upgrades aim to enable smarter tracking, greater control, and faster workflows while expanding interoperability across robotic systems
By Matthew Corr...
03/04/2025
From a trip to space to the inaugural Sports Summit, and a look behind the camera of Wicked, this is whats on the TVBEurope teams radar at the 2025 NAB Show
B...
03/04/2025
Because of the nature of this sort of sliding scale of tariffs, there are opportunities that werent there before, analyst Alice Enders tells TVBEurope
By Jenny...
03/04/2025
WASHINGTON The NAB is requesting the Federal Communications Commission make changes to Emergency Alert System (EAS) rules that would allow but not require EAS p...
03/04/2025
NEW YORK In the run-up to the 2025 NAB Show, Amagi has announced the establishment of a Broadcast Network Operations Center (NOC) in Princeton, N.J....
03/04/2025
ENGLEWOOD CLIFFS, N.J. CNBC has announced distribution deals that will launch its subscription streaming offering CNBC+ on Apple TV and Roku....
03/04/2025
BURY ST EDMUNDS, U.K. Vinten, a global provider in robotic camera support systems and a Videndum brand, will unveil significant advancements to its VEGA contro...
03/04/2025
Lightcraft Jetset Expands iPhone Virtual Production Tool with a Dozen New Featur...
03/04/2025
Celtx launches Screenplay Plugin to help editors automate post-production workfl...
03/04/2025
Maxon One Release Delivers Greater Creative Freedom and Workflow Performance for...
03/04/2025
Berklee Announces Paul Dworkis as Executive Vice President and Chief Financial O...
03/04/2025
Facebook
Twitter
LinkedIn
Thales, a global leader in advanced technologies...
03/04/2025
Facebook
Twitter
LinkedIn
The new flight training centre features Thales...
03/04/2025
RT Supporting the Arts | April 2025
This April, RT is delighted to support C irt International Festival of Literature, Incognito Art Sale for Jack and Jill ...
03/04/2025
RT has today launched Clarity, a new strand of coverage in which its journalism will work to counter the deliberate manipulation of facts and challenge false a...
03/04/2025
GeForce NOW isn't fooling around.
This month, 21 games are joining the cloud gaming library of over 2,000 titles. Whether chasing epic adventures, testing ...
02/04/2025
Pedro Pascal appears in Anna Boden and Ryan Fleck's Freaky Tales, which pr...
02/04/2025
With a focus on safeguarding premium content value and authenticity, NAGRA highlighted key areas of interest in the media and entertainment industry. Of note wa...
02/04/2025
In our latest blog, gain insights into the media industry's challenges and how NAGRA Active Streaming Protection provides a framework for holistic content p...
02/04/2025
In our latest blog Tim Pearson considers Generative AI and the opportunities it presents as well as some of the challenges it can cause for media, entertainment...
02/04/2025
Learn valuable insights into strengthening your content protection strategy and discover how multi-DRM helped transform content security for leading post-produc...
02/04/2025
This year's IBC 2024 was an incredible opportunity to connect with industry leaders and innovators, and the conversations around consumer cybersecurity were...
02/04/2025
As a lifelong sports enthusiast from the U.S., I've always been captivated by how sports can unite people. From the roar of the crowd during major events to...
02/04/2025
In our latest blog, Tim Pearson caught up with Julian Williams at Anthropic to explore the science of conversations and how the increasing adoption of generativ...
02/04/2025
In our latest blog, Tim Pearson considers recent industry successes in dismantling large-scale pirate operations and what defensive steps video service provider...
02/04/2025
In our latest blog, Laura Rognoni explores OpenTV ENTera, the latest innovation from NAGRAVISION that's designed as a blueprint for today's streaming se...
02/04/2025
Scott Alexander, President of Missile Solutions, Aerojet Rocketdyne, L3Harris, writes in Breaking Defense: L3Harris is building the factories of the future that...
02/04/2025
Calrec's Argo S ramps up Raycom's output for OTT, FAST and OTA channels North Carolina's Raycom Sports has upgraded its flagship RHD1 mobile product...
02/04/2025
Calrec expands ecosystem at NAB 2025 giving broadcasters access to dynamic workflows and ultimate flexibility Helping broadcasters meet the shifting needs of me...
02/04/2025
aconnic AG (ISIN: DE000A0LBKW6), Munich, is launching a new 10 Gigabit Carrier Ethernet system for industrial application. The ACCEED 4108 DR provides full MEF ...
02/04/2025
TV Tech: What do you anticipate will be the most significant technology trends at the 2025 NAB Show?...
02/04/2025
SKY and DGO, the streaming and live TV platforms of DIRECTV Latin America and SKY Brasil, are moving forward with consolidating the highest-level experience fo...
02/04/2025
IABM is delivering a strategic transformation at NAB Show designed to fiercely champion members amidst global, industry challenges, elevating and innovating to ...
02/04/2025
Following a well-attended February 27th-28th GovSatCom in Luxembourg, Hiltron Communications promoted its wide range of satellite communication products, system...
02/04/2025
MASV, the fastest large file transfer platform for media professionals, is revolutionizing enterprise media workflows by enabling faster, more reliable, and sca...
02/04/2025
AgileTV, a leader in TV and video technology solutions, is partnering with CANAL Germany, the leading B2B TV-licensing provider in Germany, to introduce "The E...
02/04/2025
New model leverages 20Gbps USB 3.2 Gen 2x2 interface to capture 12G SDI without a driver or external power
Magewell, developer of innovative, high-performance ...
02/04/2025
MwareTV, a leading cloud-based multi-tenant TV platform provider, is set to launch a ground-breaking new toolset at NAB 2025 (booth W3457, Las Vegas Convention ...
02/04/2025
LiveU will spotlight its latest technical collaborations around efficient story-centric workflows and cloud collaboration in its expanded EcoSystem at the upcom...
02/04/2025
Live Media Group, a leader in live broadcast solutions and event production, has named Ryan Hatch as Vice President, Strategic Accounts, effective April 1st. In...
02/04/2025
New AI Innovation in Industry-Leading Adobe Premiere Pro Empowers Video Pros to ...
02/04/2025
DigitalGlue and Symply Partner to Deliver Next-Generation Storage Solutions for ...
02/04/2025
Music Therapy Students Awarded First Internship Stipend from Children's Musi...
02/04/2025
WASHINGTON The National Association of Broadcasters (NAB) will present the Television Chairman's Award to renowned magicians and television personalities, P...
02/04/2025
MINNEAPOLIS-ST. PAUL The Minnesota Twin have inked a new, multi-year partnership with Gray Media and FOX 9, KMSP, to broadcast 10 Tuesday night regular season g...
02/04/2025
SAN JOSE Adobe today announced the official launch of its Generative Extend AI tool for Premiere Pro. The feature announced at its Adobe Max conference last fa...
02/04/2025
Sally Wallington, SVP of sales at Pebble, explores the mission-critical considerations broadcasters should make when choosing a playout provider
Sponsored Cont...
02/04/2025
TVBEurope meets Tim Claman, chief product officer at Avid, to discuss the compan...