MangoBoost

First-ever Multi-Region Heterogeneous Inference: MangoBoost Sets New Standards for Global AI Scaling in MLPerf Inference v6.0

First-ever Multi-Region Heterogeneous Inference: MangoBoost Sets New Standards for Global AI Scaling in MLPerf Inference v6.0

Key Highlights

  • First-ever multi-region GPU deployment: Orchestrated a heterogeneous mixture of GPU clusters across separate networks in the U.S. and South Korea, achieving near-perfect 95.5% performance scaling on Llama2-70B.
  • Multi-region deployment allows the system to route users to the closest regional cluster, ensuring minimal latency and global load balancing without requiring complex network overhauls.
  • First-ever deployment of heterogeneous serving using 3 GPU architectures: We successfully deployed a complex multi-node architecture across 3 generations of AMD Instinct™ GPUs: the MI355X, MI325X and MI300X.
  • Pioneering power-capped results: The first-ever submission of power-capped results, demonstrating the remarkable effectiveness of the AMD Instinct™ MI355X operating under a limited power budget.
  • Continuous Software Optimization: Delivered a 2.4% throughput increase on AMD Instinct™ MI300X online results compared to MLPerf v5.1, alongside continued, optimized support for the entire AMD Instinct™ GPU server lineup (MI300X, MI325X, and MI355X).

Pushing the Boundaries of LLM Inference

MangoBoost is proud to announce groundbreaking results in the MLPerf Inference v6.0 benchmark. This round pushes the absolute frontier of LLM inference serving, showcasing how LLMBoost™ unlocks the full potential of global, multi-region GPU clusters.

By combining advanced software orchestration, geo-tied mapping, and deep collaborations with Dell and AMD, we continue to redefine performance, scalability, and Total Cost of Ownership (TCO) for large-scale, international GenAI deployments. Our MLPerf 6.0 submission moves beyond traditional, localized data-center metrics to prove that enterprise AI can thrive across complex, disparate networks. Below is a deep dive into the technical breakthroughs that define our submission:

Shattering Geographic Boundaries with MLPerf’s First Multi-Region GPU Submission

Historically, achieving record-breaking AI inference performance has required tightly coupled, homogeneous clusters confined to a single data center. High-speed, local interconnects were considered an absolute necessity to prevent latency bottlenecks. In MLPerf Inference v6.0, MangoBoost shatters this paradigm.

For the first time in MLPerf history, we submitted results utilizing a multi-region, geographically distributed GPU system. This was not a simulation; it was a live, trans-Pacific deployment that tested the absolute limits of software orchestration, network resilience, and hardware flexibility.

The Architecture: A Trans-Pacific AI Cluster

To prove the real-world viability of LLMBoost™, we orchestrated a highly complex, heterogeneous mixture of multi-region clusters separated by thousands of miles. The deployment spanned two entirely separate networks and geographic zones:

  • The Lead Node (United States): Powered by Dell’s state-of-the-art server infrastructure, we utilized AMD’s cutting-edge AMD Instinct™ MI355X GPUs to act as the primary orchestration and compute engine.
  • The Worker Nodes (South Korea): Hosted on MangoBoost’s proprietary server infrastructure, we integrated previous-generation AMD Instinct™ MI325X and AMD Instinct™ MI300X GPUs.

This created a massive, distributed pipeline serving the Llama2-70B model, marking the first-ever heterogeneous serving system to operate simultaneously across three distinct generations of AMD Instinct™ GPU architectures.

Figure 1: MLPerf Inference v6.0 Benchmark on Trans-Pacific Heterogeneous Cluster Architecture

Achieving 95.5% Scaling on multi-region Llama2-70B deployment

When operating heterogeneous hardware, the standard risk is that the pipeline moves only as fast as its slowest component. When you add geographic distance (U.S. to South Korea) into the equation, traditional inference frameworks collapse under the weight of communication latency and pipeline bubbles.

Despite these severe topological challenges, LLMBoost™ achieved a near-perfect 95.5% performance scaling on the massive Llama2-70B model. This exceptional metric means that the combined U.S. and South Korea system delivered 95.5% of the theoretical maximum throughput that these GPUs would have achieved had they been sitting in the exact same server rack.

Figure 2: Performance scaling of Llama2-70B across a multi-region deployment (U.S./South Korea) combining MI355X, MI325X, and MI300X GPUs. LLMBoost achieves 95.5% of the theoretical maximum performance, virtually eliminating WAN latency penalties

How LLMBoost Made This Possible:

  • Heterogeneous Load Balancing across three GPU architectures and two regions: The software automatically profiled the differing compute capabilities of the AMD Instinct™ MI355X, MI325X, and MI300X in real-time. It then unevenly distributed the model layers and batch workloads, ensuring each GPU generation operated at exactly 100% of its specific capacity without bottlenecking the others.
  • Network-Resilient Orchestration: With orchestration by MangoBoost’s LLMBoost software, the system maintained stable, high-throughput token generation despite the inherent jitter of a trans-oceanic network connection.

A Proof Point for Unparalleled Enterprise Flexibility

This 95.5% scaling metric is more than just a technical flex; it is a powerful proof point for enterprise IT and AI infrastructure teams.

It demonstrates Dell’s robust server architectures and AMD’s Instinct™ hardware, proving they can operate flawlessly in highly non-traditional deployments. More importantly, it highlights LLMBoost’s ability to maintain peak performance within complex or pre-existing network architectures, allowing inference requests to be served with geolocation in mind.

Organizations no longer need to execute massive overhauls to upgrade their AI capabilities. With LLMBoost, a company can deploy Dell's newest AMD Instinct™ MI355X servers in a new data center and seamlessly link them with legacy AMD Instinct™ MI300X servers sitting in other facilities, operating them all as a single, unified, ultra-efficient system.

To learn more about our submission with Dell, click here.

Further Maximizing TCO with Effortless Integration

Unifying inference across disparate regions provides a transformative Total Cost of Ownership (TCO) advantage. This revolutionary architecture enables geo-tied mapping, where the serving system intelligently routes users to the regional cluster geographically closest to them. This provides superior global load balancing and drastically reduces latency.

Organizations with existing clusters can now "plug in" new nodes, regardless of their physical location or network isolation. By removing the need for a total network overhaul or complex re-architecting, LLMBoost ensures that teams can scale their AI capacity internationally whenever and wherever it is needed, maximizing every dollar of CapEx without sacrificing operational efficiency.

Power-Capped Efficiency and Enhanced Baseline Performance

As enterprise AI scales, data centers are increasingly hitting a hard ceiling: power availability. Compute density is no longer just about physical space; it is about how much performance can be extracted per watt. LLMBoost continues to squeeze every ounce of performance out of enterprise hardware, proving that top-tier throughput does not have to come at the expense of a blown power budget or a redesigned facility.

Pioneering Power-Capped MLPerf Submissions

In MLPerf Inference v6.0, MangoBoost broke new ground by delivering the first-ever power-capped results. We explicitly demonstrated the immense effectiveness of the AMD Instinct™ MI355X operating under a strictly limited power budget. By utilizing LLMBoost’s software, the system dynamically balances maximum token generation with strict power ceilings.

Figure 3: Power-Capped Inference Efficiency with LLMBoost on AMD Instinct™ MI355X

Why the 1000W Power Budget is a Critical Enterprise Milestone

As AI models grow exponentially, the GPUs required to run them are drawing unprecedented levels of power, with next-generation accelerators frequently pushing past the 1000W to 1200W Thermal Design Power (TDP) mark. Proving that the AMD Instinct™ MI355X can deliver state-of-the-art MLPerf inference results while strictly capped under a 1000W is a massive business enabler:

  • Maximizing Existing Rack Density: Most traditional enterprise data centers are heavily constrained by rack-level power limits (often capping out at 20kW to 40kW per rack). By keeping the GPU power under 1000W, enterprises can safely deploy dense AI nodes into existing infrastructure without tripping facility power limits or stranding expensive rack space.
  • Avoiding Costly Infrastructure Overhauls: Crossing the 1000W per-chip threshold often forces data centers into a mandatory, highly disruptive transition from advanced air cooling to complex direct-to-chip liquid cooling. Demonstrating top-tier performance safely under this limit proves that the MI355X, powered by LLMBoost, can be deployed rapidly without waiting for years-long facility retrofits.
  • Slashing OpEx and Meeting ESG Goals: Power and cooling constitute the largest ongoing Operational Expenditures (OpEx) for AI deployments. Operating efficiently within a capped 1000W budget directly translates to massive electricity savings, enabling organizations to scale their AI capabilities while strictly adhering to corporate sustainability and carbon reduction (ESG) targets.

Continuous Software Optimization: Maximizing Performance Across the AMD Instinct™ Portfolio

True enterprise software doesn't just support the newest hardware; it continuously revitalizes existing infrastructure. In an industry heavily focused on the latest silicon releases, MangoBoost remains deeply committed to pushing the cutting-edge of inference, as well as extracting every ounce of value from previous-generation deployments. Without altering the underlying hardware, we pushed our AMD Instinct™ MI300X online results further, adding 2.4% more throughput compared to our record-setting MLPerf 5.1 submission.

Figure 4: Llama2-70B throughput increase on MI300X server scenario from MLPerf Inference v5.1 to v6.0

This gain is a direct result of LLMBoost’s relentless refinement of memory management, continuous batching mechanisms, and full-stack auto-tuning. This sustained improvement highlights our commitment to maximizing our customers' hardware investments, while continuing to solidify our seamless, optimized support for the entire AMD Instinct™ GPU server portfolio, spanning the AMD Instinct™ MI300X, MI325X, and MI355X GPUs.

Our MLPerf v6.0 results are a testament that infrastructure shouldn't be left behind just because a newer chip hits the market. By continuously optimizing LLMBoost across all generations, we ensure that our customers' existing servers don't just avoid obsolescence; they actually appreciate in performance long after they are deployed.

A Record-Setting Collaboration with AMD

Our groundbreaking MLPerf v6.0 submission is a powerful testament to our close, ongoing partnership with AMD as a catalyst for redefining enterprise AI inference. AMD’s Instinct™ accelerators deliver phenomenal compute density and massive memory bandwidth. When paired with MangoBoost’s LLMBoost™ orchestration software, we create a powerful synergy that pushes the absolute boundaries of what inference solutions can achieve.

Together, we are proving that enterprises can seamlessly bridge multiple generations of AMD silicon, maximize throughput across distributed networks, and achieve unprecedented hardware utilization at a global scale. The result is an inference solution that is not only exceptionally fast, but incredibly flexible.

Meena Arunachalam, Fellow and Director, AI Workloads Performance Engineering at AMD, captures the technical magnitude of what our combined technologies have achieved in this latest MLPerf round:

"Through our continued collaboration, MangoBoost's LLMBoost is redefining what large-scale AI inference looks like in the first-ever multi-region MLPerf submission using three generations of AMD Instinct™ GPUs. Coordinating Dell’s AMD Instinct™ MI355X servers in the U.S. alongside AMD Instinct™ MI325X and AMD Instinct™ MI300X hardware in Korea to achieve a 95.5% scaling efficiency across two continents is a significant technical achievement. Beyond this, we are seeing LLMBoost's continuous throughput improvements across the entire AMD Instinct™ family. By pairing these baseline gains with new power-capped efficiencies, the combination of MangoBoost’s LLMBoost and AMD Instinct™ GPUs proves that enterprises can deploy massive, globally distributed AI pipelines while remaining strictly within their existing infrastructure and power constraints."

Turn-Key AI Infrastructure: Idea to Production in Minutes

LLMBoost provides a turn-key solution bridging the gap from idea to production, allowing developers to serve open models at any scale, instantly.

Beyond Software: MangoBoost Hardware Solutions

Alongside LLMBoost™, MangoBoost accelerates infrastructure from the ground up with advanced DPU-based hardware acceleration:

  • Mango BoostX™: A powerful FPGA-based DPU designed for offloading networking, storage, and security workloads.
  • Mango BoostX™ RNIC: A RoCEv2 NIC for ultra-fast, intra-rack GPU-to-GPU connectivity.
  • Mango Storage Server: A high-density, disaggregated storage platform built for high-bandwidth, low-latency NVMe access.
  • Mango GPU Server: A full-stack system purpose-built to maximize AI performance and optimize tokens/$.

Try LLMBoost™ Today

To experience the globally scalable, record-setting performance of MLPerf v6.0, register on our virtual demo page.