HPC clusters are still run like it's 2010: bare-metal to running cluster is a weeks-long manual slog, and silent failures burn compute-hours for hours or days before anyone notices. Every vendor's stack speaks its own dialect. Nobody's built one intelligence layer that spans all of them.
Bare-metal-to-cluster deployment is hand-run and error-prone. Every week spent provisioning is a week of amortized hardware cost with nothing running on it.
GPU ECC errors, thermal drift, interconnect degradation: invisible until a job fails hours in. Your best engineers become firefighters instead of builders.
Every scheduler and provisioner has its own alerts, its own blind spots. Nobody owns the one view that would have caught it in time.
Two acquisitions eliminated every production-grade, independent HPC management option. If your hardware isn't Nvidia, you're now on your own.
The dominant independent HPC cluster manager. Deployed at national labs, universities, and hyperscalers worldwide. Now absorbed into Nvidia Base Command. Vendor-locked to Nvidia hardware.
The leading GPU workload scheduler for AI infrastructure. Kubernetes-native, widely adopted for AI training clusters. Now Nvidia. Acquisition flagged by EU/UK regulators for anti-competitive lock-in.
OpenHPC: Open-source, DIY, no enterprise support.
Penguin Solutions: Government/DoD-focused, services-heavy, not AI-native.
The gap: No production-grade, vendor-neutral option for AMD, IBM, or mixed-hardware clusters.
Panacea orchestrates provisioning, health, and remediation across every major HPC and Kubernetes stack — with AI-powered prediction on top, not bolted on after. Vendor-neutral by design. AMD-first by conviction.
Production Go implementation with full OpenTelemetry observability: every run emits metrics, logs, and traces, not just a final score. The same framework is the continuous health signal Panacea's self-healing layer consumes.
Every benchmark ships as a typed production Go CLI with structured JSON config and result output — designed to run continuously across live fleets, not just at bring-up.
A node that regresses on STREAM or RCCL gets flagged automatically before it silently corrupts a job's wall-clock time. This is what the best HPC ops teams do manually today. Panacea does it continuously.
Core42 use case: Full benchmarking pipeline for MI300X deployment validation and ongoing health monitoring.
These are the design targets the platform is engineered against: the numbers an HPC ops team actually has to hit, not the numbers that look good in a deck.
Every major enterprise and AI lab is standing up GPU superclusters right now — most run by teams who have never operated infrastructure at this scale. The bottleneck stopped being compute. It became keeping compute alive and utilized.
GPU cluster deployment is happening faster than HPC operations expertise can be hired. Core42, G42, Humain, and national AI programs are spending billions on hardware that still needs someone to run it.
Nvidia's acquisition of Bright + Run:ai isn't about those products. It's about owning the management layer so AMD wins on silicon but loses on stickiness. Every AMD customer now needs an alternative.
Whoever becomes the default control plane for this generation of clusters owns the relationship for a decade. Infrastructure switching costs don't forgive being late. The window is now.
Every dollar of GPU capex needs ops software. Nvidia just ensured theirs is the only option for Nvidia hardware — creating a captive market for AMD, IBM, and mixed-hardware clusters that represents billions in annual software spend.
Every alternative is either Nvidia-locked, government-focused, or a DIY science project. Panacea is the only production-grade option built for AMD-first, mixed-hardware AI infrastructure at hyperscale.
| Solution | Vendor-Neutral | AMD / MI300X | AI/GPU Native | Self-Healing | Production Scale | Status |
|---|---|---|---|---|---|---|
| Panacea Ops | ✓ Yes | ✓ AMD-first | ✓ Yes | ✓ ML-driven | ✓ 10k+ nodes | Available now |
| Nvidia Base Command (Bright) | ✗ Nvidia only | ✗ No | ✓ Yes | ⚠ Partial | ✓ Yes | Nvidia hardware only |
| Run:ai (Nvidia) | ✗ Nvidia only | ✗ No | ✓ Yes | ✗ No | ✓ Yes | Nvidia hardware only |
| OpenHPC | ✓ Yes | ✓ Yes | ✗ No | ✗ No | ⚠ DIY only | Expertise required |
| Penguin Solutions | ✓ Yes | ⚠ Limited | ✗ No | ✗ No | ⚠ Gov/DoD focus | Services-heavy |
UAE's national AI infrastructure company. Operator of some of the world's largest GPU clusters — MI300X, H200, and mixed-hardware at hyperscale.
In early discussion for: benchmarking, deployment, and total system management as a pilot engagement.
Core42's worldwide ops manager is engaged. This is the exact customer profile Panacea is built for — massive AMD hardware deployments with no Nvidia-native management layer.
Infrastructure management software is the stickiest category in enterprise — switching costs are enormous and contract lengths are long. One Core42-scale customer is a multi-year, multi-million dollar relationship.
Annual software license per managed node. Tier pricing: research (<500 nodes), enterprise (500–5k), hyperscale (5k+).
Target ARR per hyperscale customer: $500K–$2M+
Recurring, predictable, grows with the cluster.
Deployment, integration, and onboarding for new cluster bring-ups. Premium support SLAs for mission-critical environments.
High-margin near-term revenue while the platform scales. This is also the foot-in-the-door: every deployment becomes a long-term license.
Full-service ops for operators who want Panacea running their cluster without staffing the expertise internally.
The highest-value offering for new entrants (Core42-adjacent customers) who have the hardware but not the HPC ops depth.
For 19 years, Kris has been the person other people call when a supercomputer goes dark. He built and ran Piz Daint (TOP500 #3) — where a mistake doesn't mean a bad quarter, it means a physics result that doesn't ship or a forecast that arrives too late.
He helped build the CERN LHC Grid infrastructure that confirmed the Higgs boson. He built BioHive-1 (TOP500 #84) in 90 days, $1.6M under budget, during a global supply chain crisis — because the deadline didn't move for anyone.
Currently at AMD as Principal Member of Technical Staff, DCGPU-Perf Group — benchmarking and optimizing MI300X for HPC and AI workloads. The exact problem Panacea solves.
Every infrastructure startup claims domain expertise. Few have actually been the last line of defense when ten-figure hardware goes silent at 3am.
Panacea is two decades of "what actually breaks at 3am" encoded into software — so the next generation of operators doesn't have to carry the pager the way he did.
The Core42 relationship isn't cold outreach. It's a peer-to-peer conversation between people who have actually run these machines. That's the moat.
The platform is real and operating: Go benchmark suite, provisioning and self-healing services, enterprise auth integrations — all implemented. This round funds Core42 deployment and go-to-market.
Nvidia locked the independent HPC management stack. AMD is winning on silicon. The operator who fills the gap between great hardware and great software wins the next decade of AI infrastructure. That's Panacea.