Seed Pitch · 2026

Your GPUs are dying quietly. Nobody's watching.

Panacea Ops is the AI-native control plane that catches it before the invoice does.
Zero-Touch Provisioning Self-Healing Architecture AMD MI300X · NCCL/RCCL · HPL · HPCG · IOR · STREAM 10k+ Node Scale Vendor-Neutral
FounderKris Howard — 30yr Unix, 19yr HPC (Piz Daint TOP500 #3; LHC Grid / Higgs boson)
Contactkris@panacea-ops.com
StageSeed · $2.5M
1 / 13
The Problem

You find out about failures from the invoice, not the dashboard.

HPC clusters are still run like it's 2010: bare-metal to running cluster is a weeks-long manual slog, and silent failures burn compute-hours for hours or days before anyone notices. Every vendor's stack speaks its own dialect. Nobody's built one intelligence layer that spans all of them.

Weeks, not minutes

Bare-metal-to-cluster deployment is hand-run and error-prone. Every week spent provisioning is a week of amortized hardware cost with nothing running on it.

Failures that hide

GPU ECC errors, thermal drift, interconnect degradation: invisible until a job fails hours in. Your best engineers become firefighters instead of builders.

A console for everything, control over nothing

Every scheduler and provisioner has its own alerts, its own blind spots. Nobody owns the one view that would have caught it in time.

2 / 13
Market Shift

Nvidia bought the competition. The independent stack is gone.

Two acquisitions eliminated every production-grade, independent HPC management option. If your hardware isn't Nvidia, you're now on your own.

Bright Computing — Acquired Jan 2022

The dominant independent HPC cluster manager. Deployed at national labs, universities, and hyperscalers worldwide. Now absorbed into Nvidia Base Command. Vendor-locked to Nvidia hardware.

Run:ai — Acquired 2024 (~$700M)

The leading GPU workload scheduler for AI infrastructure. Kubernetes-native, widely adopted for AI training clusters. Now Nvidia. Acquisition flagged by EU/UK regulators for anti-competitive lock-in.

What's left

OpenHPC: Open-source, DIY, no enterprise support.

Penguin Solutions: Government/DoD-focused, services-heavy, not AI-native.

The gap: No production-grade, vendor-neutral option for AMD, IBM, or mixed-hardware clusters.

3 / 13
The Solution

One intelligent control plane, from bare metal to running jobs.

Panacea orchestrates provisioning, health, and remediation across every major HPC and Kubernetes stack — with AI-powered prediction on top, not bolted on after. Vendor-neutral by design. AMD-first by conviction.

Zero-Touch Provisioning

  • Fully automated deployment, bare metal to running cluster
  • Template-driven workflows, no hand-tuned per-node scripts
  • PDU control, BMC integration, power cycling, thermal monitoring built in
  • Multi-system: Slurm, Bright, xCAT, Warewulf, OpenHPC, Kubernetes, PBS Pro, SGE
  • AMD MI300X native support · RCCL-optimized benchmarking pipeline

Self-Healing Architecture

  • ML-based predictive failure detection with confidence scoring
  • Automated remediation for silent errors before jobs fail
  • Automated failover, node replacement, workload redistribution
  • Real-time anomaly detection across CPU, GPU, memory, network, power, thermal
  • Continuous benchmark loop — NCCL/RCCL regressions flagged before they corrupt jobs
4 / 13
Proof, Not Promises

A real benchmark framework, not a marketing claim.

Production Go implementation with full OpenTelemetry observability: every run emits metrics, logs, and traces, not just a final score. The same framework is the continuous health signal Panacea's self-healing layer consumes.

Standard HPC suite, built in

  • NCCL multi-GPU collectives · NVIDIA
  • RCCL multi-GPU collectives · AMD
  • DeepSpeed distributed training
  • GPCNeT network congestion
  • HPL / HPCG TOP500-class compute
  • IOR / FIO storage I/O
  • STREAM memory bandwidth

Built for scale, not demos

Every benchmark ships as a typed production Go CLI with structured JSON config and result output — designed to run continuously across live fleets, not just at bring-up.

A node that regresses on STREAM or RCCL gets flagged automatically before it silently corrupts a job's wall-clock time. This is what the best HPC ops teams do manually today. Panacea does it continuously.

Core42 use case: Full benchmarking pipeline for MI300X deployment validation and ongoing health monitoring.

5 / 13
Engineering Targets

Built against real operator SLAs, not vanity metrics.

These are the design targets the platform is engineered against: the numbers an HPC ops team actually has to hit, not the numbers that look good in a deck.

10,000+
Nodes per supercluster
Primary design target
99.9%
Managed cluster uptime
Availability target
<30 min
Full cluster deployment
Bare metal to running, target
90%+
Resource allocation efficiency
Utilization target
6 / 13
Why Now

The AI buildout made this urgent. The acquisitions made it inevitable.

Every major enterprise and AI lab is standing up GPU superclusters right now — most run by teams who have never operated infrastructure at this scale. The bottleneck stopped being compute. It became keeping compute alive and utilized.

The buildout

GPU cluster deployment is happening faster than HPC operations expertise can be hired. Core42, G42, Humain, and national AI programs are spending billions on hardware that still needs someone to run it.

The lock-in play

Nvidia's acquisition of Bright + Run:ai isn't about those products. It's about owning the management layer so AMD wins on silicon but loses on stickiness. Every AMD customer now needs an alternative.

The window

Whoever becomes the default control plane for this generation of clusters owns the relationship for a decade. Infrastructure switching costs don't forgive being late. The window is now.

7 / 13
Market Size

A multi-billion dollar market with no vendor-neutral leader.

Every dollar of GPU capex needs ops software. Nvidia just ensured theirs is the only option for Nvidia hardware — creating a captive market for AMD, IBM, and mixed-hardware clusters that represents billions in annual software spend.

$8.9B
HPC software market 2024
6.5% CAGR · MarketsandMarkets
$2.1B
AI infrastructure management segment
GPU cluster ops & orchestration
$180B+
GPU cluster capex 2024–2026
Hyperscalers + national AI programs
0
Vendor-neutral production-grade alternatives
Post-Nvidia acquisitions. The gap Panacea fills.
8 / 13
Competition

Vendor-neutral. Production-grade. The only one.

Every alternative is either Nvidia-locked, government-focused, or a DIY science project. Panacea is the only production-grade option built for AMD-first, mixed-hardware AI infrastructure at hyperscale.

Solution Vendor-Neutral AMD / MI300X AI/GPU Native Self-Healing Production Scale Status
Panacea Ops ✓ Yes ✓ AMD-first ✓ Yes ✓ ML-driven ✓ 10k+ nodes Available now
Nvidia Base Command (Bright) ✗ Nvidia only ✗ No ✓ Yes ⚠ Partial ✓ Yes Nvidia hardware only
Run:ai (Nvidia) ✗ Nvidia only ✗ No ✓ Yes ✗ No ✓ Yes Nvidia hardware only
OpenHPC ✓ Yes ✓ Yes ✗ No ✗ No ⚠ DIY only Expertise required
Penguin Solutions ✓ Yes ⚠ Limited ✗ No ✗ No ⚠ Gov/DoD focus Services-heavy
9 / 13
Traction

Early signal from the operators who matter.

Core42 — Early Discussions

UAE's national AI infrastructure company. Operator of some of the world's largest GPU clusters — MI300X, H200, and mixed-hardware at hyperscale.

In early discussion for: benchmarking, deployment, and total system management as a pilot engagement.

Core42's worldwide ops manager is engaged. This is the exact customer profile Panacea is built for — massive AMD hardware deployments with no Nvidia-native management layer.

Platform Validation

  • Production Go codebase: provisioning, benchmarking, self-healing services implemented
  • Enterprise auth integrations shipped: Okta, LDAP/AD, Vault, 1Password
  • CI pipeline running on self-hosted runners
  • Lint-clean codebase · 157 tests passing / 0 failing
  • AMD MI300X as primary hardware target · RCCL pipeline on roadmap
  • IP protected under Schedule A (Poser8 LLC)
200+
Commits merged
Built in under 45 days
438
Test files
725K lines of Go · 1,713 source files
Core42
Pilot in discussion
World's largest AMD GPU deployments
10 / 13
Business Model

Enterprise software economics with infrastructure stickiness.

Infrastructure management software is the stickiest category in enterprise — switching costs are enormous and contract lengths are long. One Core42-scale customer is a multi-year, multi-million dollar relationship.

Platform License

Annual software license per managed node. Tier pricing: research (<500 nodes), enterprise (500–5k), hyperscale (5k+).

Target ARR per hyperscale customer: $500K–$2M+

Recurring, predictable, grows with the cluster.

Professional Services

Deployment, integration, and onboarding for new cluster bring-ups. Premium support SLAs for mission-critical environments.

High-margin near-term revenue while the platform scales. This is also the foot-in-the-door: every deployment becomes a long-term license.

Managed Operations

Full-service ops for operators who want Panacea running their cluster without staffing the expertise internally.

The highest-value offering for new entrants (Core42-adjacent customers) who have the hardware but not the HPC ops depth.

11 / 13
The Unfair Advantage

You can't hire this. You have to have lived it.

Kris Howard, Founder

For 19 years, Kris has been the person other people call when a supercomputer goes dark. He built and ran Piz Daint (TOP500 #3) — where a mistake doesn't mean a bad quarter, it means a physics result that doesn't ship or a forecast that arrives too late.

He helped build the CERN LHC Grid infrastructure that confirmed the Higgs boson. He built BioHive-1 (TOP500 #84) in 90 days, $1.6M under budget, during a global supply chain crisis — because the deadline didn't move for anyone.

Currently at AMD as Principal Member of Technical Staff, DCGPU-Perf Group — benchmarking and optimizing MI300X for HPC and AI workloads. The exact problem Panacea solves.

Why this matters

Every infrastructure startup claims domain expertise. Few have actually been the last line of defense when ten-figure hardware goes silent at 3am.

Panacea is two decades of "what actually breaks at 3am" encoded into software — so the next generation of operators doesn't have to carry the pager the way he did.

The Core42 relationship isn't cold outreach. It's a peer-to-peer conversation between people who have actually run these machines. That's the moat.

12 / 13
The Ask

$2.5M seed to take this from proven engine to funded company.

The platform is real and operating: Go benchmark suite, provisioning and self-healing services, enterprise auth integrations — all implemented. This round funds Core42 deployment and go-to-market.

Engineering — $1.5M (60%)

  • 2–3 senior HPC/infrastructure engineers
  • Field engineering for Core42 pilot deployment
  • RCCL + AMD MI300X optimization track
  • 18-month runway to Series A

Go-to-Market — $625K (25%)

  • Core42 pilot: deployment, integration, onboarding
  • 2–3 additional design partners (national lab, university HPC center, GPU cloud)
  • SC25 and ISC conference presence
  • Business development and partner channels

Operations — $375K (15%)

  • Owned test hardware (same discipline that built BioHive-1 under budget)
  • Legal: IP protection, Schedule A hardening, customer contracts
  • Infrastructure: CI/CD, security, compliance baseline
13 / 13
Let's Talk

Stop finding out about failures from the invoice.

Nvidia locked the independent HPC management stack. AMD is winning on silicon. The operator who fills the gap between great hardware and great software wins the next decade of AI infrastructure. That's Panacea.

Seed Round · $2.5M · Panacea Ops