$450 Billion in AI Infrastructure, Most of It Wasted: ApexData Launches RidgeScope
Most of the compute in this year's
Behind that number sits what the industry calls the MFU gap – the shortfall between compute bought and compute used. Model FLOPs utilization, the share of a GPU's theoretical compute that actually advances the model, typically sits between 30 and 40 percent; fleet studies measure averages closer to 20 percent. More than half of every dollar spent on GPU compute is lost to slow chips, starved data pipelines, communication stalls, hardware faults and runs that finish cleanly while learning nothing.
At about
"The industry is financing GPUs as if they were fully productive assets, yet most clusters deliver less than half of what the spec sheet promises," said
A lightweight agent on each server collects more than 12,000 signals per training run: GPU behavior, interconnect traffic, job logs and scheduler records. An AI engine weighs the evidence against more than 20 known failure patterns, ruling each in or out. Every case ends in a verdict: what the data shows, what it implies and what it cannot decide, each claim tied to the measurement behind it.
"GPU waste is quiet. One slow chip drags down 127 healthy ones, a job saves its progress so often that it stops making any, another sits on expensive hardware computing nothing, and every dashboard stays green," said Andrey Shamakhov, co-founder and CTO of ApexData. "We name the failure, show the evidence and say what it costs. And when the data cannot answer something, the verdict says so instead of guessing."
For GPU cloud operators, RidgeScope settles the question behind every support ticket: customer code or cluster hardware.
RidgeScope recognizes Hugging Face Transformers, Megatron-LM, PyTorch Lightning and Keras/TensorFlow.
For most teams, training data is the company's most valuable intellectual property. RidgeScope reads only system telemetry and scheduler metadata – never datasets, source code or model weights – and can run entirely inside the customer's own network: on-premise or air-gapped, with local LLM models, per-tenant isolation, SSO/SAML, and GDPR- and PCI DSS-grade encryption.
Generally available today, RidgeScope installs in minutes on Slurm and Kubernetes and returns a first verdict within 30 minutes. A public Failure Catalog and Evidence Index list every failure mode and every signal a verdict may cite. Demonstrations run live at https://ridgescope.ai.
About ApexData
ApexData Inc. (Sunnyvale, California) builds AI-powered observability for ML training and infrastructure, drawing on two decades of DevOps and distributed-systems experience.
Media Contact:
Andrei Surkov
+14083298998
[email protected]
View original content to download multimedia:https://www.prnewswire.com/news-releases/450-billion-in-ai-infrastructure-most-of-it-wasted-apexdata-launches-ridgescope-302828765.html
SOURCE ApexData Inc
Serious News for Serious Traders! Try StreetInsider.com Premium Free!
You May Also Be Interested In
- Aizen Enters into Collaboration with San Diego Biopharma to Advance Oral Biologics
- FRONTERA ANNOUNCES SECOND QUARTER 2026 RESULTS
- 77 DAYS TO GO: 7 WONDERS OF FUTURE CITIES COUNTS DOWN TO START OF VOTING
Create E-mail Alert Related Categories
PRNewswire, Press ReleasesSign up for StreetInsider Free!
Receive full access to all new and archived articles, unlimited portfolio tracking, e-mail alerts, custom newswires and RSS feeds - and more!



Tweet
Share