The artificial intelligence race is shifting from training models to running them efficiently at scale. This process, known as inference, occurs whenever a trained AI model receives new information and generates an answer, prediction, image or action. Every chatbot response, autonomous-agent decision and real-time recommendation requires inference compute.
For technology companies, the central challenge is no longer simply building the most capable model. It is delivering billions of outputs quickly, reliably and at a commercially sustainable cost.
Saudi Arabia’s national AI company, HUMAIN, underscored that priority by appointing technology entrepreneur Michael Carter as Senior Vice-President of Inference Services. Carter was an early investor in AI-chip company Groq and founded mobile gaming company Playco.
HUMAIN said Carter will lead its inference business by “turning AI compute at scale into powerful model APIs for customers around the world.” The appointment signals that inference is becoming a distinct business line capable of transforming large computing investments into accessible, revenue-generating AI services. HUMAIN announcement
The world’s largest technology companies are making the same calculation. Microsoft introduced Maia 200 specifically to improve the economics of AI inference, while Google describes Ironwood as its first TPU designed specifically for inference. Amazon Web Services has developed Inferentia chips to provide cost-efficient inference through its cloud platform. Microsoft Google AWS
Startups are also developing new infrastructure models. Hamburg-based CYANIS AI is building a distributed European inference network that places modular computing capacity at existing household, industrial and utility-scale energy sites. Its platform is designed to orchestrate these resources while supporting European data sovereignty and lower-latency AI services.
In July, CYANIS reported validating orchestration across multiple GPU types and cloud providers, including models ranging from 0.8 billion to 675 billion parameters. The company is now preparing an initial commercial rollout targeting 10 MW of capacity. CYANIS AI Technical brief
Is inference emerging as the commercial and infrastructure layer through which AI will be delivered at scale?

