About the project
An enterprise AI company develops generative AI systems for document intelligence, enterprise knowledge management, and domain-specific language models. Its ML engineering team performs model training, fine-tuning, evaluation, synthetic-data generation, embedding generation, and large-scale inference. Initially, most GPU workloads ran in public cloud environments.
Challenge
Cloud GPUs worked well while GPU usage was intermittent. As the company moved from experimentation to continuous production and R&D workloads, its infrastructure requirements changed. The engineering team increasingly needed eight H100-class GPUs simultaneously for distributed workloads.
Key requirements included:
- 8x NVIDIA H100 GPUs in one system;
- NVLink connectivity;
- hundreds of gigabytes of GPU memory;
- large system memory capacity;
- dedicated GPU availability;
- fast storage;
- high-speed networking;
- predictable monthly cost.
Cloud capacity remained useful for temporary bursts, but the company’s baseline GPU consumption had become almost continuous.
Solution
The company moved its baseline AI compute capacity to a dedicated multi-GPU platform.
AZM-12 Multi-GPU Dedicated Server – Netherlands: 2x Intel Xeon Gold 6448Y, 8x NVIDIA H100 NVLink, 640 GB total HBM2e, 2 TB RAM, 2 TB NVMe, 10 Gbps network.
Used for: distributed LLM training, fine-tuning, preference optimization, model evaluation, synthetic-data generation, embeddings, and large-scale inference.
The eight H100 GPUs can be assigned to one distributed workload or divided between multiple smaller jobs.
The company maintains its ML software environment permanently on the server, including:
- NVIDIA drivers and CUDA;
- PyTorch;
- distributed training frameworks;
- containerized ML environments;
- model caches;
- dataset caches;
- monitoring;
- job scheduling.
Public cloud GPUs remain part of the architecture, but primarily for temporary capacity beyond the dedicated baseline.
Hybrid GPU architecture
- Dedicated 8x H100 server: continuous model training, fine-tuning, production inference, evaluation, development workloads.
- Cloud GPUs: short-term additional capacity, unusually large experiments, temporary demand spikes.
Results
Average utilization of paid GPU capacity increased from approximately 35 – 45% to more than 70%. A persistent GPU environment also reduced the operational overhead associated with repeatedly creating cloud instances, synchronizing datasets, restoring checkpoints, downloading containers, and rebuilding model caches. Large jobs can use all eight H100 GPUs continuously, while unused GPUs can be reassigned to inference, evaluation, or development workloads. Most importantly, baseline AI compute spending became predictable. For workloads requiring continuous H100 usage, dedicated GPU infrastructure provides a fundamentally different cost model from purchasing GPU capacity by the hour.
What’s next
The company plans to separate training and production inference as workloads increase. A second 8-GPU system could also form the basis of a multi-node GPU cluster. Results at a glance: 8x NVIDIA H100 NVLink, 640 GB GPU memory, 70%+ GPU utilization, persistent training environment, predictable GPU capacity.