All Courses

Planning and Optimizing AI Infrastructure

Chapter 1: Foundations of AI Compute Infrastructure

Introduction to AI Workloads

The Role of CPUs in AI Systems

The Role of GPUs in Accelerating AI

Comparing CPU and GPU Architectures for ML

Introduction to TPUs and other ASICs

Memory and its Importance for Large Models

Storage Solutions for AI Datasets

Networking Considerations for Distributed Systems

Hands-on Practical: Benchmarking CPU vs GPU

Chapter 2: Designing On-Premise AI Infrastructure

Assessing Workload Requirements

Selecting Server Hardware for AI

GPU Interconnect Technologies

High-Speed Storage Configurations

Networking for Data and Model Transfer

Power and Cooling Requirements

Building a Bare-Metal AI Server

Practice: Creating a Hardware Specification Sheet

Chapter 3: Leveraging Cloud Platforms for AI

Overview of Major Cloud Providers for AI

Comparing Managed AI Services vs IaaS

Selecting Virtual Machine Instances for Training

Choosing Instances for Inference and Serving

Object Storage Services for Datasets

Understanding Cloud Networking and VPCs

Security Considerations in the Cloud

Hands-on Practical: Launching a GPU Cloud Instance

Chapter 4: Containerization and Orchestration for ML

Introduction to Docker for Reproducible Environments

Building a Docker Image with ML Libraries

Introduction to Kubernetes for Managing ML Workloads

Kubernetes Components: Pods, Services, Deployments

Managing GPU Resources in a Kubernetes Cluster

Using Kubeflow for ML Pipelines

Hands-on Practical: Deploying a Model on Kubernetes

Chapter 5: Strategies for Performance Optimization

Identifying Performance Bottlenecks

Techniques for Distributed Training

Using Mixed-Precision Training

Model Quantization for Efficient Inference

Optimizing Data Loading and Preprocessing Pipelines

Profiling GPU and CPU Usage

Hands-on Practical: Applying Mixed-Precision Training

Chapter 6: Cost Management and Optimization

Analyzing On-Premise Total Cost of Ownership

Understanding Cloud Pricing Models

Strategies for Reducing Cloud Compute Costs

Managing Data Storage and Transfer Costs

Implementing Cost Monitoring and Alerting

Right-Sizing Infrastructure for Workloads

Practice: Calculating and Comparing Job Costs

Choosing Instances for Inference and Serving

Was this section helpful?

References

Designing Machine Learning Systems, Chip Huyen, 2022 (O'Reilly Media) - A comprehensive guide covering the full lifecycle of machine learning systems, including architecture design for model serving, hardware selection, and optimization techniques.
NVIDIA Triton Inference Server Documentation, NVIDIA Corporation, 2023 (NVIDIA Corporation) - Official documentation for NVIDIA's open-source inference serving software, detailing how to deploy and optimize models on various hardware, including GPU batching strategies.
AWS Inferentia and AWS Neuron SDK Documentation, Amazon Web Services, 2023 - Official resources explaining AWS Inferentia accelerators, their architecture, and how to use the AWS Neuron SDK for compiling and deploying models for cost-efficient inference at scale.
Designing and deploying a machine learning prediction service, Google Cloud, 2023 (Google Cloud) - An architectural guide from Google Cloud offering considerations and best practices for building scalable and reliable machine learning prediction services, encompassing various deployment options and hardware choices.

© 2025 ApX Machine LearningEngineered with