Get in Touch
 Duration 21 hours

Course Outline

Introduction to Scaling Ollama

  • Ollama’s architecture and key scaling factors
  • Typical bottlenecks in multi-user setups
  • Best practices for preparing infrastructure

Resource Allocation and GPU Optimization

  • Strategies for maximizing CPU and GPU efficiency
  • Considerations for memory and bandwidth management
  • Applying resource constraints at the container level

Deployment with Containers and Kubernetes

  • Packaging Ollama using Docker
  • Deploying Ollama within Kubernetes clusters
  • Managing load balancing and service discovery

Autoscaling and Batching

  • Formulating autoscaling policies for Ollama
  • Using batch inference to boost throughput
  • Balancing latency against throughput requirements

Latency Optimization

  • Analyzing inference performance via profiling
  • Implementing caching and model warm-up techniques
  • Minimizing I/O and communication costs

Monitoring and Observability

  • Connecting Prometheus for metrics collection
  • Creating dashboards using Grafana
  • Setting up alerts and incident response for Ollama infrastructure

Cost Management and Scaling Strategies

  • Allocating GPUs with cost awareness
  • Evaluating cloud versus on-premises deployment options
  • Developing strategies for sustainable growth

Summary and Next Steps

Requirements

  • Hands-on experience with Linux system administration
  • Solid understanding of containerization and orchestration principles
  • Knowledge of deploying machine learning models

Audience

  • DevOps engineers
  • ML infrastructure teams
  • Site reliability engineers

Number of participants


Price per participant

Upcoming Courses

Related Categories