Get in Touch

Course Outline

Foundations of Agentic Systems in Production

  • Agentic architecture components: loops, tools, memory, and orchestration layers
  • The agent lifecycle: from development and deployment to continuous operation
  • Key challenges in managing agents at production scale

Infrastructure and Deployment Strategies

  • Deploying agents within containerized and cloud-based environments
  • Scaling methodologies: horizontal versus vertical scaling, concurrency, and throttling
  • Orchestration of multi-agent systems and workload distribution

Monitoring and Observability

  • Essential metrics: latency, success rates, memory consumption, and agent call depth
  • Tracing agent activities and visualizing call graphs
  • Implementing observability through Prometheus, OpenTelemetry, and Grafana

Logging, Auditing, and Compliance

  • Centralized logging and structured event aggregation
  • Ensuring compliance and auditability within agentic workflows
  • Creating audit trails and replay mechanisms to facilitate debugging

Performance Tuning and Resource Optimization

  • Minimizing inference overhead and refining agent orchestration cycles
  • Utilizing model caching and lightweight embeddings for enhanced retrieval speeds
  • Conducting load testing and stress scenario analysis for AI pipelines

Cost Governance and Control

  • Analyzing cost drivers for agents: API calls, memory, compute resources, and external integrations
  • Monitoring agent-specific expenses and applying chargeback models
  • Implementing automated policies to curb agent sprawl and eliminate idle resource usage

CI/CD and Rollout Tactics for Agents

  • Embedding agent pipelines into CI/CD ecosystems
  • Strategies for testing, versioning, and rolling back iterative agent updates
  • Executing progressive rollouts and secure deployment mechanisms

Failure Recovery and Reliability Engineering

  • Architecting for fault tolerance and graceful degradation
  • Applying retry, timeout, and circuit breaker patterns to ensure agent reliability
  • Establishing incident response and post-mortem frameworks for AI operations

Capstone Project

  • Constructing and deploying an agentic AI system with comprehensive monitoring and cost tracking
  • Simulating load, assessing performance, and optimizing resource consumption
  • Presenting the final architecture and monitoring dashboard to peers

Recap and Future Directions

Requirements

  • A robust grasp of MLOps and production machine learning systems
  • Hands-on experience with containerized deployments (Docker/Kubernetes)
  • Proficiency with cloud cost optimization and observability tools

Target Audience

  • MLOps Engineers
  • Site Reliability Engineers (SREs)
  • Engineering Managers overseeing AI infrastructure
 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories