Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI for Operations
- Transitioning from static runbooks to reasoning agents: the evolution of IT automation
- Agent architecture: reasoning loops, tool utilization, memory, and planning
- Determining when to automate versus when to maintain human involvement
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: supervisor, hierarchical, and swarm models
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agents
- Constructing your first operational agent: querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Linking agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Agent-based log querying: integration with Elasticsearch, Loki, and Splunk
- Utilizing infrastructure tools via agent actions: kubectl, Terraform, and Ansible
- Designing secure tool interfaces through parameter validation and idempotency
Incident Response Automation
- Automated incident triage: severity classification and routing
- Generating root cause hypotheses and collecting evidence
- Automated remediation: executing restarts, scaling, rollbacks, and failovers
- Creating an incident runbook agent with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Classifying actions: read-only, low-risk, high-risk, and destructive
- Establishing approval gates and escalation policies for critical operations
- Guardrail strategies: action allowlists, blast radius limits, and rollback guarantees
- Maintaining audit trails and decision provenance for compliance
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialized agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose contradictory actions
- Simulating major incidents end-to-end with multi-agent response
Observability and Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Assessing agent decision quality: precision, recall, and time-to-resolution
- Implementing feedback loops: learning from operator overrides and outcomes
- Tracking costs and managing token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: APIs, webhooks, and scheduled jobs
- Implementing gradual autonomy rollout: from shadow mode to full auto-remediation
- Establishing runbooks for agent failures: protocols for when the agent itself encounters issues
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE practices.
- Proficiency in Python scripting and REST APIs.
- Fundamental understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers developing self-healing infrastructure.
- IT operations leaders evaluating agentic AI for incident management.
14 Hours