Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Multimodal LLMs in Vertex AI
- Overview of multimodal capabilities within Vertex AI
- Gemini models and their supported data modalities
- Enterprise and research use cases
Establishing the Development Environment
- Configuring Vertex AI for multimodal operations
- Handling datasets across different modalities
- Practical lab: setting up the environment and preparing datasets
Long Context Windows and Advanced Reasoning
- Comprehending long-context workflows
- Applications in planning and decision-making processes
- Practical lab: implementing long-context analysis techniques
Designing Cross-Modal Workflows
- Synthesizing text, audio, and image analysis
- Chaining multimodal steps within automated pipelines
- Practical lab: architecting a multimodal pipeline
Managing Gemini API Parameters
- Configuring multimodal inputs and outputs
- Enhancing inference speed and operational efficiency
- Practical lab: fine-tuning Gemini API parameters
Advanced Applications and Integrations
- Developing interactive multimodal agents and assistants
- Integrating with external APIs and tools
- Practical lab: constructing a complete multimodal application
Evaluation and Iteration
- Assessing multimodal performance metrics
- Tracking accuracy, alignment, and drift indicators
- Practical lab: evaluating the effectiveness of multimodal workflows
Recap and Future Directions
Requirements
- Solid proficiency in Python programming
- Practical experience in developing machine learning models
- Working familiarity with multimodal data types, including text, audio, and images
Target Audience
- AI researchers
- Senior developers
- Machine learning scientists
14 Hours