Get in Touch

Course Outline

Introduction to Multimodal LLMs in Vertex AI

  • Overview of multimodal features within Vertex AI
  • Introduction to Gemini models and their supported modalities
  • Enterprise and research use cases

Setting Up the Development Environment

  • Configuring Vertex AI for multimodal processing
  • Managing datasets across different modalities
  • Hands-on lab: environment configuration and dataset preparation

Long Context Windows and Advanced Reasoning

  • Exploring long-context workflow capabilities
  • Applications in planning and decision-making processes
  • Hands-on lab: executing long-context analysis

Cross-Modal Workflow Design

  • Synthesizing text, audio, and image analysis
  • Linking multimodal steps within a pipeline
  • Hands-on lab: constructing a multimodal pipeline

Working with Gemini API Parameters

  • Setting up multimodal inputs and outputs
  • Enhancing inference speed and resource efficiency
  • Hands-on lab: adjusting Gemini API parameters

Advanced Applications and Integrations

  • Creating interactive multimodal agents and assistants
  • Incorporating external APIs and tools
  • Hands-on lab: developing a comprehensive multimodal application

Evaluation and Iteration

  • Assessing multimodal performance
  • Key metrics for accuracy, alignment, and drift detection
  • Hands-on lab: evaluating multimodal workflow effectiveness

Summary and Next Steps

Requirements

  • Strong proficiency in Python programming
  • Background experience in developing machine learning models
  • Working knowledge of multimodal data types, including text, audio, and images

Target Audience

  • AI researchers
  • Senior developers
  • Machine learning scientists
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories