AI-901 Study Guide
Module 6 of 812 min read
Computer Vision
Learn how AI interprets visual information from images and video. This chapter covers the four core CV tasks (image classification, object detection, semantic segmentation, and multimodal analysis), how CNN and Vision Transformer architectures learn from data, diffusion-based image and video generation, and Azure AI Vision capabilities in Microsoft Foundry.
These study notes summarise Microsoft Learn material for Exam AI-901. For the official skills measured, see the Microsoft Learn study guide for Exam AI-901.
On this pageShow
- What Is Computer Vision?
- How Images Are Represented
- Core Computer Vision Tasks
- Image Classification
- Object Detection
- Semantic Segmentation
- Contextual Image Analysis (Multimodal)
- How Computer Vision Models Learn
- Convolutional Neural Networks (CNNs)
- Vision Transformers (ViTs)
- Multimodal Models
- Image Generation with Diffusion
- Computer Vision in Microsoft Foundry
- Multimodal Models for Image Analysis
- Image Generation Models
- Video Generation Models
- Working with Vision in Foundry
- Foundry Portal Workflow for Computer Vision
- Key Takeaways for the Exam
What Is Computer Vision?
Computer vision is the area of AI that enables software to interpret and understand visual input: photographs, videos, and live camera feeds. To a computer, an image is simply a grid of pixel values. Computer vision models learn to extract meaning from those grids by training on large volumes of labeled images.
In Module 1 you saw three examples: a manufacturing quality control system (image classification), self-driving cars identifying pedestrians and traffic lights (object detection), and a product designer generating photorealistic sneaker mockups from a text description (diffusion). This module covers how all three work in depth.
How Images Are Represented
Every digital image is a multi-dimensional array of numbers. A grayscale image is a 2D array where each value between 0 (black) and 255 (white) represents one pixel. A color image adds a third dimension: three channels (Red, Green, Blue), each a 2D array of values. The combination of RGB values at each pixel position determines the color that appears.
This numerical representation is what makes images processable by machine learning models. The model learns which patterns of pixel values correspond to which objects or scenes.
Core Computer Vision Tasks
Image Classification
Image classification is the task of predicting a single label for an entire image. A model trained on labeled examples learns to identify the main subject of an image. For example, a checkout system trained on images of fruit can identify whether a customer placed an apple, banana, or orange on the scale. The model outputs a probability for each possible class; the class with the highest probability is the prediction.
Object Detection
Object detection goes further by locating and labeling multiple objects within a single image. Instead of one label for the whole image, the model outputs a set of bounding boxes, each with a class label and a confidence score. A retail self-checkout camera using object detection could identify several items placed on the belt simultaneously.
Semantic Segmentation
Semantic segmentation provides the most precise localization. Rather than drawing a box around each object, the model classifies every individual pixel in the image according to which object it belongs to. This produces a detailed "map" of the scene and is used in applications like medical imaging, autonomous vehicles, and robotics.
Contextual Image Analysis (Multimodal)
The most capable modern models combine visual and language understanding in a single multimodal architecture. These models can generate natural language descriptions of images, answer questions about visual content, and understand complex scenes. For example, a multimodal model shown a photo of a park can respond: "A person sitting on a bench reading a book, with trees and a fountain in the background."
How Computer Vision Models Learn
Convolutional Neural Networks (CNNs)
For many years, convolutional neural networks (CNNs) were the dominant architecture for computer vision. A CNN works by applying filters (also called kernels) across an image to extract features. A filter is a small grid of weights (for example, 3x3) that slides across the image. At each position, it multiplies its weights by the corresponding pixel values and sums them up to produce a single output value.
Different filters detect different features: one filter might detect horizontal edges, another vertical edges, another specific color patterns. The process of sliding a filter across an image is called convolution. Stacking multiple convolutional layers lets the network learn increasingly abstract features: early layers detect edges and textures, middle layers detect shapes and parts, and later layers recognize entire objects.
The training process adjusts the filter weights iteratively until the model's predictions match the known labels. These learned weights encode what makes each object class visually distinctive.
Vision Transformers (ViTs)
More recently, the transformer architecture that powers LLMs has been adapted for images. A vision transformer (ViT) divides an image into a grid of fixed-size patches (for example, 16x16 pixel tiles), converts each patch into a linear vector, and processes all patches through a transformer with attention.
Attention allows the model to learn which patches are contextually related to which others. Just as transformer language models learn that "bark" and "dog" are related, a vision transformer learns that patches containing "hat" shapes tend to appear near patches containing "head" shapes.
Multimodal Models
When a model is trained on both images with associated text descriptions, the vision encoder and language encoder can be combined into a multimodal model. A technique called cross-modal attention creates a shared embedding space where visual features and language concepts are aligned. This alignment is what allows a multimodal model to:
- Generate captions for unseen images
- Answer questions about visual content
- Search for images using natural language queries
Image Generation with Diffusion
Modern AI can also generate images from text descriptions, applying the same generative AI principles covered in Module 3, but adapted for visual output rather than text. The dominant approach is called diffusion. A diffusion model starts with a completely random image (pure noise) and iteratively removes noise in a direction guided by the text prompt. At each step, the model asks: "Given this prompt and the image so far, what should the next, slightly less noisy version look like?" After many iterations, a coherent image matching the description emerges.
The same diffusion principle is extended to video generation, where the model additionally accounts for physical plausibility (objects move consistently with physics) and temporal progression (the sequence of frames tells a logical story).
Computer Vision in Microsoft Foundry
Azure provides vision capabilities through Foundry Tools as part of Microsoft Foundry. Three key services cover the main computer vision scenarios:
Multimodal Models for Image Analysis
Foundry's model catalog includes multimodal models (such as GPT-4.1 with vision) that can analyze images, generate descriptions, and answer questions about visual content. You interact with these models through the same prompt-based interface used for text - simply including the image in the request via the Chat playground.
Image Generation Models
Foundry provides access to image generation models (such as DALL-E and others in the catalog) that create images from text prompts. These can be used for content creation, design prototyping, and generating training data for other vision models.
Video Generation Models
Foundry also provides video generation capabilities, enabling the creation of short video clips from text descriptions. This extends the diffusion-based approach from static images to temporal sequences. For processing and extracting information from existing videos, see Module 7 which covers Azure Content Understanding's video extraction capabilities.
Working with Vision in Foundry
To use vision capabilities in a project:
- Choose a multimodal or image generation model from the Foundry catalog and deploy it
- Test it in the Playground by uploading an image and asking a question about it
- For specialized tasks like reading text from images or detecting specific objects, use Azure AI Vision from Foundry Tools, which provides pre-built models for OCR, object detection, and image analysis
Azure AI Vision also provides face detection (locating faces in an image) and face verification (confirming whether two images show the same person), useful for identity-based applications like access control or photo organization.
Connecting Computer Vision from an Application
Every application that calls Azure AI Vision or a deployed multimodal model needs the same three ingredients:
| Ingredient | What it is | Where to find it |
|---|---|---|
| Endpoint URL | The HTTPS address of your Azure AI Vision resource or deployed model | Foundry portal โ Foundry Tools โ Azure AI Vision โ Keys and Endpoint (or Deployments for multimodal models) |
| API Key | A secret credential that authorises your app to call the service | Same location as the endpoint |
| SDK (code library) | A pre-built package that handles image encoding and communication details for you | Added to your code project |
The call sequence always follows the same three steps: connect (set up a connection using your endpoint URL and API key), send an image (pass the image as a URL or file), then read results (the service returns structured data - labels, bounding boxes, captions, or extracted text - that your app can use).
The Vision playground's View code button shows this exact pattern with your real endpoint pre-filled.
Exam tip: Azure AI Vision covers image analysis, dense captioning, OCR, and face detection - all through one endpoint. For multimodal vision (asking questions about images in natural language), you use a deployed multimodal model (such as GPT-4.1-mini or GPT-4o) through the Chat playground.
Foundry Portal Workflow for Computer Vision
In Microsoft Foundry, vision tasks use models deployed from the model catalog. Multimodal models (such as gpt-4.1-mini) handle image understanding via the Chat playground. Text-to-image models (such as DALL-E or FLUX) generate new images from text prompts. Video generation models (such as Sora) create short video clips from text descriptions. These are three separate model types with distinct use cases.
๐ Note: Computer vision tasks in the current Foundry portal use the New Foundry interface. The classic Foundry interface is used for other services such as Language and Content Understanding.
๐ Hands-On Lab - Lab 5: AI Eyes for Meridian - Computer Vision Use GPT-4o to analyse real property inspection photos and generate urgency reports, then use DALL-E to create hotel marketing images - all in a live Azure environment. Already enrolled? Labs are launching shortly - stay tuned!. New here? Get the AI-901 Lab Bundle โ
Key Takeaways for the Exam
- Image classification: returns a single label for the whole image (what is it?).
- Object detection: returns labels + bounding box coordinates for each object (what and where?).
- Semantic segmentation: classifies every pixel individually (which pixel belongs to which object?).
- Multimodal analysis: models trained on both images and text, can describe, answer questions about, or classify image content.
- CNNs process images with sliding filters that detect features layer by layer (edges โ shapes โ objects). Architectures include ResNet, EfficientNet, YOLO.
- Vision Transformers (ViTs) divide images into patches, treat each as a token, and apply attention across patches. Power models like CLIP and Florence.
- Diffusion models generate images by learning to reverse a noise-adding process. DALL-E 3 and GPT Image use this approach.
- Azure AI Vision handles: image analysis (tags, captions), dense captioning, OCR, spatial analysis, face detection/verification.
- Foundry portal flow (New Foundry): Deploy a multimodal model (e.g. gpt-4.1-mini) โ Chat playground โ attach image โ ask question. For generation: deploy a text-to-image model (e.g. DALL-E, FLUX) โ image generation playground. For video: deploy a video generation model (e.g. Sora) โ video playground.
- Key distinction: Multimodal models understand and describe existing images; image generation models create new images from text; video generation models create short video clips.
Official exam information from Microsoft
- Study guide for Exam AI-901: Microsoft Azure AI Fundamentals (skills measured, weights and passing score), and the Microsoft Certified: Azure AI Fundamentals certification page.