AI-901 Study Guide
Module 7 of 814 min read
AI-Powered Information Extraction
Learn how AI unlocks structured data from unstructured sources. This chapter covers the full OCR pipeline, field extraction and mapping techniques, and how Azure Document Intelligence and Azure Content Understanding in Foundry automate document processing at scale.
These study notes summarise Microsoft Learn material for Exam AI-901. For the official skills measured, see the Microsoft Learn study guide for Exam AI-901.
On this pageShow
- What Is Information Extraction?
- Stage 1: Optical Character Recognition (OCR)
- The OCR Pipeline
- Stage 2: Field Extraction and Mapping
- Approaches to Field Extraction
- What the Field Extraction Layer Produces
- Information Extraction in Microsoft Foundry
- Azure Document Intelligence
- Azure Content Understanding
- Choosing the Right Service
- Extracting Information from Audio and Video
- Foundry Portal Workflow for Information Extraction
- Key Takeaways for the Exam
What Is Information Extraction?
In Module 1, information extraction appeared as the fifth of the six core AI workloads. The description was brief: AI pulls structured data from unstructured sources like scanned forms, receipts, invoices, and contracts, and, as Module 1's updated overview now notes, from audio and video recordings too. This module covers how all of that works in practice, and what Azure provides to do it at scale.
Information extraction is the process of automatically pulling structured data values out of unstructured sources such as scanned documents, photographs, PDF forms, invoices, receipts, and audio or video recordings. Businesses deal with enormous volumes of unstructured documents every day: expense receipts, contracts, medical records, invoices, ID documents, and more. Processing these manually is slow and error-prone. AI-powered information extraction automates this work by reading the document, identifying the relevant fields, and outputting structured data that can flow directly into business systems.
The pipeline typically involves two stages:
- OCR - detecting and reading the text in an image or document
- Field extraction - interpreting that text and mapping it to specific data fields
These two stages are distinct problems. The first is about reading; the second is about understanding. Let's start with reading.
Stage 1: Optical Character Recognition (OCR)
Optical character recognition (OCR) is the computer vision technique that identifies and reads text from images. It is the foundation of almost every information extraction solution. Without accurate OCR, the field extraction layer has nothing reliable to work with.
The OCR Pipeline
OCR works through five steps:
1. Image Acquisition The source image enters the system. This could be a scanned document, a smartphone photo, a video frame, or a PDF rendered as an image. Image quality at this stage directly affects the accuracy of everything that follows.
2. Pre-processing and Enhancement Before text detection begins, the image is cleaned up:
- Noise reduction removes dust spots, scanner artifacts, and digital noise
- Contrast adjustment enhances the difference between text and background
- Skew correction detects and straightens documents that were scanned at an angle
- Resolution optimization adjusts the image to the optimal size for recognition models
3. Text Region Detection The system analyzes the enhanced image to find areas that contain text. It distinguishes text from images, graphics, and whitespace, groups characters into words, and groups words into lines and paragraphs.
4. Character Recognition The core step. A neural network reads the segmented text regions and predicts what character each region contains. Modern OCR uses transformer-based models trained on enormous datasets of printed and handwritten text across many languages and fonts.
5. Output Generation The recognized characters are assembled into words, sentences, and structured text. The output includes the text content along with positional information (bounding boxes for each word), confidence scores for each recognition, and often a preserved layout that maps back to the original document structure.
Good OCR gets the words right. But a document like an invoice may contain many lines of text, and only some are relevant (vendor name, date, total). That is where field extraction comes in.
Stage 2: Field Extraction and Mapping
Field extraction is the intelligence layer on top of OCR. It takes the raw text that OCR produced and figures out which pieces of text correspond to which data fields in your target schema.
Approaches to Field Extraction
Template-based detection For documents with a fixed, known layout (like a standard form), templates define where each field is located. The system looks for known labels like "Invoice Number:" or "Date:" and extracts the value next to them using pattern matching and regular expressions. This approach is fast and accurate for standardized documents but breaks down when layouts vary.
Machine learning-based detection Transformer-based models like LayoutLM combine text content, visual layout, and positional information to understand document structure without relying on fixed templates. Trained on labeled examples, these models learn to identify fields across varied layouts and formats.
Generative AI for schema-based extraction Large language models can extract fields from documents by following natural language instructions. You describe the schema you want (extract vendor name, invoice date, line items, and total), provide the OCR output as context, and the model extracts and formats the data accordingly. This works well for complex, unstructured documents where rules and templates are impractical.
What the Field Extraction Layer Produces
Regardless of approach, field extraction produces several types of output:
Key-value pairs represent individual fields: vendor_name, invoice_date, total_amount, and so on, each with their extracted value.
Table extraction identifies structured tables in documents and outputs them with their rows, columns, and cell values preserved. This is essential for invoices with line items or financial statements with multiple columns.
Confidence scores accompany each extracted value to indicate how certain the model is. Values below a threshold can be flagged for human review - this is the Reliability and Safety principle from responsible AI applied directly to production workflows. Module 8 covers all six principles with exam-style scenarios.
Data normalization ensures consistency: dates are converted to a standard format, currency symbols are stripped and values stored as numbers, and text is normalized for case and encoding. This ensures data flowing into business systems is clean and consistent.
With the two-stage pipeline clear, let's look at how Azure implements all of this.
Information Extraction in Microsoft Foundry
Microsoft Foundry and Azure provide information extraction through three main services. The right choice depends on your document type, whether you have labeled training data, and how much customization you need.
Azure Document Intelligence
Azure Document Intelligence (formerly Form Recognizer) is the specialized service for extracting data from documents and forms. It provides:
- Pre-built models for common document types: invoices, receipts, ID documents, tax forms, business cards, and more. These models are trained on millions of real-world documents and work out of the box.
- Custom models trained on your own labeled documents for proprietary forms specific to your organization.
- Layout analysis that extracts the structural elements of any document: text, tables, selection marks, and their positions, without needing a pre-built or custom model for the specific form type.
Azure Document Intelligence handles the full pipeline automatically. You send a document (image, PDF, or URL), and the service returns the extracted fields, confidence scores, and positional data in a structured JSON response.
Azure Content Understanding
Azure Content Understanding extends extraction beyond text documents to other media types:
- Audio extraction transcribes spoken content and extracts structured information from meetings, call center recordings, and interviews
- Video extraction analyzes video frames and audio simultaneously, extracting both visual and spoken content
- Multimodal extraction combines visual, audio, and text signals for complex documents that mix media types
This is what the exam means when it lists audio and video as sources for information extraction - not just scanned documents.
Choosing the Right Service
| Use Case | Recommended Service |
|---|---|
| Common document types out of the box | Azure Document Intelligence pre-built models |
| Proprietary forms specific to your organization | Azure Document Intelligence custom models |
| Processing audio recordings or video content | Azure Content Understanding |
| Building a searchable document index | Azure AI Search with cognitive skills |
| Flexible extraction with generative AI reasoning | Foundry model with schema-based prompting |
Extracting Information from Audio and Video
Azure Content Understanding is not limited to documents and images. It can process audio and video files to extract structured information - this is a distinct capability the exam tests separately from document extraction.
From audio files, Content Understanding can:
- Transcribe all spoken words to a full text transcript
- Identify different speakers and label their segments (speaker diarization)
- Extract key topics discussed throughout the recording
- Generate a structured summary of the conversation
From video files, Content Understanding can:
- Transcribe all dialogue and on-screen text (OCR on video frames)
- Identify and label speakers across the video
- Detect scene changes and chapter boundaries
- Extract visual objects and actions appearing on screen
- Generate a time-stamped summary of what happens throughout the video
This capability is useful for processing meeting recordings, customer service call logs, video tutorials, lecture recordings, and broadcast media - all without custom model training.
The key exam point: the same Content Understanding service handles documents, images, audio, and video through a unified API. The difference is only in which file type you submit and which output fields the analyzer returns.
Connecting Information Extraction from an Application
Every application that calls Azure Document Intelligence or Azure Content Understanding needs the same three ingredients:
| Ingredient | What it is | Where to find it |
|---|---|---|
| Endpoint URL | The HTTPS address of your Document Intelligence or Content Understanding resource | Foundry portal → Foundry Tools → the relevant service → Keys and Endpoint |
| API Key | A secret credential that authorises your app to call the service | Same location as the endpoint |
| SDK (code library) | A pre-built package that handles file submission and polling for results for you | Added to your code project |
The call sequence follows the same three steps: connect (set up a connection using your endpoint URL and API key), submit a document or file (pass the file along with the analyzer you want to use), then read results (because extraction takes time, your app polls - checks back - until the job is done, then reads the extracted fields and their confidence scores).
The playground's View code button shows this exact pattern with your real endpoint pre-filled.
Exam tip: Both Document Intelligence and Content Understanding process files asynchronously - meaning you submit the file, get a job ID back immediately, and then check back (poll) until the job is done before reading the results. Think of it like dropping clothes at a dry cleaner: you get a ticket, come back later, and collect the finished items. This submit → poll → retrieve sequence is directly exam-tested.
Foundry Portal Workflow for Information Extraction
Azure AI Content Understanding is available within Microsoft Foundry. It offers prebuilt analyzers for common document types (invoices, receipts, ID documents, contracts) as well as custom analyzers you train on your own document schemas. The same service handles documents, images, audio files, and video through a unified API.
Extracted fields and their confidence scores are visible in the results panel before you write any code.
For the exam, the key architectural fact is that Content Understanding always uses an asynchronous API pattern: you submit the file, receive an Operation-Location URL in the response header, poll that URL until the status equals succeeded, then read the result. This two-step submit-then-poll pattern applies identically whether you are processing a PDF invoice or a video recording.
📌 Note: The Azure Content Understanding workspace is accessed through the classic Foundry interface (New Foundry toggle disabled), not the New Foundry interface used for agents and generative AI.
🔒 Hands-On Lab - Lab 6: Automate Meridian's Document Processing Build an async invoice processing pipeline using Content Understanding, extract supplier fields from real PDFs, and process audio files with speaker diarization - all in a live Azure environment. Already enrolled? Labs are launching shortly - stay tuned!. New here? Get the AI-901 Lab Bundle →
Key Takeaways for the Exam
- Information extraction converts unstructured content (documents, images, audio, video) into structured, queryable data - without manual data entry.
- OCR is the foundation: it recognises text in images and scanned documents. Modern OCR handles handwriting, tables, and multi-column layouts.
- Azure Document Intelligence offers prebuilt models for invoices, receipts, ID documents, contracts, and tax forms. Custom models can be trained on your own document types.
- Azure Content Understanding is the Foundry-native service. It unifies document, image, audio, and video extraction through a single API.
- Audio extraction: transcription, speaker diarization (who said what), topic detection, conversation summary.
- Video extraction: dialogue transcription, speaker labeling, scene/chapter detection, on-screen text OCR, visual object detection, time-stamped summary.
- Foundry portal flow: Azure Content Understanding is accessed via the classic Foundry interface. It offers prebuilt and custom analyzers for documents, images, audio, and video through a unified API.
- API pattern (always asynchronous): POST file to analyze endpoint → get Operation-Location header → GET that URL repeatedly → when status equals "succeeded" → read extracted fields.
- The same two-step submit-then-poll pattern applies to documents, images, audio files, and video files.
- Confidence scores appear on every extracted field. Low-confidence fields should trigger human review in production workflows.
Official exam information from Microsoft
- Study guide for Exam AI-901: Microsoft Azure AI Fundamentals (skills measured, weights and passing score), and the Microsoft Certified: Azure AI Fundamentals certification page.