Skip to content
GetHandsOn.ai

AI-901 Study Guide


Module 7 of 814 min read

AI-Powered Information Extraction

Learn how AI unlocks structured data from unstructured sources. This chapter covers the full OCR pipeline, field extraction and mapping techniques, and how Azure Document Intelligence and Azure Content Understanding in Foundry automate document processing at scale.

These study notes summarise Microsoft Learn material for Exam AI-901. For the official skills measured, see the Microsoft Learn study guide for Exam AI-901.

On this pageShow

What Is Information Extraction?

In Module 1, information extraction appeared as the fifth of the six core AI workloads. The description was brief: AI pulls structured data from unstructured sources like scanned forms, receipts, invoices, and contracts, and, as Module 1's updated overview now notes, from audio and video recordings too. This module covers how all of that works in practice, and what Azure provides to do it at scale.

Information extraction is the process of automatically pulling structured data values out of unstructured sources such as scanned documents, photographs, PDF forms, invoices, receipts, and audio or video recordings. Businesses deal with enormous volumes of unstructured documents every day: expense receipts, contracts, medical records, invoices, ID documents, and more. Processing these manually is slow and error-prone. AI-powered information extraction automates this work by reading the document, identifying the relevant fields, and outputting structured data that can flow directly into business systems.

The pipeline typically involves two stages:

  1. OCR - detecting and reading the text in an image or document
  2. Field extraction - interpreting that text and mapping it to specific data fields

These two stages are distinct problems. The first is about reading; the second is about understanding. Let's start with reading.


Stage 1: Optical Character Recognition (OCR)

Optical character recognition (OCR) is the computer vision technique that identifies and reads text from images. It is the foundation of almost every information extraction solution. Without accurate OCR, the field extraction layer has nothing reliable to work with.

The OCR Pipeline

OCR works through five steps:

1. Image Acquisition The source image enters the system. This could be a scanned document, a smartphone photo, a video frame, or a PDF rendered as an image. Image quality at this stage directly affects the accuracy of everything that follows.

2. Pre-processing and Enhancement Before text detection begins, the image is cleaned up:

  • Noise reduction removes dust spots, scanner artifacts, and digital noise
  • Contrast adjustment enhances the difference between text and background
  • Skew correction detects and straightens documents that were scanned at an angle
  • Resolution optimization adjusts the image to the optimal size for recognition models

3. Text Region Detection The system analyzes the enhanced image to find areas that contain text. It distinguishes text from images, graphics, and whitespace, groups characters into words, and groups words into lines and paragraphs.

4. Character Recognition The core step. A neural network reads the segmented text regions and predicts what character each region contains. Modern OCR uses transformer-based models trained on enormous datasets of printed and handwritten text across many languages and fonts.

5. Output Generation The recognized characters are assembled into words, sentences, and structured text. The output includes the text content along with positional information (bounding boxes for each word), confidence scores for each recognition, and often a preserved layout that maps back to the original document structure.

Good OCR gets the words right. But a document like an invoice may contain many lines of text, and only some are relevant (vendor name, date, total). That is where field extraction comes in.


Stage 2: Field Extraction and Mapping

Field extraction is the intelligence layer on top of OCR. It takes the raw text that OCR produced and figures out which pieces of text correspond to which data fields in your target schema.

Approaches to Field Extraction

Template-based detection For documents with a fixed, known layout (like a standard form), templates define where each field is located. The system looks for known labels like "Invoice Number:" or "Date:" and extracts the value next to them using pattern matching and regular expressions. This approach is fast and accurate for standardized documents but breaks down when layouts vary.

Machine learning-based detection Transformer-based models like LayoutLM combine text content, visual layout, and positional information to understand document structure without relying on fixed templates. Trained on labeled examples, these models learn to identify fields across varied layouts and formats.

Generative AI for schema-based extraction Large language models can extract fields from documents by following natural language instructions. You describe the schema you want (extract vendor name, invoice date, line items, and total), provide the OCR output as context, and the model extracts and formats the data accordingly. This works well for complex, unstructured documents where rules and templates are impractical.

What the Field Extraction Layer Produces

Regardless of approach, field extraction produces several types of output:

Key-value pairs represent individual fields: vendor_name, invoice_date, total_amount, and so on, each with their extracted value.

Table extraction identifies structured tables in documents and outputs them with their rows, columns, and cell values preserved. This is essential for invoices with line items or financial statements with multiple columns.

Confidence scores accompany each extracted value to indicate how certain the model is. Values below a threshold can be flagged for human review - this is the Reliability and Safety principle from responsible AI applied directly to production workflows. Module 8 covers all six principles with exam-style scenarios.

Data normalization ensures consistency: dates are converted to a standard format, currency symbols are stripped and values stored as numbers, and text is normalized for case and encoding. This ensures data flowing into business systems is clean and consistent.

With the two-stage pipeline clear, let's look at how Azure implements all of this.


Information Extraction in Microsoft Foundry

Microsoft Foundry and Azure provide information extraction through three main services. The right choice depends on your document type, whether you have labeled training data, and how much customization you need.

Azure Document Intelligence

Azure Document Intelligence (formerly Form Recognizer) is the specialized service for extracting data from documents and forms. It provides:

  • Pre-built models for common document types: invoices, receipts, ID documents, tax forms, business cards, and more. These models are trained on millions of real-world documents and work out of the box.
  • Custom models trained on your own labeled documents for proprietary forms specific to your organization.
  • Layout analysis that extracts the structural elements of any document: text, tables, selection marks, and their positions, without needing a pre-built or custom model for the specific form type.

Azure Document Intelligence handles the full pipeline automatically. You send a document (image, PDF, or URL), and the service returns the extracted fields, confidence scores, and positional data in a structured JSON response.

Azure Content Understanding

Azure Content Understanding extends extraction beyond text documents to other media types:

  • Audio extraction transcribes spoken content and extracts structured information from meetings, call center recordings, and interviews
  • Video extraction analyzes video frames and audio simultaneously, extracting both visual and spoken content
  • Multimodal extraction combines visual, audio, and text signals for complex documents that mix media types

This is what the exam means when it lists audio and video as sources for information extraction - not just scanned documents.

Choosing the Right Service

Use CaseRecommended Service
Common document types out of the boxAzure Document Intelligence pre-built models
Proprietary forms specific to your organizationAzure Document Intelligence custom models
Processing audio recordings or video contentAzure Content Understanding
Building a searchable document indexAzure AI Search with cognitive skills
Flexible extraction with generative AI reasoningFoundry model with schema-based prompting

Extracting Information from Audio and Video

Azure Content Understanding is not limited to documents and images. It can process audio and video files to extract structured information - this is a distinct capability the exam tests separately from document extraction.

From audio files, Content Understanding can:

  • Transcribe all spoken words to a full text transcript
  • Identify different speakers and label their segments (speaker diarization)
  • Extract key topics discussed throughout the recording
  • Generate a structured summary of the conversation

From video files, Content Understanding can:

  • Transcribe all dialogue and on-screen text (OCR on video frames)
  • Identify and label speakers across the video
  • Detect scene changes and chapter boundaries
  • Extract visual objects and actions appearing on screen
  • Generate a time-stamped summary of what happens throughout the video

This capability is useful for processing meeting recordings, customer service call logs, video tutorials, lecture recordings, and broadcast media - all without custom model training.

The key exam point: the same Content Understanding service handles documents, images, audio, and video through a unified API. The difference is only in which file type you submit and which output fields the analyzer returns.

Connecting Information Extraction from an Application

Every application that calls Azure Document Intelligence or Azure Content Understanding needs the same three ingredients:

IngredientWhat it isWhere to find it
Endpoint URLThe HTTPS address of your Document Intelligence or Content Understanding resourceFoundry portal → Foundry Tools → the relevant service → Keys and Endpoint
API KeyA secret credential that authorises your app to call the serviceSame location as the endpoint
SDK (code library)A pre-built package that handles file submission and polling for results for youAdded to your code project

The call sequence follows the same three steps: connect (set up a connection using your endpoint URL and API key), submit a document or file (pass the file along with the analyzer you want to use), then read results (because extraction takes time, your app polls - checks back - until the job is done, then reads the extracted fields and their confidence scores).

The playground's View code button shows this exact pattern with your real endpoint pre-filled.

Exam tip: Both Document Intelligence and Content Understanding process files asynchronously - meaning you submit the file, get a job ID back immediately, and then check back (poll) until the job is done before reading the results. Think of it like dropping clothes at a dry cleaner: you get a ticket, come back later, and collect the finished items. This submit → poll → retrieve sequence is directly exam-tested.


Foundry Portal Workflow for Information Extraction

Azure AI Content Understanding is available within Microsoft Foundry. It offers prebuilt analyzers for common document types (invoices, receipts, ID documents, contracts) as well as custom analyzers you train on your own document schemas. The same service handles documents, images, audio files, and video through a unified API.

Extracted fields and their confidence scores are visible in the results panel before you write any code.

For the exam, the key architectural fact is that Content Understanding always uses an asynchronous API pattern: you submit the file, receive an Operation-Location URL in the response header, poll that URL until the status equals succeeded, then read the result. This two-step submit-then-poll pattern applies identically whether you are processing a PDF invoice or a video recording.

📌 Note: The Azure Content Understanding workspace is accessed through the classic Foundry interface (New Foundry toggle disabled), not the New Foundry interface used for agents and generative AI.

🔒 Hands-On Lab - Lab 6: Automate Meridian's Document Processing Build an async invoice processing pipeline using Content Understanding, extract supplier fields from real PDFs, and process audio files with speaker diarization - all in a live Azure environment. Already enrolled? Labs are launching shortly - stay tuned!. New here? Get the AI-901 Lab Bundle →


Key Takeaways for the Exam

  • Information extraction converts unstructured content (documents, images, audio, video) into structured, queryable data - without manual data entry.
  • OCR is the foundation: it recognises text in images and scanned documents. Modern OCR handles handwriting, tables, and multi-column layouts.
  • Azure Document Intelligence offers prebuilt models for invoices, receipts, ID documents, contracts, and tax forms. Custom models can be trained on your own document types.
  • Azure Content Understanding is the Foundry-native service. It unifies document, image, audio, and video extraction through a single API.
  • Audio extraction: transcription, speaker diarization (who said what), topic detection, conversation summary.
  • Video extraction: dialogue transcription, speaker labeling, scene/chapter detection, on-screen text OCR, visual object detection, time-stamped summary.
  • Foundry portal flow: Azure Content Understanding is accessed via the classic Foundry interface. It offers prebuilt and custom analyzers for documents, images, audio, and video through a unified API.
  • API pattern (always asynchronous): POST file to analyze endpoint → get Operation-Location header → GET that URL repeatedly → when status equals "succeeded" → read extracted fields.
  • The same two-step submit-then-poll pattern applies to documents, images, audio files, and video files.
  • Confidence scores appear on every extracted field. Low-confidence fields should trigger human review in production workflows.

Official exam information from Microsoft

Get the full AI-901 study guide as a PDF, freeAll 8 modules in one printable file. Enter your email on the guide page and it is yours.

Keep going

Get the full AI-901 guide as a PDFEvery module in one file. Free after you enter your email.
Practice AI-901 questionsExam-style questions with explanations, free to start.
Hands-on AI-901 labsApply this in a real Azure environment.
AI-901 guide overviewAll modules, pick what to read next.