AI-901 Study Guide
Module 4 of 813 min read
Natural Language Processing and Text Analysis
Explore how AI makes sense of human language. This chapter covers tokenization, how text is statistically analyzed, semantic language models, and how Azure Language in Microsoft Foundry enables practical text analysis tasks like sentiment analysis, entity extraction, and summarization.
These study notes summarise Microsoft Learn material for Exam AI-901. For the official skills measured, see the Microsoft Learn study guide for Exam AI-901.
On this pageShow
- What Is Natural Language Processing?
- From Words to Numbers: Tokenization and Preprocessing
- Pre-processing Techniques
- Statistical Text Analysis
- Frequency Analysis
- TF-IDF (Term Frequency - Inverse Document Frequency)
- Bag-of-Words
- Semantic Language Models
- Word Embeddings
- Contextualized Embeddings
- Text Analysis Capabilities in Azure
- Language Detection
- Sentiment Analysis
- Named Entity Recognition (NER)
- PII Detection
- Text Summarization
- Text Classification
- Using Azure Language in Foundry
- Foundry Portal Workflow for Text Analysis
- Key Takeaways for the Exam
What Is Natural Language Processing?
Natural language processing (NLP) is the branch of AI that deals with understanding, analyzing, and generating human language. It is the foundational technology behind generative AI, search engines, chatbots, and a wide range of text analysis tools.
In Module 1 you saw three quick NLP examples: Amazon automatically tagging customer reviews as positive or negative (sentiment analysis), a flight chatbot extracting "London" and "next Tuesday" from a user message (named entity recognition), and a hospital redacting patient names from documents before sharing them (PII detection). This module covers how each of those works under the hood.
NLP is not one single technique. It is a collection of methods, ranging from simple statistical counting to deep learning, that together allow software to derive meaning from text. Before an AI system can analyze text, it needs to convert raw language into a form it can work with. That process starts with tokenization and preprocessing, the same foundational step you saw applied to LLMs in Module 3, now examined from the classical NLP perspective.
From Words to Numbers: Tokenization and Preprocessing
Computers cannot process raw text directly. The first step in any NLP pipeline is tokenization: breaking a body of text into smaller units called tokens. A token is typically a word, but it can also be a sub-word fragment, a punctuation mark, or a number. For example, the phrase "We choose to go to the moon" becomes seven tokens: We, choose, to, go, to, the, moon. Each token is assigned a unique numeric ID so the model can work with numbers rather than strings.
Pre-processing Techniques
Before tokenizing, text is usually cleaned up so the model sees consistent, meaningful input:
| Technique | What it does | Example |
|---|---|---|
| Normalization | Converts to lowercase, removes punctuation, standardizes whitespace | "Hello, World!" becomes "hello world" |
| Stop word removal | Removes common words that carry little meaning | "the", "is", "and" are removed |
| Stemming | Strips words to their root form, even if the result is not a real word | "running", "runs", "runner" all become "run" |
| Lemmatization | Reduces words to their dictionary base form | "better" becomes "good", "ran" becomes "run" |
| N-grams | Creates multi-word tokens to preserve some context | "New York" becomes a single token rather than two separate words |
Once text is cleaned and tokenized, the next question is: how do you extract meaning from it? The simplest approach is statistics.
Statistical Text Analysis
Statistical methods analyze word frequencies and patterns across documents. They do not require deep learning and are still widely used in search, filtering, and classification tasks.
Frequency Analysis
The simplest approach is to count how often each token appears. Frequently occurring terms generally indicate the main topics of a document. For example, a document where "AI", "model", and "training" appear most often is likely about machine learning.
TF-IDF (Term Frequency - Inverse Document Frequency)
When you have multiple documents, simple frequency counts fall short because common words like "cloud" or "model" appear in every Azure-related document. TF-IDF solves this by scoring words based on how often they appear in a specific document relative to how often they appear across all documents. A high TF-IDF score means a word is frequent in that document but rare elsewhere, making it a strong signal for what that document is specifically about. This technique is widely used in search engines and document classification.
Bag-of-Words
Bag-of-words represents each document as a vector of word frequencies, ignoring grammar and word order. Despite its simplicity, it is effective for classification tasks where the presence of certain words is more important than their order.
Statistical methods work well for many tasks, but they treat every occurrence of a word the same regardless of context. The word "bank" in "river bank" and "bank account" would be handled identically. Semantic language models fix that problem.
Semantic Language Models
Semantic models capture the meaning of words, not just their frequency. They do this by representing words as vectors in a high-dimensional space.
Word Embeddings
A word embedding is a numeric vector that represents a word's meaning. Words with similar meanings are placed close together in vector space. For example, "king" and "queen" have similar vectors, as do "Paris" and "London". This spatial relationship enables vector arithmetic. The classic example: the vector for "king" minus "man" plus "woman" gives a vector very close to "queen". The model has learned semantic relationships from patterns in training data.
Contextualized Embeddings
Early embedding models like Word2Vec assigned a single fixed vector to each word regardless of context. The word "bank" had the same embedding whether it meant a financial institution or a river bank. Modern transformer-based models produce contextualized embeddings: the vector for each word changes based on the words around it. This is how GPT and similar models understand that "bank" in "I deposited money at the bank" means something different from "bank" in "I sat on the bank of the river."
Contextualized embeddings are the foundation of modern NLP capabilities including text summarization, named entity recognition, and text classification.
These techniques, from statistical counting to contextual embeddings, are what power the pre-built text analysis services Microsoft offers through Azure.
Text Analysis Capabilities in Azure
Azure provides text analysis through Azure Language, available as part of the Foundry Tools suite in Microsoft Foundry. These are production-ready services you can call via API without building or training models yourself.
Language Detection
Identifies the language a piece of text is written in. Useful when your application receives content from global users and needs to route or process it differently by language.
Sentiment Analysis
Classifies text as positive, negative, or neutral, with a confidence score. Azure Language can also identify sentiment at the sentence level, not just for the whole document. Common applications include customer feedback analysis, social media monitoring, and support ticket prioritization.
Named Entity Recognition (NER)
Identifies and classifies specific entities within text: people, organizations, locations, dates, phone numbers, email addresses, and more. For example, "Satya Nadella joined Microsoft in 1992" would extract "Satya Nadella" as a person, "Microsoft" as an organization, and "1992" as a date.
PII Detection
Identifies personally identifiable information (names, addresses, email addresses, national IDs) so it can be redacted before text is stored or shared. This is a practical tool for meeting data privacy requirements like GDPR.
Text Summarization
Generates a condensed version of a document, either by extracting the most representative sentences (extractive summarization) or by generating new language that captures the key themes (abstractive summarization). Abstractive summarization is powered by generative AI under the hood, the same generative AI principles covered in Module 3.
Text Classification
Assigns a document to one or more pre-defined categories. Custom classification lets you train a model on your own labeled data to categorize documents according to your specific taxonomy.
You do not need to implement any of these techniques from scratch. In Foundry, they are available as ready-to-use services.
Using Azure Language in Foundry
In Microsoft Foundry, Azure Language is accessible through Foundry Tools. You can:
- Use it directly via the Foundry portal to analyze sample text interactively in the Language playground
- Connect it as a tool in an AI agent, so the agent can perform text analysis as part of a broader workflow
A lightweight pattern for NLP-powered applications is to combine Azure Language with a generative model: use Azure Language for structured extraction tasks like NER and PII detection, then pass the cleaned, structured output to a generative model for summarization or response generation.
Connecting Azure Language from an Application
Every application that calls Azure AI Language needs the same three ingredients:
| Ingredient | What it is | Where to find it |
|---|---|---|
| Endpoint URL | The HTTPS address of your Azure AI Language resource | Foundry portal โ Foundry Tools โ Azure AI Language โ Keys and Endpoint |
| API Key | A secret credential that authorises your app to call the service | Same location as the endpoint |
| SDK (code library) | A pre-built package that handles the technical communication details for you - so you call simple functions rather than crafting raw web requests | Added to your code project |
The call sequence for any NLP feature always follows the same three steps: connect (set up a connection using your endpoint URL and API key), send text (call the relevant method - for example, analyse sentiment or extract entities), then read results (the service returns a structured response with scores and labels you can use in your app).
The Language playground's View code button shows this exact pattern with your real endpoint pre-filled.
Exam tip: Azure AI Language is a multi-capability service - one endpoint and one API key give you access to sentiment analysis, NER, PII detection, key phrase extraction, and summarisation. You do not need separate credentials per feature.
Foundry Portal Workflow for Text Analysis
Azure AI Language is accessible within Microsoft Foundry. It provides sentiment analysis, key phrase extraction, named entity recognition, PII detection, and summarization - all through a single endpoint and resource connection.
The Language playground lets you paste text and instantly see labels, confidence scores, and entity categories before writing any code. The View code button then generates the exact SDK snippet for whichever feature you tested.
For the exam, the key fact is that one Azure AI Language resource handles all NLP tasks - you do not create separate deployments for sentiment versus entity extraction.
๐ Note: The Language playground for Azure AI Language is accessed through the classic Foundry interface (New Foundry toggle disabled), not the New Foundry interface used for agents and generative AI.
๐ Hands-On Lab - Lab 3: Understand Guest Reviews with Azure AI Language Apply sentiment analysis, key phrase extraction, NER, and opinion mining to real hotel review data in a live Azure environment. Available with course access - Get the AI-901 Lab Course
Key Takeaways for the Exam
- NLP is how AI understands human language. It works on text that humans write or speak naturally.
- Tokenization splits text into processable units. Stop word removal filters noise. Stemming and lemmatization reduce words to root forms.
- TF-IDF scores word importance by frequency within a document minus frequency across all documents.
- Semantic models (like BERT and sentence transformers) represent meaning as vectors, enabling similarity search beyond keyword matching.
- Azure AI Language provides: sentiment analysis (positive/negative/neutral/mixed with confidence scores), key phrase extraction, named entity recognition (Person, Location, Organization, DateTime), language detection, and abstractive summarization.
- Foundry portal flow: Foundry Tools โ Azure AI Language โ Connect โ Language playground โ test โ View code.
- One Azure AI Language resource handles all NLP tasks - no separate model deployments per feature.
Official exam information from Microsoft
- Study guide for Exam AI-901: Microsoft Azure AI Fundamentals (skills measured, weights and passing score), and the Microsoft Certified: Azure AI Fundamentals certification page.