Skip to content
GetHandsOn.ai

AI-901 Study Guide


Module 5 of 813 min read

AI and Speech

Learn how AI converts spoken language to text and text back to speech. This chapter covers the full speech recognition and synthesis pipelines, real-world use cases, and how to use Azure Speech in Microsoft Foundry to build voice-enabled applications.

These study notes summarise Microsoft Learn material for Exam AI-901. For the official skills measured, see the Microsoft Learn study guide for Exam AI-901.

On this pageShow

What Is AI Speech?

AI speech capabilities allow software applications to understand spoken language and respond in spoken language. In Module 1 you saw two quick examples: Microsoft Teams generating live captions during meetings (speech-to-text), and a virtual assistant reading calendar reminders aloud (text-to-speech). This module covers how both work under the hood and how to build them in Azure.

Two core capabilities define this domain:

  • Speech recognition (speech-to-text): converts an audio signal of spoken words into written text
  • Speech synthesis (text-to-speech): converts written text into a natural-sounding audio signal

Together these form the foundation of voice-driven applications, from virtual assistants and meeting transcription tools to accessible interfaces and AI agents you can speak to.

The two capabilities are essentially mirror images of each other. Let's start with the recognition side: how AI listens.


Speech Recognition: How AI Listens

Speech recognition turns the physical phenomenon of sound into words a computer can process. The pipeline has six stages.

1. Audio Capture

A microphone converts sound waves into a digital signal, sampled thousands of times per second (typically 16,000 samples per second for speech). The quality of the microphone and the level of background noise directly affect accuracy at every stage that follows.

2. Feature Extraction

The raw audio signal is divided into short frames (roughly 25 milliseconds each). For each frame, the system computes Mel-frequency cepstral coefficients (MFCCs), a compact mathematical representation of the sound's spectral characteristics. MFCCs strip out irrelevant variation like recording volume while preserving the features that distinguish different phonemes (the basic units of sound in a language).

3. Acoustic Modeling

The acoustic model takes the MFCC features and maps them to phoneme probabilities. It answers the question: given this audio frame, what sound is most likely being produced? Modern acoustic models are deep neural networks trained on thousands of hours of labeled speech.

4. Language Modeling

The language model takes the phoneme sequence and estimates the most likely word sequence, using statistical knowledge of how words follow each other in real language. This is why speech recognition systems handle "I scream" versus "ice cream" correctly depending on context.

5. Decoding

The decoder combines the acoustic model output and the language model output to find the most probable word sequence across all possible interpretations. Beam search is the standard algorithm for this, efficiently exploring the most promising options without evaluating every possible combination.

6. Post-processing

The raw transcript is cleaned up: punctuation is added, numbers are formatted correctly, filler words may be removed, and speaker labels added if the system supports diarization (identifying who said what).

Common Speech Recognition Scenarios

  • Customer service: transcribing calls for quality assurance and analytics
  • Meeting transcription: generating searchable notes and action items
  • Voice-activated assistants: hands-free device control
  • Healthcare documentation: doctors dictating notes directly into electronic health records

Now you know how AI converts sound to text. The companion process, text to speech, works in the opposite direction.


Speech Synthesis: How AI Speaks

Speech synthesis converts written text into spoken audio through a four-stage pipeline.

1. Text Normalization

Raw text contains abbreviations, numbers, dates, and symbols that need to be expanded into spoken form before pronunciation is attempted. "Dr. Smith ordered 3 items for $25.50" becomes "Doctor Smith ordered three items for twenty-five dollars and fifty cents."

2. Linguistic Analysis

The normalized text is broken into phonemes and the system determines how to pronounce each word. Grapheme-to-phoneme (G2P) conversion maps written letters to their spoken sound representations. This handles exceptions like the different pronunciations of "read" in "I will read the book" versus "I have read the book."

3. Prosody Generation

Prosody refers to the rhythm, stress, and intonation of speech. This stage determines which syllables to emphasize, where to pause, and how to modulate pitch across a sentence. Good prosody is what separates natural-sounding speech from robotic output.

4. Waveform Generation

A neural vocoder converts the phoneme and prosody information into an actual audio waveform. Modern neural vocoders (such as WaveNet and its successors) produce speech that is nearly indistinguishable from a human voice. The resulting audio can be streamed in real time or saved as a file.

Common Speech Synthesis Scenarios

  • Accessibility tools: reading web content aloud for visually impaired users
  • Audiobooks and content narration: converting written content to audio at scale
  • IVR systems: phone menus and automated response systems
  • Multilingual applications: reading back content in the user's language with natural-sounding pronunciation

Recognition and synthesis each work independently, but the real power comes when you combine them.


Combining Recognition and Synthesis

Many real-world applications need both directions of speech processing. A voice-driven AI agent works like this:

  1. The user speaks a question (recognition converts it to text)
  2. The application processes the text and generates a response
  3. The response is spoken back to the user (synthesis converts text to audio)

This pattern is used in customer service IVR systems, language learning apps, and voice-driven AI agents.

The Multimodal Voice Pattern (Exam Objective)

The AI-901 exam tests a specific architectural pattern that combines Azure AI Speech with a deployed generative model to create a full voice-to-voice AI experience. This is distinct from using Azure AI Speech alone.

The pattern works in four steps:

  1. Capture audio - the user speaks; your application captures the audio stream
  2. Transcribe - Azure AI Speech (STT) converts the audio to a text transcript
  3. Reason - the transcript is sent as the user message to a deployed model (e.g. GPT-4o) via its endpoint; the model generates a text response
  4. Speak - Azure AI Speech (TTS) converts the response text back to audio and plays it to the user

This is called the multimodal voice pattern because the overall pipeline accepts speech input and produces speech output, with a language model doing the reasoning in between.

Why this is exam-relevant: The AI-901 study guide explicitly lists "use deployed multimodal models to enable AI solutions that process and generate content across multiple types of data" as an exam objective. The multimodal voice pattern is a primary example of this objective in practice. Expect exam questions that ask which services are involved (Azure AI Speech for audio, a deployed GPT-4o for reasoning) or how the components connect.

Voice Live in Foundry simplifies this by bundling the STT → model → TTS pipeline into a single agent configuration - but the underlying pattern is the same.


Key Considerations for Speech Applications

  • Audio quality: background noise, microphone distance, and bandwidth all affect recognition accuracy. Test with realistic audio conditions, not just clean studio recordings.
  • Language and dialect support: verify that your target languages and regional accents are supported by the service you choose.
  • Privacy and compliance: audio data containing personal speech requires careful handling to meet data protection regulations.
  • Latency: real-time conversation requires low-latency processing; batch transcription can tolerate delays. Know which you need before choosing a deployment pattern.
  • Accessibility: always provide a text-based alternative for users who prefer or require it. This directly supports the Inclusiveness principle - one of the six responsible AI principles Module 8 explores in full.

Azure Speech in Foundry handles the underlying complexity for all of this.


Azure Speech in Microsoft Foundry

Azure Speech is the Azure service that provides both speech-to-text and text-to-speech capabilities. In Microsoft Foundry, it is available as a pre-configured service within your Foundry project.

Speech-to-Text

Accepts a streaming audio input or an audio file and returns a transcript. Supports real-time transcription and batch transcription of pre-recorded audio. Features include:

  • Punctuation and capitalization insertion
  • Profanity filtering
  • Speaker diarization (labeling who said what in multi-speaker audio)
  • Custom vocabulary for domain-specific terms

Text-to-Speech

Accepts text and returns an audio stream. Key features include:

  • Multiple voices and languages
  • SSML support (Speech Synthesis Markup Language) for fine-grained control over prosody, rate, and pitch
  • Custom Neural Voice for creating branded voices trained on a small amount of sample audio

Voice Live

Foundry also includes Voice Live, which enables speech-capable agents. An agent configured with Voice Live can accept spoken input and respond with synthesized speech, creating a complete conversational voice experience without needing to wire together the recognition and synthesis steps manually.

Connecting Azure Speech from an Application

Every application that calls Azure AI Speech needs the same three ingredients:

IngredientWhat it isWhere to find it
Endpoint URLThe HTTPS address of your Azure Speech resourceFoundry portal → Foundry Tools → Azure AI Speech → Keys and Endpoint
API KeyA secret credential that authorises your app to call the serviceSame location as the endpoint
SDK (code library)A pre-built package that handles audio streaming and communication details for youAdded to your code project

The call sequence always follows the same three steps: connect (set up a speech connection using your endpoint URL and API key), send audio or text (stream audio to get a transcript back, or pass text to get audio back), then read results (your app receives the transcript text or the synthesised audio).

The Speech playground's View code button shows this exact pattern with your real endpoint pre-filled.

Exam tip: One Azure AI Speech resource covers both speech-to-text and text-to-speech. You do not need separate endpoints or keys per capability.


Foundry Portal Workflow for Speech

Azure AI Speech is available through Foundry Tools in your Foundry project. It provides both speech-to-text (STT) and text-to-speech (TTS) through the Speech playground, where you can upload an audio file or speak directly into the microphone to test recognition, or type text and preview synthesis with any of the available neural voices.

The View code button generates the SDK snippet for whichever speech feature you tested.

For the exam, remember that neural voice names follow the pattern <locale>-<Name>Neural (for example en-US-AriaNeural), and that real-time STT and batch transcription are distinct modes with different latency trade-offs.

📌 Note: In the New Foundry interface, Azure Speech is not a standalone playground - speech capability is enabled via the Speech mode toggle on an agent, which integrates Azure Speech Voice Live (STT + TTS) directly into the agent.

The exam also tests a specific pattern: using Azure Speech together with a deployed multimodal model to create a voice-enabled AI experience. The flow is:

  1. Azure Speech (STT) converts the user's spoken audio into a text transcript
  2. The transcript is passed as the user message to a deployed model like GPT-4o
  3. The model's response is passed to Azure Speech (TTS) which converts it back to audio

This pattern lets a multimodal model respond to spoken input without the model itself needing to process audio directly.

🔒 Hands-On Lab - Lab 4: Add Voice to Meridian's Platform Build a speech-to-text transcription pipeline and a complete voice loop (speak in, Aria responds aloud) in a live Azure environment using Azure AI Speech. Already enrolled? Labs are launching shortly - stay tuned!. New here? Get the AI-901 Lab Bundle →


Key Takeaways for the Exam

  • Speech-to-text (STT): converts spoken audio into a text transcript. Used for voice commands, transcription services, and live captioning.
  • Text-to-speech (TTS): converts text into natural-sounding audio. Neural voices (e.g. en-US-AriaNeural) are indistinguishable from human speech.
  • Real-time vs batch: real-time STT processes a live microphone stream; batch transcription handles pre-recorded audio files asynchronously.
  • Speaker recognition: identifies who is speaking. Speaker verification confirms a claimed identity; speaker diarization labels each speaker segment.
  • Speech translation: transcribes audio in one language and translates to another in a single step (e.g. Spanish audio → English text).
  • Azure AI Speech is accessed via Azure Foundry Tools or directly as an Azure AI service.
  • Multimodal voice pattern: Azure Speech (STT) → transcript as user message → deployed model (GPT-4o) → response text → Azure Speech (TTS) → audio output.
  • Foundry portal flow: Foundry Tools → Azure AI Speech → Connect → Speech playground → upload audio or use microphone → View code.
  • Neural voice names follow the pattern <locale>-<Name>Neural (e.g. en-US-AriaNeural).

Official exam information from Microsoft

Get the full AI-901 study guide as a PDF, freeAll 8 modules in one printable file. Enter your email on the guide page and it is yours.

Keep going

Get the full AI-901 guide as a PDFEvery module in one file. Free after you enter your email.
Practice AI-901 questionsExam-style questions with explanations, free to start.
Hands-on AI-901 labsApply this in a real Azure environment.
AI-901 guide overviewAll modules, pick what to read next.