Welcome to Part 7 of our complete study series for the AI-103 Certification. After focusing on performance optimization and operational governance in Part 6, we now tackle: Computer Vision: Image/Video Generation & Multimodal Understanding. In this module, you will evaluate visual architectures across Azure AI Vision 4.0 (Read API, Dense Captioning, Spatial Analysis), multimodal reasoning, text-to-image synthesis and face-preserving editing with GPT-image-1, text-to-video workflows with Sora (preview), and enterprise media processing with Azure AI Video Indexer.
| Computer Vision, Generative Media & Multimodal AI: AI-103 Practice Questions |
In this free AI-103 practice exam module, you will solve 20 realistic scenario questions covering model selection, OCR pipelines, generative media APIs, edge model export (CoreML and TensorFlow), and video intelligence. Each question features a comprehensive explanation that details why the correct answer succeeds and why alternative options fail on exam day.
Computer Vision & Generative Media Practice Questions (Part 7)
Q1: High-Volume Binary Classification (Custom Vision vs. Multimodal LLMs)
A social media platform needs to apply one simple, narrow tag — "contains a cat" or "does not contain a cat" — to roughly 50 million uploaded photos per day, as cheaply and quickly as possible. The task never requires open-ended reasoning or explaining why a photo was tagged.
Which approach best fits this specific requirement?
• Why D is correct: For an extremely high-volume, narrow, binary classification task with no need for open-ended reasoning, a small dedicated Custom Vision classifier runs far cheaper and faster per image than invoking a large multimodal reasoning model on every single one of 50 million daily photos — the same cost/latency logic behind choosing a small, specialized model over a frontier LLM for simple, repetitive tasks.
• Why other options are incorrect:
- Why A is incorrect: Running a large multimodal model with a custom prompt on 50 million images daily is dramatically more expensive and slower than necessary for a single narrow yes/no tag.
- Why B is incorrect: Sora generates video from prompts; it has no role in classifying existing uploaded photos.
- Why C is incorrect: Regenerating each photo doesn't produce a classification label at all, and is far more computationally expensive than the task needs.
Q2: Azure Face API Limited Access Governance (Yes/No)
For each of the following statements about Azure's Face API, select Yes if the statement is true. Otherwise, select No.
1. Basic face detection (locating a face's bounding box and general attributes like head pose) is available without special approval.
2. Face identification and verification (matching a detected face against a specific named individual) are Limited Access features requiring Microsoft's approval before use.
3. Any developer can enable face identification for a public-facing app immediately after creating an Azure AI Services resource, with no additional review.
• Statement 1 is Yes: Basic detection-level capabilities are broadly available as part of standard Azure AI Vision face detection.
• Statement 2 is Yes: Because identification and verification involve matching a face to a specific person's identity, Microsoft gates these capabilities as Limited Access, requiring an application describing the intended use case before access is granted.
• Statement 3 is No: Limited Access features are not self-service — they require submitting a use-case application and receiving Microsoft's approval first, precisely to prevent the statement 3 scenario.
Q3: Printed & Handwritten Text Extraction (Read API OCR)
An automated claims processing solution must ingest scanned paper receipts, handwritten claims, and multi-page PDF documents to extract all textual content for downstream search indexing.
Which Azure AI Vision capability is designed to extract both printed and handwritten text from images and documents?
• Why C is correct: The Read API is the dedicated Optical Character Recognition (OCR) engine in Azure AI Vision. It is optimized to recognize and extract printed text, handwritten notes, and mixed typography from complex image files (JPEG, PNG, TIFF) and multi-page PDFs, providing word-level bounding boxes and detected language codes.
• Why other options are incorrect:
- Why A is incorrect: Image captioning generates a brief natural-language sentence describing the visual content of a photograph (e.g., "a person standing at a counter"); it does not transcribe individual lines of text.
- Why B is incorrect: Face detection detects human faces, landmarks, and spatial coordinates; it cannot detect or extract alphanumeric text.
- Why D is incorrect: Speaker recognition is an audio capability in Azure AI Speech used to authenticate or verify human voices from sound recordings; it does not process visual images or PDFs.
Q4: Few-Shot Custom Classifier Training (Transfer Learning & Foundation Models)
A quality inspection workflow in a manufacturing facility must detect rare surface flaws on specialized mechanical components. Because defective parts occur infrequently, the engineering team has gathered only 20 labeled images per defect category.
Which approach allows the team to build an accurate domain-specific classifier while requiring the smallest amount of labeled training data?
• Why B is correct: Azure AI Vision Custom Models utilize transfer learning atop pre-trained vision foundation models (such as Florence). Because the foundation model has already learned generalized visual features (edges, contours, surface textures, object geometry) from billions of images, fine-tuning requires only a tiny fraction of labeled domain samples (often as few as 10 to 20 images per class) to achieve production-grade classification accuracy.
• Why other options are incorrect:
- Why A is incorrect: Training a deep neural network from scratch with random initialization requires tens or hundreds of thousands of labeled images. With only 20 samples per class, the model will severely overfit and fail to generalize.
- Why C is incorrect: While data augmentation (rotations, crops, flips) artificially expands training variation, applying it without a pre-trained base model still leaves the architecture starved of initial visual representation data.
- Why D is incorrect: Pure heuristic scripts (evaluating raw RGB pixel thresholds) are brittle, cannot adapt to lighting variations or complex mechanical surface textures, and do not represent a machine learning classifier.
Q5: Video Analysis Architecture Limitations (Frame Sampling in Multimodal LLMs)
A team feeds a 45-minute training video directly into GPT-4o Vision, asking it to identify every distinct product shown throughout the entire recording. The model's summary misses several products that appear only briefly.
What is the most likely explanation for these gaps?
• Why A is correct: When analyzing video, vision-enabled chat models select a limited number of frames across the whole duration rather than processing every frame. Consequently, a product appearing only briefly between sampled frames can be missed entirely. This is a known architectural trade-off for long or detail-dense recordings, where a purpose-built tool like Video Indexer (which processes the full video timeline) is better suited.
• Why other options are incorrect:
- Why B is incorrect: The model does accept video input; the issue is sampling coverage across the timeline, not a total lack of support.
- Why C is incorrect: The scenario describes missed content due to gaps in time coverage, not a resolution rejection error.
- Why D is incorrect: Temperature affects output randomness in generated text; it doesn't control how many video frames are sampled for analysis.
Q6: Enterprise Media Processing & Speaker Diarization (Azure AI Video Indexer)
A media company needs to analyze a large archive of broadcast news footage to automatically identify topic changes, extract named speakers, generate transcripts, and produce searchable metadata for each news segment. The archive contains thousands of hours of video, and the solution must support semantic search across all content without requiring custom model development.
Which approach should you recommend?
• Why C is correct:
- Turnkey Multimodal Solution: Azure AI Video Indexer is specifically designed for enterprise media processing. It natively combines computer vision, speech-to-text, facial recognition, and natural language processing into a single automated pipeline.
- All-in-One Capabilities: Out of the box, Video Indexer extracts closed captions/transcripts, identifies named individuals via visual face databases and voice acoustics (speaker diarization), segments video into scenes and shots, and extracts topical themes.
- Zero Custom Model Training: It provides built-in search APIs and index structures, satisfying the core constraint: "without requiring custom model development."
• Why other options are incorrect:
- Why A is incorrect: Training and deploying a custom video segmentation model on Azure ML directly violates the constraint requiring a solution "without requiring custom model development." Furthermore, Azure AI Language processes text, not raw audio streams, so it cannot generate speech transcripts from video.
- Why B is incorrect: While Azure AI Vision Video Retrieval generates vector embeddings for multimodal search, it does not provide comprehensive speaker diarization/identification, full spoken audio transcription, or high-level structured topic-change metadata out of the box.
- Why D is incorrect: Extracting and feeding individual video frames from thousands of hours of video into GPT-4o Vision is financially prohibitive, introduces immense processing latency, requires building complex custom orchestration glue code, and completely ignores the audio track needed for transcription and speaker voice identification.
Q7: Programmatic Video Clip Rendering (Azure AI Video Indexer Projects)
A developer is building a content creation pipeline using Azure AI Video Indexer. The team has uploaded a 10-minute marketing video and needs to programmatically create an edited highlight reel by extracting specific time segments. When constructing the API request to generate the edited video output, which request body parameter or approach should the developer configure to define the specific time ranges that should be included in the final edited clip?
• Why this answer is correct:
- Video Indexer Projects API: In Azure AI Video Indexer, creating a customized or edited video clip from one or more existing indexed videos requires creating a Project (formerly playlists).
- Video Ranges Schema: When defining or updating a project, the payload requires an array of video ranges. Each range explicitly specifies the source videoId and a time interval (e.g., start: "00:01:15", end: "00:02:30").
- Render Operation: Once the project and its desired time intervals are configured, the client invokes the Render Project API. The service asynchronously stitches the specified cuts together into a new consolidated media file, which can then be downloaded via a temporary SAS URL.
• Why other options are incorrect:
- Why setting summarizationStyle is incorrect: Azure AI Video Indexer does not possess a parameter named summarizationStyle="highlights" that autonomously cuts and renders edited MP4 files. Physical video trimming and stitching must be explicitly defined via project time ranges.
- Why includeAudioEffects and keyframeInterval is incorrect: keyframeInterval controls how frequently keyframes are sampled during indexing, and audio effects settings relate to audio transcription/sound classification. Neither parameter specifies start/end clip boundaries or renders a new video file.
- Why thumbnailId is incorrect: The thumbnailId parameter is used with the Get Thumbnail API to download a single static image (JPEG/PNG) of a specific extracted frame. It cannot generate, trim, or output video clips.
Q8: Text-to-Image Synthesis from Scratch (GPT-image-1)
A marketing team needs to generate brand-new, photorealistic product mockup images from written descriptions (e.g., "a matte black water bottle on a marble countertop, studio lighting"). No existing photo is being edited — the image must be created from scratch.
Which model should be deployed for this task?
• Why C is correct: GPT-image-1, available through Azure OpenAI in Microsoft Foundry, performs text-to-image generation — synthesizing new images directly from a natural-language description, which is exactly the scenario described.
• Why other options are incorrect:
- Why A is incorrect: Custom Vision object detectors locate and label objects in existing images; they do not generate new images.
- Why B is incorrect: The Read API extracts text from images; it has no image generation capability.
- Why D is incorrect: Video Indexer analyzes and extracts insights from existing video content; it does not generate static images.
Q9: Visual Similarity Search Without Text (Multimodal Embeddings)
An e-commerce catalog with 500,000 product photos needs a "find visually similar items" feature: a shopper uploads a photo of a couch, and the system returns catalog items that look similar — without any of the products having manually written, searchable text descriptions.
To power this feature, which technique should the team use to vectorize every catalog image?
• Why B is correct: Multimodal embeddings convert an image's visual content into a dense vector, letting the system compute similarity between the uploaded photo and catalog images purely from visual appearance — no text descriptions required.
• Why other options are incorrect:
- Why A is incorrect: The Read API (OCR) only extracts visible text; most product photos have little or no readable text, and visual similarity has nothing to do with printed characters.
- Why C is incorrect: Face detection identifies human faces; it has no relevance to matching furniture or general product appearance.
Q10: Automated Text-to-Video Production Pipeline (Select 2)
A media production pipeline requires automating the creation of marketing video reels. The solution must generate original short video footage directly from descriptive natural language text prompts, then programmatically trim, splice, and assemble the resulting footage into finished video clips.
Which TWO capabilities are required to build this automated pipeline? (Each correct selection presents part of the solution. Select TWO.)
• Why C is correct: Generative text-to-video models (such as Sora deployed through Azure OpenAI / Azure AI Foundry) accept natural language prompts and synthesize photorealistic or stylized dynamic video scenes.
• Why E is correct: Generative models output raw video files. Achieving the requirement to trim, splice, and modify the generated clips requires an automated media processing and video-editing workflow pipeline (such as containerized media workflows or Azure Media processing functions).
• Why other options are incorrect:
- Why A is incorrect: Optical Character Recognition (OCR) detects, extracts, and reads printed text from images or video frames. It cannot generate visual video frames or edit multimedia assets.
- Why B is incorrect: Layout extraction parses text, tables, and document structures from business forms and PDFs; it has no application in video generation or video editing.
- Why D is incorrect: Real-time speech translation translates spoken audio streams between human languages; it does not synthesize visual scenes or manipulate video footage.
Q11: Generative Text-to-Video Workflow & Asynchronous Jobs (Sora Preview)
A content team needs to generate short promotional video clips from written scene descriptions and retrieve the finished MP4 files programmatically from an Azure-hosted service. A solutions architect must identify both the correct Azure service and the request pattern that service uses. Which recommendation is correct?
• Why this answer is correct:
- Sora in Microsoft Foundry (Preview): Microsoft Foundry / Azure OpenAI provides first-party access to OpenAI’s Sora text-to-video generation model (currently in preview), allowing developers to synthesize high-definition video clips (MP4) from descriptive natural-language prompts.
- Asynchronous Polling Pattern: Due to the heavy compute requirements of generative video diffusion models, rendering takes anywhere from dozens of seconds to several minutes. Therefore, the API is architected strictly as an asynchronous job pattern:
1. The client submits a POST request containing the prompt, resolution, and duration to create a video generation job.
2. The service responds immediately with a job ID (HTTP 202 Accepted).
3. The client periodically polls (e.g.,
GET /openai/v1/video/generations/jobs/{job_id}) to check the status until it returns succeeded.4. The client retrieves the SAS download URL to stream or save the final MP4 file.
• Why other options are incorrect:
- Why A is incorrect: Microsoft Foundry natively integrates text-to-video foundation models (such as Sora) as managed endpoints directly within the platform, eliminating the need to set up custom third-party infrastructure on Azure Machine Learning.
- Why C is incorrect: Azure AI Video Indexer is an analytical service designed to extract insights from existing media files; it does not synthesize new video footage from text prompts. Furthermore, video generation is far too resource-intensive to return synchronously in an HTTP response body.
- Why D is incorrect: GPT-4o accepts video frames and audio as inputs for visual reasoning, but it cannot render raw MP4 video streams as an output. Returning binary MP4 files inline inside a synchronous chat completion payload violates REST API limits.
Q12: Computer Vision & Generative Media Capabilities (Matching)
Match each requirement to the correct computer vision or generative capability. Each capability is used exactly once.
1. Sort scanned invoices into "paid" or "unpaid" based on their overall visual layout, using only a handful of labeled examples.
2. Locate and draw a bounding box around every scratch on a phone screen in an inspection photo.
3. Enable a "shop this look" feature that finds visually similar clothing items from an image, without relying on product text.
4. Create an entirely new hero image for a website from a text description of the desired scene.
• 1 → Custom Vision (Classification): Assigning one overall label to an entire image (paid vs unpaid) from a small labeled set is a classification task, not one requiring localized bounding boxes.
• 2 → Custom Vision (Object Detection): Locating and outlining each individual scratch with a bounding box requires object detection, which returns both a label and spatial coordinates per instance.
• 3 → Multimodal embeddings: Visual similarity search without relying on text descriptions is exactly what image-to-image vector embeddings enable.
• 4 → GPT-image-1: Generating a brand-new image from a text prompt is text-to-image generation, rather than analytical processing of an existing picture.
Q13: Region-Based Image Inpainting & Editing (GPT-image-1)
A real estate marketing team has an existing photo of an empty living room and wants to add a specific piece of furniture into a chosen area of that same photo, guided by a text description, while keeping the rest of the room untouched.
Which capability should the team use?
• Why B is correct: GPT-image-1 accepts an existing image as input and supports editing/inpainting — modifying a specified region of that image according to a text prompt while leaving the rest of the image intact, which is exactly the "add furniture to one part of this existing photo" requirement.
• Why other options are incorrect:
- Why A is incorrect: Sora generates new video from a prompt; it does not edit a specific region of an existing still photo.
- Why C is incorrect: The Read API extracts text; it has no image editing or generation capability.
- Why D is incorrect: Finding a similar existing stock photo does not modify the team's own specific room photo — it returns an unrelated image entirely.
Q14: Facial Feature Preservation in Image Editing (input_fidelity)
A team uses GPT-image-1 to edit a real customer's headshot photo — swapping the background and clothing style, while critically ensuring the person's actual face remains clearly recognizable and unaltered in the output.
Which parameter should the team configure to control this?
• Why D is correct: The
input_fidelity parameter in GPT-image-1's editing API controls how closely the output preserves the style and features of the subject in the original input image (supporting values such as high and low). Setting it to high ensures that facial structures and biometric characteristics remain faithful to the original reference photo during clothing or background modifications.• Why other options are incorrect:
- Why A is incorrect: Temperature affects text generation randomness in language models; it does not control visual preservation in image editing.
- Why B is incorrect:
top_p is a nucleus text-sampling parameter, not an image-editing visual fidelity control.- Why C is incorrect:
max_tokens governs completion token lengths; it has no mechanism to preserve facial details during generative image inpainting.
Q15: Synthetic Video Presenters (Text to Speech Avatar)
A company wants to turn written training scripts into short videos of a photorealistic 2D presenter speaking the script aloud, with lip movements synced to the generated speech — without filming an actual human presenter.
Which capability fits this requirement?
• Why A is correct: The Text to Speech Avatar feature in Azure AI Speech synthesizes video of a photorealistic 2D avatar speaking written text, with neural models generating lip movements that naturally synchronize with the synthesized voice audio — designed specifically for corporate training, virtual presenters, and automated broadcasting without physical cameras.
• Why other options are incorrect:
- Why B is incorrect: Video Indexer analyzes and extracts insights from pre-existing video; it cannot render or generate speaking avatar videos.
- Why C is incorrect: Sora generates general visual video scenes from text prompts; it is not architected for deterministic script-based lip-synced speech delivery.
- Why D is incorrect: GPT-image-1 generates and edits static images, not animated video clips with synchronized audio streams.
Q16: Real-Time Live Stream Queue Monitoring (Spatial Analysis)
A retail store wants to monitor, in real time, how many customers are currently in a checkout queue and trigger an alert when the queue exceeds a defined length, using existing live security camera feeds.
Which capability fits this requirement?
• Why C is correct: Spatial Analysis (part of Azure AI Vision) is specifically built to ingest real-time video streams (via RTSP feeds on IoT Edge), tracking people's positions, physical movement, and zone occupancy (such as checkout line length or customer dwell time) to trigger immediate alerts.
• Why other options are incorrect:
- Why A is incorrect: Video Indexer processes and indexes uploaded video files post-event; it is not designed for continuous live stream monitoring or instantaneous alerts.
- Why B is incorrect: The Read API extracts text characters; it cannot track human presence or measure queue occupancy.
- Why D is incorrect: Periodic still-image processing with generative models is far too slow, expensive, and unsuited for continuous physical space surveillance and real-time event alerting.
Q17: Visual Content Moderation Governance (Yes/No)
For each of the following statements about moderating images in a user-generated content platform, select Yes if the statement is true. Otherwise, select No.
1. Azure AI Content Safety can analyze uploaded images for harmful categories such as violence, sexual content, and self-harm, similar to how it analyzes text.
2. Image moderation and text moderation in Azure AI Content Safety are evaluated using entirely separate harm category definitions with no conceptual overlap.
3. An image generated by GPT-image-1 can still be independently checked with Azure AI Content Safety's image moderation after generation, in addition to the model's own built-in output moderation.
• Statement 1 is Yes: Azure AI Content Safety provides image moderation across harmful content categories analogous to its text moderation categories, letting a platform screen visual uploads the same way it screens text.
• Statement 2 is No: Image and text moderation in Content Safety are built around the same core harm categories (such as violence, sexual content, and self-harm) applied to each respective modality — they are not completely disconnected classification taxonomies.
• Statement 3 is Yes: While GPT-image-1 includes built-in output moderation, an enterprise application can still pass generated visual assets through Azure AI Content Safety's image moderation API as an independent, secondary defense layer before rendering them to end users.
Q18: Accessibility & Multi-Region Scene Descriptions (Dense Captioning)
An accessibility team wants to automatically generate a separate short description for each distinct object or region within a single complex photo — not just one caption for the whole image — so a screen reader can describe a busy scene in detail, region by region.
Which capability fits this requirement?
• Why A is correct: Dense captioning detects multiple prominent objects and regions across a single image, generating a distinct natural-language caption and spatial bounding box for each one. This provides the granular, region-by-region visual detail required by assistive screen readers.
• Why other options are incorrect:
- Why B is incorrect: A single whole-image caption summarizes the entire scene in one high-level sentence, losing granular regional distinctions.
- Why C is incorrect: OCR (Read API) extracts visible printed or handwritten text characters; it does not describe visual objects or regions that contain no text.
- Why D is incorrect: Face detection locates human faces; it cannot caption general non-human objects or scenery.
Q19: Offline Edge Model Export for Mobile Devices (Select 2)
A company wants to run a trained Custom Vision image classifier fully offline on both iOS and Android mobile devices, with the model embedded directly in each app and no server round-trip needed at inference time.
Which TWO export options support this requirement? (Each correct selection covers one of the two platforms. Select TWO.)
• Why A and B are correct: Custom Vision (when trained using a "Compact" domain) supports exporting trained models into on-device formats: CoreML for native Apple iOS hardware acceleration and TensorFlow / TensorFlow Lite for Android applications. This enables real-time on-device classification without internet connectivity.
• Why other options are incorrect:
- Why C is incorrect: Calling a cloud-hosted REST endpoint requires active internet round-trips for every prediction, violating the offline edge requirement.
- Why D is incorrect: A text document containing weights is not an executable, compiled machine learning runtime package.
- Why E is incorrect: Raw training images provide zero inference functionality without a compiled model architecture and execution engine.
Q20: Defect Localization & Bounding Boxes (Object Detection)
An automated quality assurance system analyzes photos of manufactured components on an assembly line. The software must pinpoint the exact location of specific defects, outline each defective area using bounding boxes, and assign a specific category label to each detected defect.
Which computer vision capability is required for this specific task?
• Why B is correct: Object detection is specifically designed to locate one or more items within an image, assign a classification label to each item, and output spatial coordinates defining a bounding box around each detected area. Because the scenario demands outlining localized defect boundaries, object detection is the required architecture.
• Why other options are incorrect:
- Why A is incorrect: Image classification assigns one or more labels to the image as a whole (e.g., "defective" or "flawless"). It does not generate spatial coordinates or identify where defects are located inside the image.
- Why C is incorrect: Optical Character Recognition (OCR) extracts alphanumeric text from scanned documents and signage; it cannot identify physical manufacturing component defects.
- Why D is incorrect: Spatial analysis is an Azure AI Vision feature optimized for real-time video feeds to monitor human movements, occupancy, and physical space interactions (e.g., tracking dwell time or queue lengths), making it unsuitable for static component defect inspection.
Microsoft Foundry Computer Vision & Generative Media Cheat Sheet
Earning the official Developing AI Apps and Agents on Azure (AI-103) credential requires knowing when to pick a specialized computer vision service over an expensive multimodal model, and how to configure generative media tools correctly. Use this comparison table to quickly review the core capabilities before test day:
| Vision Capability | Input → Output | Key Architectural Differentiator | Best-Fit Scenario & Status |
|---|---|---|---|
| Custom Vision (Classification) | Image → Single Label | Uses transfer learning on foundation models; trains with as few as 10–20 images per class. | High-volume binary tags (e.g., 50M photos/day) where cost and latency matter. |
| Custom Vision (Object Detection) | Image → Labels + Bounding Boxes | Pinpoints exact spatial coordinates and boundaries for multiple items in one image. | Manufacturing quality checks, defect localization, and inventory counting. |
| Read API (Azure AI Vision OCR) | Image / Multi-Page PDF → Extracted Text | Extracts printed characters, cursive handwriting, and mixed fonts with word coordinates. | Scanning receipts, paper claims, and multi-page compliance documents. |
| Dense Captioning | Image → Multi-Region Descriptions | Generates separate natural-language captions for different areas of a busy image. | Accessibility tools and screen readers that need detailed, region-by-region context. |
| Multimodal Embeddings | Image → Dense Vector | Encodes visual features directly into vector space for cosine similarity searches. | Zero-text visual search (e.g., “find visually similar clothing or furniture”). |
| Spatial Analysis | Live RTSP Video Stream → Real-Time Events | Runs on IoT Edge containers to track physical movements, dwell times, and zones. | Real-time checkout queue alerts and store floor monitoring. |
| GPT-image-1 (Generate & Edit) | Text Prompt (+ Image) → New / Edited Image | Supports inpainting and uses the input_fidelity parameter (high/low) to protect face details. |
Product mockups from scratch and photo edits requiring facial preservation. |
| Sora (Generative Video) | Text Prompt → High-Definition MP4 | Preview: Asynchronous polling job pattern (POST job → poll status → download SAS URL). |
Synthesizing original marketing video footage directly from scene descriptions. |
| Azure AI Video Indexer | Recorded Video → Rich Metadata | Turnkey media pipeline: speech-to-text, speaker diarization, OCR, face detection, and Project rendering. | Archive indexing, automatic transcripts, and programmatic highlight reel rendering. |
| Mobile Edge Export | Compact Custom Model → Local File | Exports trained models to CoreML (iOS) and TensorFlow Lite (Android). | Fully offline, on-device classification with zero network latency. |
Key Takeaways for AI-103 Domain 2 (Vision, Video & Generative Media)
• Specialized Models vs. Multimodal LLMs: A major exam trap is choosing GPT-4o Vision for simple, high-volume tasks. If an app needs to classify 50 million images a day with a simple yes/no label, deploying a lightweight Custom Vision classifier is dramatically cheaper and faster than prompting a large frontier model.
• GPT-image-1 and
input_fidelity: When using GPT-image-1 to edit photos (such as swapping backgrounds or outfits), settinginput_fidelityto high ensures the person’s facial features and identity stay unchanged. Remember:temperatureandmax_tokensdo not preserve visual features in image editing.• Sora Asynchronous Polling (Preview): Video generation diffusion models require heavy compute time. Sora in Microsoft Foundry uses an asynchronous job pattern: your app submits a
POSTrequest, receives a job ID, pollsGET /openai/v1/video/generations/jobs/{job_id}until the status shows succeeded, and then retrieves the download URL.• Video Indexer vs. GPT-4o Vision Limitations: When you pass a long video to GPT-4o Vision, it samples only a limited number of frames across the timeline. Short events between sampled frames can be missed completely. For full-timeline audio transcription, speaker diarization, and scene-change detection, Azure AI Video Indexer is the correct enterprise choice.
• Dense Captioning vs. Standard Image Captioning: Standard image captioning returns one sentence summarizing the whole picture. Dense captioning detects multiple regions across the scene and returns distinct bounding boxes and captions for each, making it the right pick for detailed accessibility screen readers.
• Offline Mobile Deployment: To run an image classifier completely offline inside mobile apps without internet round-trips, train a Custom Vision model using a Compact domain and export it to CoreML for iOS or TensorFlow for Android.
• Face API Responsible AI Guardrails: Basic face detection (locating a face box) is widely available. However, face identification and verification (matching a face to a named person) are Limited Access features that strictly require Microsoft use-case approval before they can be enabled.
Frequently Asked Questions (AI-103 Computer Vision & Media FAQ)
Here are concise, straightforward answers to the most common computer vision and generative media questions tested on the AI-103 exam:
Are these free AI-103 practice exam questions updated for the 2026 syllabus?
Yes. This free AI-103 practice test covers the latest 2026 exam updates, including GPT-image-1 inpainting controls, the input_fidelity parameter, Sora video generation workflows, Azure AI Vision 4.0 Read API, and Azure AI Video Indexer project architectures.
What does the input_fidelity parameter do in GPT-image-1?
The input_fidelity parameter controls how closely an edited image preserves the original subject's style and facial features during inpainting or modifications. Setting it to high ensures that a person's face remains recognizable when changing clothing or backgrounds.
Is Sora text-to-video generally available in Microsoft Foundry?
No, Sora is currently in preview in Microsoft Foundry / Azure OpenAI. The service exposes an asynchronous job API where developers submit video generation requests, poll the job status endpoint, and download the finished MP4 file once processing completes.
Why does GPT-4o Vision miss short events in long videos?
Multimodal chat models sample a limited budget of keyframes across a video rather than reading every single frame. If a product or event appears briefly between sampled frames, the model will miss it. Dedicated tools like Azure AI Video Indexer process the entire timeline and are better suited for comprehensive video analysis.
Can Custom Vision models run offline on mobile devices without an internet connection?
Yes. When you train a Custom Vision project using a Compact domain, you can export the compiled model directly to CoreML for iOS devices and TensorFlow / TFLite for Android. This allows applications to perform real-time image classification fully on-device without cloud API calls.
Next Step in Your AI-103 Certification Journey
Congratulations on completing this free AI-103 practice exam module on Computer Vision, Generative Media, and Multimodal Understanding! You now know how to design cost-effective vision pipelines, run generative image and video workflows, and apply proper governance controls.
With visual data covered, we turn our attention to language processing and voice interfaces. In Part 8 of our series, we focus on Text Analysis & Speech Solutions, diving into Azure AI Language, conversational language understanding (CLU), sentiment analysis, speech-to-text diarization, and custom voice models.
Continue your preparation by jumping to the next part using the navigation below: