Changing the engine while the plane is flying: migrating 60,000 apps under live load
July 24, 2026For the first time, real-time transcription goes multilingual
July 24, 2026One of the core principles in Microsoft Foundry is giving customers flexibility to choose the right model for the right job. As we’ve continued to expand the Microsoft AI (MAI) model family, we’ve heard a consistent theme from developers and enterprises: they want more choice based on their specific workload requirements. Some need the highest possible quality, fidelity, and consistency for professional production workflows. Others prioritize responsiveness and cost-efficiency for frequent, real-time responses.
Today, we’re introducing two new additions to the MAI portfolio:
- MAI-Image-2.5 Pro, designed for customers who need maximum visual fidelity and creative control
- MAI-Voice-2 Flash, built for low-latency voice applications where every millisecond matters.
Together, these models extend the MAI family with specialized options that help developers optimize for the jobs they need to do. Let’s dive in.
MAI-Image-2.5 Pro: Built for professional creative workflows
MAI-Image-2.5 Pro is our newest high-fidelity image generation model, designed for scenarios where visual accuracy, consistency, and creative control are critical. The model delivers stronger object consistency, improved alignment with creative intent, and enhanced visual reasoning and world knowledge, making it the ideal choice when quality and fidelity matter more than throughput.
As generative AI moves from experimentation to production, creative teams increasingly need models that can reliably generate assets that meet professional standards without extensive manual editing. MAI-Image-2.5 Pro was built to address those needs. Here are some of the scenarios it best fits:
- Campaign Hero Assets & Multi-Frame Storytelling. Creative agencies and marketing teams can create polished campaign visuals while maintaining consistent products, characters, and brand elements across every asset. Whether producing launch campaigns, social media variants, retail signage, or digital advertising, Image-2.5 Pro helps ensure visual consistency throughout the customer journey.
- Product Photography & E-Commerce Catalogs. Retailers and consumer brands can generate high-quality product imagery with accurate rendering of packaging, labels, materials, and reflections. The model’s improved object consistency helps maintain identical product representation across multiple angles, color variations, and lifestyle settings.
- Storyboarding & Pre-Visualization. Film studios, game developers, and creative production teams can rapidly develop visual concepts while maintaining consistency across characters, environments, props, and scenes. Enhanced visual reasoning enables more accurate interpretation of camera direction, lighting, and staging instructions.
- Regulated Industry Content Creation. Organizations in healthcare, financial services, manufacturing, and other regulated industries can generate imagery that requires greater real-world accuracy and domain understanding, reducing the effort needed to correct inaccuracies before customer-facing use
When should customers use MAI-Image-2.5 Pro vs. MAI-Image-2.5?
Both models deliver high-quality image generation capabilities, but they are optimized for different priorities.
Choose MAI-Image-2.5 Pro when:
- You need maximum image fidelity and creative quality.
- Object consistency across multiple images is critical.
- Your workflow depends on visual reasoning, world knowledge, and close adherence to creative direction.
- You are producing professional marketing, advertising, product design, or enterprise creative assets.
MAI-Voice-2 Flash: Real-time voice experiences at scale
We’re also introducing MAI-Voice-2 Flash, a new low-latency text-to-speech model that extends MAI-Voice-2 with faster response times and greater cost efficiency across more than 15 supported languages.
As voice becomes a critical interface for AI applications, responsiveness is increasingly important to the end-user experience. Whether a customer is speaking with an AI-powered support agent, interacting with a voice assistant, or navigating a self-service phone system, long pauses can make experiences feel slow and unnatural. MAI-Voice-2 Flash was built to address those scenarios. Here are some of the scenarios on when to choose MAI-Voice-2 Flash:
- Call Center Agents: Customer support organizations can generate spoken responses in real time, minimizing delays between conversation turns and enabling more natural customer interactions. Low latency helps improve customer experiences while supporting AI-powered service, support, and sales workflows.
- Conversational Voice Assistants: Developers can build voice-enabled copilots, assistants, and intelligent applications that respond nearly instantly. The result is a more fluid, natural conversation that feels interactive rather than turn-based.
- Interactive Voice Response (IVR) Systems: Organizations can modernize traditional phone systems with dynamic AI-generated speech that responds contextually to customer requests while maintaining a responsive user experience.
When should customers use MAI-Voice-2 Flash vs. MAI-Voice-2?
The distinction between the two voice models comes down to whether customers are optimizing for voice identity or real-time responsiveness.
Choose MAI-Voice-2 Flash when:
- Low latency is a primary requirement.
- You are building conversational assistants, IVR systems, or call center experiences.
- Users expect immediate spoken responses as part of a live interaction.
- Cost efficiency and responsiveness are more important.
Get started today in Microsoft Foundry
MAI-Image-2.5 Pro is available through Microsoft Foundry, providing developers with access to Microsoft’s latest advancements in image generation. Pricing starts at $5 per 1M tokens for text input, and $106 per 1M tokens for image output.
MAI-Voice-2 Flash is available through Azure Speech, allowing customers to leverage Azure Speech’s enterprise-grade reliability, scalability, and ecosystem while benefiting from Microsoft’s latest voice technology. Pricing starts at $15 per 1M characters.