Microsoft released two new models for public preview on July 23: MAI-Image-2.5-Pro and MAI-Voice-2-Flash. It also shared internal data showing that after the Customer Support Center switched to MAI-Voice-2-Flash, Hashrate costs fell by 89%. Microsoft CEO Satya Nadella said that when its own models match or surpass frontier alternatives, it routes traffic to MAI.
Function positioning and pricing of MAI-Image-2.5-Pro and MAI-Voice-2-Flash
MAI-Image-2.5-Pro is positioned for high-end image generation. Pricing is $5 per million text input tokens and $106 per million image output tokens. Bing Image Creator has fully switched to MAI-Image-2.5. WPP global creative director Rob Reilly called it “a major leap forward for generative media tools” in a Microsoft announcement.
MAI-Voice-2-Flash is positioned for large-scale voice processing scenarios (such as customer support centers). It is twice as fast as its predecessor and 32% lower in cost. After the customer support center switched to this model, Hashrate costs dropped by 89%. The announcement also disclosed deployment results for Bing Image Creator, PowerPoint (GPU costs lowered by 84% vs. GPT-Image-2), OneDrive (storage efficiency +26%, latency -25%), and Dragon Copilot (error rate -50%).
Concrete results of GPU cost savings
Microsoft disclosed performance results across these scenarios:
MAI-Voice-2-Flash (customer support center voice): Hashrate cost down 89%; twice as fast as the predecessor, with costs 32% lower
MAI-Image-2.5-Pro (PowerPoint): GPU costs lowered by 84% vs. OpenAI GPT-Image-2
OneDrive (switching models): Storage efficiency increased by 26%, latency lowered by about 25%
Dragon Copilot (MAI-Transcribe-1.5, medical transcription): Serves 170k medical professionals; handled 28 million patient records last quarter; transcription and language recognition error rates across 58 languages are relatively down 50%
MAI-Code-1-Flash (GitHub Copilot): Code adoption is about 10% higher than GPT-5.4 Mini and Claude Haiku 4.5; token consumption is down 10%; can run on H100 and even A100 GPUs
Nadella’s flywheel strategy and MAI’s traffic logic for replacing frontier models
On X, Nadella posted titled “Frontier diffusion and control,” outlining Microsoft’s model strategy: when its models perform on par with or better than frontier alternative solutions, it directs traffic to MAI to deliver services at larger scale with lower costs. The strategy is supported by a “hill-climbing” methodology: instead of determining life or death through a single training run, it makes data, models, and products “work together” to form a continuously iterating flywheel.
He also said that in the evaluation system, “even if you remove any one model, it should still keep climbing”—keeping memory and skills outside the model to form the real control Microsoft owns. After MAI-Code-1-Flash is further trained in Excel’s reinforcement learning environment, it can run on older H100 and A100 GPUs, match GPT-5.6 across most Excel tasks, and free up the latest GB200 cluster for training rather than service requests.
FAQ
What are the prices for MAI-Image-2.5-Pro and MAI-Voice-2-Flash?
MAI-Image-2.5-Pro is priced at $5 per million text input tokens and $106 per million image output tokens, and is available in Azure public preview. MAI-Voice-2-Flash is twice as fast as its predecessor and 32% lower in cost; pricing details are subject to Azure’s official announcement.
How much does Microsoft say GPU costs drop after switching to MAI models?
According to internal data disclosed by Microsoft, after the customer support center switched to MAI-Voice-2-Flash, Hashrate costs dropped by 89%; after PowerPoint switched to an MAI image model, GPU costs were lowered by 84% compared with GPT-Image-2; after Dragon Copilot switched to MAI-Transcribe-1.5, transcription and language recognition error rates were relatively down 50%. However, all the above figures come from Microsoft internal evaluations, not independent third-party benchmark tests.
How can enterprise customers use Microsoft’s MAI training approach through Azure?
Microsoft uses Azure’s Foundry and Frontier Tuning products to let enterprise customers train models using their own data for dedicated models. Microsoft emphasizes that these models are trained on “clean, auditable enterprise-grade data” and “do not undergo third-party model distillation,” to address compliance concerns about the sources of training data.