Microsoft releases MAI-Voice-2-Flash, cutting GPU costs by up to 89%

MSFT-2.24%
Key Takeaways
  • Microsoft released MAI-Voice-2-Flash and MAI-Image-2.5-Pro into public preview on July 23 with significant cost reductions.
  • Microsoft's customer service centers achieved 89% compute cost reduction after adopting MAI-Voice-2-Flash for speech processing.
  • MAI-Code-1-Flash can run on older H100 and A100 GPUs while matching GPT-5.6 performance on Excel tasks.

Microsoft released two new models for public preview on July 23: MAI-Image-2.5-Pro and MAI-Voice-2-Flash. It also shared internal data showing that after the Customer Support Center switched to MAI-Voice-2-Flash, Hashrate costs fell by 89%. Microsoft CEO Satya Nadella said that when its own models match or surpass frontier alternatives, it routes traffic to MAI.

Function positioning and pricing of MAI-Image-2.5-Pro and MAI-Voice-2-Flash

MAI-Image-2.5-Pro is positioned for high-end image generation. Pricing is $5 per million text input tokens and $106 per million image output tokens. Bing Image Creator has fully switched to MAI-Image-2.5. WPP global creative director Rob Reilly called it “a major leap forward for generative media tools” in a Microsoft announcement.

MAI-Voice-2-Flash is positioned for large-scale voice processing scenarios (such as customer support centers). It is twice as fast as its predecessor and 32% lower in cost. After the customer support center switched to this model, Hashrate costs dropped by 89%. The announcement also disclosed deployment results for Bing Image Creator, PowerPoint (GPU costs lowered by 84% vs. GPT-Image-2), OneDrive (storage efficiency +26%, latency -25%), and Dragon Copilot (error rate -50%).

Concrete results of GPU cost savings

Microsoft disclosed performance results across these scenarios:

MAI-Voice-2-Flash (customer support center voice): Hashrate cost down 89%; twice as fast as the predecessor, with costs 32% lower

MAI-Image-2.5-Pro (PowerPoint): GPU costs lowered by 84% vs. OpenAI GPT-Image-2

OneDrive (switching models): Storage efficiency increased by 26%, latency lowered by about 25%

Dragon Copilot (MAI-Transcribe-1.5, medical transcription): Serves 170k medical professionals; handled 28 million patient records last quarter; transcription and language recognition error rates across 58 languages are relatively down 50%

MAI-Code-1-Flash (GitHub Copilot): Code adoption is about 10% higher than GPT-5.4 Mini and Claude Haiku 4.5; token consumption is down 10%; can run on H100 and even A100 GPUs

Nadella’s flywheel strategy and MAI’s traffic logic for replacing frontier models

On X, Nadella posted titled “Frontier diffusion and control,” outlining Microsoft’s model strategy: when its models perform on par with or better than frontier alternative solutions, it directs traffic to MAI to deliver services at larger scale with lower costs. The strategy is supported by a “hill-climbing” methodology: instead of determining life or death through a single training run, it makes data, models, and products “work together” to form a continuously iterating flywheel.

He also said that in the evaluation system, “even if you remove any one model, it should still keep climbing”—keeping memory and skills outside the model to form the real control Microsoft owns. After MAI-Code-1-Flash is further trained in Excel’s reinforcement learning environment, it can run on older H100 and A100 GPUs, match GPT-5.6 across most Excel tasks, and free up the latest GB200 cluster for training rather than service requests.

FAQ

What are the prices for MAI-Image-2.5-Pro and MAI-Voice-2-Flash?

MAI-Image-2.5-Pro is priced at $5 per million text input tokens and $106 per million image output tokens, and is available in Azure public preview. MAI-Voice-2-Flash is twice as fast as its predecessor and 32% lower in cost; pricing details are subject to Azure’s official announcement.

How much does Microsoft say GPU costs drop after switching to MAI models?

According to internal data disclosed by Microsoft, after the customer support center switched to MAI-Voice-2-Flash, Hashrate costs dropped by 89%; after PowerPoint switched to an MAI image model, GPU costs were lowered by 84% compared with GPT-Image-2; after Dragon Copilot switched to MAI-Transcribe-1.5, transcription and language recognition error rates were relatively down 50%. However, all the above figures come from Microsoft internal evaluations, not independent third-party benchmark tests.

How can enterprise customers use Microsoft’s MAI training approach through Azure?

Microsoft uses Azure’s Foundry and Frontier Tuning products to let enterprise customers train models using their own data for dedicated models. Microsoft emphasizes that these models are trained on “clean, auditable enterprise-grade data” and “do not undergo third-party model distillation,” to address compliance concerns about the sources of training data.

Disclaimer: The information on this page may come from third-party sources and is for reference only. It does not represent the views or opinions of Gate and does not constitute any financial, investment, or legal advice. Virtual asset trading involves high risk. Please do not rely solely on the information on this page when making decisions. For details, see the Disclaimer.
Comment
0/400
No comments