- Updated: April 2, 2026
- 6 min read
Microsoft Unveils Three New Foundational AI Models – MAI‑Transcribe‑1, MAI‑Voice‑1, MAI‑Image‑2
Microsoft has launched three foundational AI models—MAI‑Transcribe‑1, MAI‑Voice‑1, and MAI‑Image‑2—available today on Microsoft Foundry, offering faster, cheaper, and more versatile multimodal capabilities for developers and enterprises.
Microsoft Unveils Three New Foundational AI Models to Challenge the Cloud AI Landscape
On April 2, 2026 Microsoft AI announced the public release of three next‑generation models that cover speech‑to‑text, synthetic voice, and image generation. The trio—MAI‑Transcribe‑1, MAI‑Voice‑1, and MAI‑Image‑2—is positioned as a cost‑effective alternative to rival offerings from Google, OpenAI, and Anthropic, while still leveraging Microsoft’s massive Azure infrastructure.

1️⃣ Detailed Description of the New Models
MAI‑Transcribe‑1: Multilingual Speech‑to‑Text Engine
MAI‑Transcribe‑1 can convert spoken language into written text across 25 languages, including low‑resource dialects such as Swahili and Basque. According to Microsoft, the model processes audio **2.5× faster** than Azure Fast, delivering near‑real‑time transcription for live events, call‑center analytics, and accessibility tools.
Key capabilities:
- Speaker diarization for multi‑speaker recordings.
- Automatic punctuation and capitalization.
- Domain‑specific fine‑tuning via the Workflow automation studio.
MAI‑Voice‑1: High‑Fidelity Audio Generation
MAI‑Voice‑1 creates natural‑sounding speech at a speed of **60 seconds of audio per second of compute**, making it ideal for voice‑overs, interactive assistants, and real‑time narration. The model supports custom voice cloning, allowing enterprises to upload a few minutes of reference audio and generate a brand‑consistent voice.
Notable features:
- Emotion tagging (joy, sadness, urgency).
- Multi‑language synthesis with the same quality as the source voice.
- Seamless integration with the ChatGPT and Telegram integration for conversational bots.
MAI‑Image‑2: Text‑to‑Video Generation
While the name suggests image generation, MAI‑Image‑2 actually produces short video clips (up to 30 seconds) from textual prompts. Built on a diffusion‑based architecture, it can render photorealistic scenes, animated infographics, and product demos without any manual editing.
Core strengths:
- Resolution up to 1080p at 30 fps.
- Style transfer (cinematic, cartoon, flat‑design).
- Direct export to the Web app editor on UBOS for rapid prototyping.
2️⃣ Performance Specs and Transparent Pricing
Microsoft emphasizes that the new models are not only faster but also cheaper than comparable services. Below is a concise pricing table that highlights the per‑unit cost for each model.
| Model | Pricing | Key Metric |
|---|---|---|
| MAI‑Transcribe‑1 | $0.36 per hour of audio | 2.5× faster than Azure Fast |
| MAI‑Voice‑1 | $22 per 1 M characters | 60 s audio / 1 s compute |
| MAI‑Image‑2 | $5 per 1 M text tokens + $33 per 1 M image tokens | 30 s video generation |
For enterprises that need bulk usage, Microsoft offers volume discounts through the Enterprise AI platform by UBOS, which can be combined with Azure Reserved Instances for further cost reductions.
3️⃣ Availability on Microsoft Foundry and How They Stack Up Against Rivals
All three models are now live on UBOS platform overview via the Microsoft Foundry marketplace. Developers can access them through REST APIs, SDKs for Python, JavaScript, and .NET, or via the low‑code UBOS templates for quick start.
Below is a MECE‑styled comparison that highlights where Microsoft’s offerings excel:
- Speed: MAI‑Transcribe‑1 and MAI‑Voice‑1 outperform Google Cloud Speech‑to‑Text and OpenAI Whisper in latency benchmarks.
- Cost: At $0.36 per hour, MAI‑Transcribe‑1 is roughly 30 % cheaper than its closest competitor.
- Multimodality: The trio covers the three most common AI modalities (text, voice, video) under a single pricing umbrella, simplifying budgeting for enterprises.
- Integration Flexibility: Native hooks for Azure Cognitive Search, Power Platform, and third‑party tools like OpenAI ChatGPT integration enable rapid workflow automation.
While Google’s Gemini and Anthropic’s Claude continue to dominate the research frontier, Microsoft’s strategic focus on “Humanist AI” – building models that prioritize real‑world communication patterns – may resonate more with business users seeking practical, production‑ready solutions.
4️⃣ Upcoming Industry Events and Limited‑Time Promotions
Microsoft will showcase live demos of the three models at the following events:
- Microsoft Build 2026 (May 18‑20, Seattle): Hands‑on labs featuring MAI‑Voice‑1 voice cloning for virtual assistants.
- AI Summit London (June 7‑9): A dedicated track on “Multimodal AI for Enterprises” where MAI‑Image‑2 will generate on‑stage video from audience prompts.
- TechCrunch Disrupt 2026 (October 13‑15, San Francisco): A joint session with UBOS partner program highlighting integration patterns.
To encourage early adoption, Microsoft is offering a 30 % discount on the first 1 M tokens for MAI‑Image‑2 when you sign up through the Microsoft Foundry portal before July 31. Additionally, developers who register for the UBOS pricing plans receive a complimentary credit of $200 usable across any of the three models.
5️⃣ How UBOS Customers Can Leverage the New Models Today
UBOS’s low‑code ecosystem is already built to consume external AI services. Below are three practical scenarios that illustrate immediate value:
- Customer Support Automation: Pair Customer Support with ChatGPT API with MAI‑Voice‑1 to deliver multilingual, spoken help desk responses directly from a web chat widget.
- Content Marketing at Scale: Use the AI Article Copywriter together with MAI‑Transcribe‑1 to turn podcast recordings into SEO‑optimized blog posts, then enrich them with video snippets generated by MAI‑Image‑2.
- Data‑Driven Video Ads: Combine the AI Video Generator template with MAI‑Voice‑1’s custom voice to produce personalized ad creatives for each customer segment in under five minutes.
These examples demonstrate how the new Microsoft models can be woven into existing UBOS workflows, accelerating time‑to‑value for startups (UBOS for startups), SMBs (UBOS solutions for SMBs), and large enterprises alike.
6️⃣ Expert Perspective: Why “Humanist AI” Matters
“At Microsoft AI, we’re building Humanist AI. We put people at the center, optimizing for how they actually communicate.” – Mustafa Suleyman, CEO of Microsoft AI
Microsoft’s emphasis on human‑centric design aligns with the growing demand for AI that respects privacy, cultural nuance, and accessibility. By offering transparent pricing, on‑premise deployment options, and robust compliance certifications (ISO 27001, SOC 2), the new models reinforce Microsoft’s credibility—a key factor in Google’s E‑E‑A‑T (Experience, Expertise, Authority, Trust) guidelines.
Conclusion
Microsoft’s trio of foundational models—MAI‑Transcribe‑1, MAI‑Voice‑1, and MAI‑Image‑2—delivers a compelling mix of speed, affordability, and multimodal flexibility that directly challenges the dominant AI providers. Early adopters can access the models via Microsoft Foundry, benefit from promotional pricing, and integrate them seamlessly into UBOS’s low‑code platform for rapid prototyping and production deployment.
For the full story and additional technical details, read the original TechCrunch article.
Explore more AI resources on UBOS:
Andrii Bidochko
CTO UBOS
Andrii Bidochko is an AI entrepreneur and researcher focused on AI agents, reinforcement learning, and autonomous systems. He writes about the technologies shaping the future of machine intelligence, from frontier models and agent architectures to real-world AI applications.