C:\CHANGELOG> type v1-38-0-the-price-is-temporary-the-integration-is-not.md
v1.38.0 · released · 5 min read · by

The price is temporary. The integration is not.

Microsoft AI on September 3 released MAI-Transcribe-2, a speech-recognition model the company says is faster, more accurate, and cheaper than anything OpenAI, Google, or ElevenLabs currently sells, then priced the thing at 10 cents per hour of audio. When Microsoft AI shipped the first model in this line just five months ago, it charged $0.36 an hour; the early-bird price cuts that by roughly 72%. The $0.10-per-hour rate is an introductory price that runs through the end of 2026. The company did not explain what the price would jump to in January. That is the entire problem. The model is optimized for the demo. The integration is optimized for three years of stable invoices. Those two time horizons do not overlap.

I run transcription on thepunchlist.ai, the construction dispute-resolution system. Contractors upload voicemails, recordings, and site-walk audio. The system transcribes the files, timestamps every claim, extracts named parties, and generates a structured dispute log. The feature has been in production for eight months. It calls Deepgram. The bill is predictable. I could switch to MAI-Transcribe-2 and cut the bill by half. But I cannot switch to MAI-Transcribe-2 and assume the bill stays cut by half, because the $0.10-per-hour offer is scheduled to expire at the end of 2026, and Microsoft has not disclosed what the model will cost afterward. Microsoft Learn classifies MAI-Transcribe as a public preview without a formal service-level agreement and does not recommend it for production workloads. The API exists. The guarantee does not.

The benchmarks are strong. The contract is weak.

The model ranks first on the FLEURS benchmark across 60 languages with an average Word-Error-Rate of 5.2%, defines the Pareto Frontier for accuracy and latency on Artificial Analysis, and ranks second on the Artificial Analysis Word-Error-Rate leaderboard. All this performance also comes at the best price on the market, just $0.10 per hour of audio. The announcement calls it the fastest, most accurate and cheapest speech recognition model in the world. That is marketing. It is also, for the moment, defensible marketing.

On Artificial Analysis's non-streaming word-error-rate leaderboard it sits at 2.0% AA-WER, in second place, and at roughly 411× real time, also second, which together put it on the accuracy-speed Pareto frontier: on that board, no model is both more accurate and faster. MAI-Transcribe-2 is designed to take on real-world applications with faster inference with substantially lower latency, especially for long-form audio, with up to 10× faster processing than leading competitors; speaker diarization distinguishes between speakers and attributes words to the right person within a recording, word-level timestamps provide precise timing for every word. These are the features a production system needs. The promotional price is not one of them.

The main uncertainty is economic rather than technical: $0.10 per hour is explicitly a limited-time offer through the end of 2026, and Microsoft has not published the permanent price. That is a cost structure you can test. It is not a cost structure you can budget. A promotional price is a pricing experiment until it is not: teams that build a product on the $0.10 rate should know the standard rate is undisclosed, and the honest planning number for a 2027 budget is somewhere between the launch offer and whatever Microsoft names when the offer expires. I have written before about bills that drop 72% with no warning. A bill that drops 72% and then climbs back to an undisclosed number is worse. The uncertainty compounds in both directions.

The migration window is three months. The model will not be.

MAI-Transcribe-1 launched in April and was already deprecated by August — five months. That is not a model family. That is a release cadence optimized for benchmark farming. Microsoft shipped three transcription models in five months. Each one deprecated the previous version. Each one changed the price. In April 2026, Microsoft released MAI-Transcribe-1, which supported 25 languages; this model was deprecated toward the end of last month. The API exists to call today. It may not exist to call in six months. The promotional price exists to bill today. It will not exist to bill in four.

The construction ERP at coenconstruction.com processes payroll calls, site instructions, and change-order disputes. The transcription runs every morning. The feature has been stable for two years. I could rip out Deepgram, point the webhook at MAI-Transcribe-2, and cut the bill in half before October. If the January price doubles the September price, the net savings vanish. If the January price holds the September price, every competitor matches it by March. If the model gets deprecated in Q2 the way its two predecessors did, I migrate three times in eight months. That is not a cost optimization. That is a vendor integration treadmill.

The model will keep getting better. The invoice will keep getting shorter. The small print will keep saying the price is temporary and the SLA does not apply.

Whether the price holds, or whether competitors match it before the year ends, will determine if this is a permanent reset of transcription costs or a temporary land grab; either way, the floor for what a client expects to pay for a plain transcript has moved, and it is unlikely to move back. That is the real shift. Microsoft did not just ship a cheaper model. It shipped a cheaper model with an expiration date printed in the announcement. Every procurement conversation for the next four months will anchor on ten cents an hour. Every budget forecast for next year will carry a question mark where the transcription line used to be. The invoice volatility is now a feature of the API landscape, not a bug in one vendor's pricing.

I am not switching to MAI-Transcribe-2. The benchmarks are real. The price is unreal. I need a model I can call in March and bill in April and forecast in June without checking whether the vendor deprecated it, doubled it, or pulled the SLA while I was shipping something else. The promotional price optimizes for the demo. The production integration optimizes for the invoice three years from now.


— Cole Ciprari · Business Systems Architect · Worcester, MA
my résumé is an operating system → ciprari.ai · linkedin.com/in/coleos · cole@ciprari.ai
WAS THIS ANY GOOD?
Anonymous, one tap, no account. Tap again to undo.
▚▞ GET THE NEXT RELEASE
A new release every day, plus the Sunday Rollout Report — the week's AI and tech news, summarized by a human with production access. No spam. Unsubscribe by emailing a mildly disappointed cole@ciprari.ai.
PHOSPHOR