The news
On August 11, 2026, Microsoft AI released MAI-Code-1.1-Flash, an in-house coding model, in GitHub Copilot and VS Code. Microsoft (MSFT) said MAI-Code-1.1-Flash writes better code than the model it launched at Microsoft Build in June, at a quarter of the cost.
GitHub's changelog put the saving at a 73% lower list price than MAI-Code-1-Flash, the earlier model. GitHub's pricing documentation lists the new model at $0.20 per million input tokens, $0.02 per million cached input tokens and $1.20 per million output tokens. LLM Stats, an independent model-tracking site, listed the predecessor at $0.75 and $4.50 per million input and output tokens. GitHub said annual subscribers are charged at a 0.25x premium request multiplier.
Microsoft said it focused on command-line and .NET work after developer feedback. It reported a 22% improvement on the Terminal-Bench 2.1 benchmark in GitHub Copilot CLI, a 15% improvement on .NET tasks, output that streams 25% faster and 25% fewer tokens per task. In production use, it said, code survival rose 4% and return visits rose 9%. The post did not define either metric.
According to GitHub, the model is rolling out to Copilot Free and Student users through automatic model selection, while Pro, Pro+, Max, Business and Enterprise subscribers can also pick it by hand. It works in Copilot CLI, the cloud agent, VS Code, Visual Studio, JetBrains IDEs, Eclipse, Xcode, GitHub Mobile and Copilot Chat on GitHub, and adds native vision support for reading images.
For Copilot Business and Copilot Enterprise, GitHub said administrators must enable a MAI-Code-1.1-Flash policy in Copilot settings, and that the policy is off by default.
The numbers
- List price cut vs MAI-Code-1-Flash (GitHub)
- 73%
- Input price per million tokens
- $0.20
- Output price per million tokens
- $1.20
- Premium request multiplier for annual subscribers
- 0.25x
- Terminal-Bench 2.1 improvement in Copilot CLI (Microsoft)
- 22%
- Fewer tokens per task (Microsoft)
- 25%
Why CEOs should care
For buyers of AI coding tools, the model list now matters as much as the seat price. GitHub's billing documentation says usage beyond plan allowances is billed in GitHub AI Credits, priced per token by model, so routing routine work to a cheaper model lowers the bill. Engineering leaders should find out which model automatic selection picks for their teams and decide when developers may choose pricier models.
CFOs should ask for token and credit reports by team and set spending limits before model changes shift costs without notice. Treat Microsoft's quality gains as the vendor's own measurements, and test the model on your own codebase before making it a default. CISOs get a built-in checkpoint, since the Business and Enterprise policy starts switched off. Microsoft said in June that the predecessor was trained without distillation from third-party models; ask whether the same holds for version 1.1 and what data-retention terms apply.
Boards should see the release as Microsoft building its own lower-cost models alongside the OpenAI and Anthropic models it also offers. That gives Microsoft more control over its costs, and it gives customers a new lever in price talks with every AI coding vendor.
The bigger picture
The release followed Microsoft's July 29, 2026 earnings call, where CEO Satya Nadella said millions of developers had used MAI-Code-1-Flash in GitHub Copilot, with higher code acceptance rates and 10% lower median token usage, while still having access to OpenAI and Anthropic models. He said GitHub Copilot had 50 million users and had introduced usage-based billing during the quarter.
Microsoft AI introduced MAI-Code-1-Flash on June 2, 2026, and claimed it beat Anthropic's Claude Haiku 4.5 on price-to-performance across coding benchmarks. The 1.1 release extends that pitch: smaller, cheaper in-house models for everyday coding, with frontier models from OpenAI and Anthropic still available in Copilot.
What happened next
On September 25, 2026, Microsoft introduced a redesigned Copilot that includes Code, a feature for building small tools from plain-language descriptions, and said agentic work, including Code and its new Autopilot agent, would be billed by usage.
As of September 29, 2026, GitHub's models-and-pricing page listed MAI-Code-1.1-Flash at $0.20 input, $0.02 cached input and $1.20 output per million tokens. The same page listed OpenAI's GPT-5.6 Luna at the same input and output rates and Anthropic's Claude Haiku 4.5 at $1.00 and $5.00, so buyers should compare quality as well as price.




