Microsoft stands at a pivotal point with Windows. The company has spent years layering artificial intelligence into its flagship operating system. Now the focus sharpens on developers. They gain direct access to powerful local models. They run them without cloud dependency. And they do it on the hardware already sitting on their desks.
The Verge first outlined the scope in a report published hours before the conference kicks off. Microsoft will unveil new AI models in Windows, a new reasoning model from Microsoft AI, and a Copilot “super app,” according to sources. The story also flags a new Windows 11 developer optimized experience. It promises a distraction-free environment stocked with pre-installed apps, tools and scripts. Pavan Davuluri, corporate vice president for Windows and devices, teased it simply: “something new is coming for developers.”
From Copilot Runtime to a Full AI Foundry
That something builds on announcements from Build 2025. Then Microsoft renamed and expanded its Copilot Runtime into Windows AI Foundry. The platform now supports the entire model lifecycle. Developers select, optimize, fine-tune and deploy across client hardware or cloud. Windows ML forms its foundation. It serves as the built-in inferencing runtime. No extra drivers or runtimes to bundle. It runs models across silicon from AMD, Intel, NVIDIA and Qualcomm. CPU. GPU. NPU. All of them.
Foundry Local pulls in open-source models from catalogs including Ollama and NVIDIA NIMs. Developers browse, download and test them in minutes. The system auto-detects available hardware acceleration. A CLI tool installs via WinGet. An SDK lets apps call the models directly. Microsoft detailed the approach in its Windows Developer Blog post from May 2025.
Documentation on Microsoft Learn brings the picture up to date as of April 2026. Microsoft Foundry on Windows offers ready-to-use APIs for tasks that once required custom work. Text summarization. Image description. OCR. Semantic search over local files. Video super-resolution. Speech recognition via Whisper models. All of it runs locally. Many features target Copilot+ PCs with their dedicated NPUs. Yet support stretches back to Windows 10 machines for broader reach. The overview page lays out a clear decision tree: start with the inbox APIs if they fit. Fall back to Foundry Local for LLMs or voice. Turn to full Windows ML when you need to bring your own model from Hugging Face or train one yourself.
Performance matters here. Models run without sending data off-device. Privacy improves. Latency drops. Cost stays flat since there are no API calls to pay for. Developers can fine-tune small language models like Phi or Mistral using the Foundry Toolkit extension for Visual Studio Code. The extension provides a playground, REST testing endpoint and cloud or local fine-tuning options. Quantization tools prepare models for NPU use. DirectML handles the hardware acceleration underneath.
AI Dev Gallery complements the toolkit. This open-source Windows app ships more than 25 interactive samples. It lets developers explore, download and run models from Hugging Face. Source code exports easily into their own projects. Microsoft positioned it as the practical playground for experimentation. The gallery also demonstrates Windows ML in production scenarios now that the runtime reached general availability in September 2025.
But. Local AI alone does not tell the full story. Microsoft also pushes agentic capabilities. Model Context Protocol support arrived in preview. It lets AI agents interact with native Windows apps. Apps expose their functions. Agents call them. Partners including Anthropic, Perplexity and OpenAI praised the move in the 2025 blog. Kevin Weil, chief product officer at OpenAI, said the protocol “paves the way for ChatGPT to connect to Windows tools and services that users rely on every day.”
App Actions add another layer. Developers define actions their apps can perform. The system surfaces them to agents or users. A testing playground helps validate them. Early adopters include Zoom, Filmora and Goodnotes. Security received attention too. VBS Enclave SDK creates trusted execution environments. Post-quantum cryptography support strengthens data protection. These features address enterprise concerns about running AI at the edge.
The developer experience extends beyond models. Microsoft open-sourced Windows Subsystem for Linux. Community contributions can now improve performance and add features. WinGet Configuration simplifies environment setup with declarative scripts. PowerToys gained a command palette. A new command-line editor called Edit joined the insiders program. Advanced Windows Settings centralizes customization, including deeper GitHub integration in File Explorer.
All this sets the stage for Build 2026. The conference opens Tuesday in San Francisco. Satya Nadella will deliver the keynote. Expect details on the developer optimized mode. It could rewrite parts of Windows 11 for better performance and reduced distractions. Customization options shown in recent previews hint at a more tailored interface for coders. Local model improvements will likely dominate. The next generation of Microsoft’s smaller models, including the MAI-Thinking-1 reasoning model from Mustafa Suleyman’s team, should appear. Unlike some prior work, this one avoids distillation techniques. It targets enterprise needs.
A Copilot super app also looms. It would merge various Copilot experiences into one interface. A mockup leaked last week. The full preview arrives later in summer. Early looks at Microsoft Scout, an AI agent drawing on OpenClaw research, may surface as well. GitHub Copilot faces pressure from rivals like Cursor and Claude. New coding models could help Microsoft regain ground. Reuters reported on May 28 that the company plans to release a homegrown coding model at the event along with models for transcription, reasoning, speech and images.
PCMag previewed the conference four days ago. It noted the shift to a smaller San Francisco venue focused on AI developers and technical leaders. The article captured the mood: flashy hardware announcements may take a back seat to deeper platform work.
Windows ML now ships in the Windows App SDK. Developers can deploy production apps without packaging extra dependencies. Model catalog features allow sharing optimized models across applications. Semantic search APIs support retrieval-augmented generation over local content. Fine-tuning with LoRA keeps resource use low even on laptops. These advances arrived steadily since last year’s conference. They reflect a deliberate shift from cloud-first AI to hybrid approaches that respect device capabilities.
Yet challenges remain. Not every PC carries a powerful NPU. Performance still varies. Developers must test across hardware configurations. Responsible AI practices receive strong emphasis in the documentation. Microsoft urges teams to consider bias, transparency and data handling before shipping features. The guidance appears prominently on the Foundry overview page.
So the picture sharpens. Microsoft no longer treats Windows as just a consumer OS with some AI sprinkled on top. It positions the platform as a first-class environment for building AI applications. Local inference. Agent integration. Streamlined tooling. A dedicated developer mode. The combination could attract teams wary of cloud costs or data residency rules. It could also keep Windows relevant as AI workloads move closer to the user.
Build will fill in the blanks. New model releases. Deeper details on the optimized developer experience. Updates to the Foundry stack. Expect practical demos, not just slides. The audience of developers and technical leaders wants code they can run on Monday morning. Microsoft appears ready to deliver it. The next few days will show exactly how far that ambition reaches.


WebProNews is an iEntry Publication