This is the Trace Id: 2ef451d12231750698d4446f6c516fa1
10/07/2026

From Plain English to Passing Tests: Momentic on Microsoft Foundry with Claude

AI coding assistants generate code faster than teams can verify. Legacy frameworks (Selenium, Cypress, Playwright) rely on brittle selectors that break with every UI change, so teams maintain tests instead of shipping. Reliability stalls near 95%, and false failures teach engineers to ignore alerts.

Momentic AI lets engineers write tests in plain English. Locator, assertion, visual, and extraction agents call Claude Opus and Sonnet 4.6 plus OpenAI models via Foundry, run on Azure, and report through GitHub Actions and an MCP server for Copilot, Claude Code, and Cursor.

Momentic runs 200M+ test steps a month for 2,600+ users at 99.2% reliability, stopping 390,000+ bugs from reaching production monthly. Teams automate about 70% faster than on legacy frameworks. On Microsoft Foundry, onboarding a new model provider dropped from weeks of billing, access, and rate-limit tuning to two or three hours.

Momentic

The verification gap

Software teams have never shipped code faster. AI coding assistants—GitHub Copilot, Claude Code, Cursor—now write, refactor, and review large portions of a codebase, and the volume of change reaching a repository each week has climbed to match. Verification has not kept pace. Every additional pull request creates additional surface area that has to be exercised, and the test layer most engineering organizations depend on was built for a slower era.

That layer runs on frameworks like Selenium, Cypress, and Playwright, which find elements on a page using CSS selectors and hard-coded paths. Those frameworks support a wide range of testing patterns, but where tests depend on selectors, those references can break when a designer moves a button or a developer renames a class. Teams in that position can spend significant engineering time repairing tests instead of writing features, and the tests that survive can be flaky. In Momentic's own experience with customer test suites, end-to-end reliability commonly sits near 95 percent—an internal estimate rather than an independent industry study, and a number that sounds high until a team runs thousands of tests a day and quietly learns to ignore the failures.

The gap widens as the coding tools get better. More code, written faster, against a verification layer that was already the slowest part of the pipeline.

Plain-English tests that adapt to UI changes

Wei-Wei Wu and Jeff An founded Momentic AI in San Francisco in late 2023 to close that gap. The company came out of Y Combinator’s W24 batch and has raised $19.2 million to date—a $3.7 million seed, followed eight months later by a $15 million Series A announced in November 2025, led by Standard Capital with participation from Dropbox Ventures, Y Combinator, FCVC, Transpose Platform, and Karman Ventures.

Momentic replaces selectors with intent. An engineer describes a flow in ordinary language, and the platform works out what the test is meant to prove, drives the browser to prove it, and updates its own element references when the interface underneath changes, with the updates surfaced for the team to review.

Underneath that interface is a multi-agent pipeline. A plain-English test is decomposed and routed to specialized components, each handling one part of the job: one for locating elements, one for evaluating assertions, one for interpreting visual state, one for extracting data. Those components reason over multi-modal signals including screenshots, the accessibility tree, and network traffic, rather than a single fragile string. Tests execute on Azure infrastructure and report back into the customer’s delivery pipeline through GitHub Actions, or through Momentic’s Model Context Protocol server, which wires the platform directly into GitHub Copilot, Claude Code, and Cursor.

That last piece is the one that closes the loop. An agent writing code can call Momentic to run the team's tests against its own work before a human opens the pull request. The tests, permissions, and pass criteria are configured by the engineering team, and results return to that team for review—Momentic reports on the code, it does not approve or merge it.

Why Claude, and why Microsoft Foundry

A platform that routes work across several model families pays a tax every time it adds one. Each new provider brings its own billing, its own access control, its own rate limits, and a stretch of engineering time spent wiring all of it up before a single test benefits. For a company Momentic’s size, that overhead competes directly with product work.

“In the past we would spend weeks tuning things like billing, access control, rate limits for every new model we onboarded. Foundry makes it easy to set up a new provider in a matter of 2 to 3 hours. We found that Claude’s Opus series were particularly suited to these tasks. Now we’re serving millions of tokens per minute through Foundry.” says Jeff An, Co-Founder and CTO of Momentic AI.
Weeks to a few hours changes what a small team can afford to try. Momentic’s agents now call Claude Opus and Claude Sonnet 4.6 as their primary models, alongside OpenAI models, all through Microsoft Foundry—one enterprise-grade service carrying both model families, with the applicable Azure service-level agreements and compliance certifications, and capacity that has sustained Momentic's workload of hundreds of millions of inference calls a month.

Azure’s compliance certifications do commercial work too. Momentic reports that building on a platform its enterprise prospects already have under review has helped shorten security and procurement conversations that would otherwise sit between a trial and a contract. Customers remain responsible for their own vendor, security, and procurement requirements.

Jeff An, Co-Founder and CTO, Momentic AI

“In the past we would spend weeks tuning things like billing, access control, rate limits for every new model we onboarded. Foundry makes it easy to set up a new provider in a matter of 2 to 3 hours. We found that Claude’s Opus series were particularly suited to these tasks. Now we’re serving millions of tokens per minute through Foundry.”

Jeff An, Co-Founder and CTO, Momentic AI

200 million test steps a month

Momentic now processes more than 200 million test steps every month for over 2,600 users, and prevents more than 390,000 bugs from reaching production each month. Test reliability runs at 99.2 percent, and false positives are down 99 percent—the difference between a signal engineers act on and one they mute. Teams reach a working automated suite roughly 70 percent faster than they would on a legacy framework.

Individual customers report similar shifts: manual QA checklists largely replaced by automated runs, and daily test cycles that previously took hours completing in minutes. Because the tests are written in plain language, QA staff without a coding background can author and review many tests without waiting on an engineer, which widens who in an organization can contribute to quality.
CB Insights has named Momentic a Challenger in the test design automation market, alongside companies including mabl, Leapwork, and Autify.

Wei-Wei Wu, Co-Founder and CEO, Momentic AI

“Momentic is critical infrastructure for most of our customers. We’re blocking deploys. We’re blocking releases. We’re able to achieve that reliability and uptime with Microsoft Azure Foundry as our main provider.”

Wei-Wei Wu, Co-Founder and CEO, Momentic AI

The verification layer for AI-era software

Momentic’s ambition is to sit alongside the coding tools rather than behind them, to be the layer that checks what AI writes, at the speed AI writes it. The MCP server is the clearest expression of that: verification becomes something an agent can invoke mid-task, not a stage a human remembers to run afterward.

The bet underneath is that as more code is generated, trust in that code becomes the scarce resource. A team that ships far more often is only better off if it can still tell, quickly and reliably, whether what it shipped works. That puts Momentic in an unusual position for a young company: directly in the path of its customers’ releases, where being unavailable is not an option.

“Momentic is critical infrastructure for most of our customers. We’re blocking deploys. We’re blocking releases. We’re able to achieve that reliability and uptime with Microsoft Azure Foundry as our main provider.” says Wei-Wei Wu, Co-Founder and CEO of Momentic AI.

Learn more about Momentic AI at momentic.ai.

Take the next step

Fuel innovation with Microsoft

Explore more customer stories

Find out how customers are achieving more with Microsoft products and solutions.
A man wearing headphones and smiling.

Talk to an expert about custom solutions

Let us help you create customized solutions and achieve your unique business goals.
Three people in a meeting room.

Transform work with Microsoft AI

Bring intelligence into the flow of work and help your organization achieve its goals with secure, scalable AI solutions.

Follow Microsoft