Skip to content
Innovation & Growth

Shopify builds copy-cat test stores so shopping AI agents can practice

ShopGym turns live storefronts into resettable sandboxes. The tests show how far a copy can stand in for the real thing.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
Shopify builds copy-cat test stores so shopping AI agents can practice

AI-generated image for WebPulse. About our images

In brief
  • Shopify Engineering describes ShopGym, which turns live storefronts into anonymized, resettable sandbox shops and generates shopping tasks so developers can test agents under consistent conditions.
  • In Shopify's study, agent results on synthetic shops tracked live-store results, most clearly on multi-step tasks. Some single tasks showed larger gaps, so the copies are not exact.
  • Ask any vendor selling shopping agents how they test them, and how far sandbox results sit from live-store results.

The hard part of building a shopping agent is not the shopping. It is getting the same store twice. Prices change, products sell out and pages get redesigned. If the exam keeps changing, nobody can tell whether the student improved. Shopify Engineering's new ShopGym project is built around that problem.

Why live stores make poor test rooms

Shopify describes a familiar task: an agent finds a jacket within budget, picks a size, checks the return policy and adds it to the cart. Testing that on live stores is hard, the engineers write, because products, prices and layouts keep changing. Sites may also block or slow automated visitors, which can cut a test run short. Stores built by hand give testers control, but they cover only a narrow slice of real shopping experiences.

ShopGym is Shopify's answer. It turns live storefronts into self-contained sandbox shops that can be reset, then generates shopping tasks tied to each shop's products, features and policies. Developers can reproduce a failure, compare agents and check whether a change helped.

224
Tasks in Shopify's validation study
Source: Shopify Engineering (October 1, 2026)
6
Sandbox shops tested: three synthetic, three built from real data
Source: Shopify Engineering (October 1, 2026)
7
Skill categories covered by generated tasks
Source: Shopify Engineering (October 1, 2026)

How a store becomes a test

The system has two parts. ShopArena builds the sandbox. ShopGuru writes the tasks. Each is a small group of coding agents that hand work to one another through shared files.

ShopArena starts with a live storefront's public surfaces: homepage, sitemap, search and cart endpoints, catalog and policy pages. A planner agent splits the work. Fresh agents with browser tools then explore the store and write short fragments describing it. Those fragments cover the shop's visual style, its features such as filters and pagination, and catalog statistics such as product counts and price spread.

The agents leave out brand names, product names, people's names and full web addresses as they write. Anonymity is built in rather than added afterward. A merging step combines the fragments into one readable specification. That specification is the only link between exploring and building.

The build step never sees the original store, only the specification. From it, the system first makes up a product catalog. It then assembles the site in stages. The first stage is the page frame and homepage. Collection pages, product pages, the cart, search and filters, and policy pages follow. A last pass checks that the parts work together.

Each stage repeats the same loop. A new agent edits the code. Rule-based checks confirm the build and the routes. A visual agent opens the shop in a browser and writes feedback. Using a new agent every round keeps its working context small. A failed feature can be repaired without rebuilding the shop.

ShopGuru then writes tasks. Rule-based generators produce simple ones, such as finding a product or a return policy. An LLM writes multi-step journeys, such as filtering a collection by color, sorting by price and adding an item to the cart. Automated checks confirm that every product, collection and filter a task mentions exists in the shop. Failing tasks go back to the LLM for revision.

What the results show, and what they do not

To compare structure, Shopify drew each shop as a graph. Every screen or open panel, such as a menu or cart drawer, became a point. Every action that moves between them became a link.

The copies had about as many distinct states as the originals, and their pages were about as complex. They had fewer links between states. Shopify attributes much of that to missing external links, marketing pages and other side routes.

Shopify then ran agents through two test setups. One is BrowserGym, where the agent reads a structured outline of each page (an accessibility tree). The other is Shopify's own setup, which also gives the agent screenshots. Agent performance on synthetic shops positively correlated with performance on the live stores they mirror. Results lined up best on multi-step journeys. There, both setups also ranked the models in the same order as on live stores.

The limits are in the report too. Gaps were wider on some single-skill tasks, especially in Shopify's internal setup. Shopify acknowledges that a twin is not a perfect copy of how hard the real store is. The study is also small: six shops, with three built on real data. Shopify says work on using these environments to train agents is ongoing, and the report does not include training results.

The lesson: the test is now part of the product

This is one company's engineering report, not an industry trend. But it shows an idea worth holding. If agents are graded in copies of stores, whoever builds that exam decides what counts as a successful purchase.

The sandbox omits the clutter of external links and marketing pages. That makes it a cleaner store than the real one. A fair question is how an agent that passes the clean copy behaves on the messy original. The report does not answer that, and the single-task gaps suggest it matters.

For a store owner, the task design is also informative. ShopGuru grounds each task in the catalog, navigation, policies and interaction options. In our view, that means an easy-to-find return policy, working filters and sorting are exactly what an agent test would probe. Plain storefront hygiene can be measured.

Questions to put to your team

If you buy or build shopping agents, ask how they are tested. Can a failed run be reproduced exactly? Are results reported for multi-step journeys, single tasks, or both? What is the measured gap between sandbox and live-store results? If you run a store, ask whether a person or an agent could find your return policy, filter by size and complete checkout without help. Shopify's report also does not say how its seed storefronts are chosen. That is worth asking any vendor who builds sandboxes from live shops.

A fair test is not a real store. It is a repeatable one that tells you, with stated limits, how the real one will behave.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: Shopify Engineering.

Share this insight