Skip to content

Articles / Research and Testing

What is iterative testing?

What we found

Iterative testing means each A/B test starts from what the test before it taught. In our tests, star ratings on every product card moved conversion 15.6% without a clear result, and the next test, 1 star rating at the top of the page, lifted conversion 11.3%.

Iterative testing means each A/B test starts from what the test before it taught you. You run a test, read what moved and by how much, and use that to design a sharper next test on the same page or the same idea. A loss or a flat result is not the end of an idea. It is the starting point for the next round.

Most pages about iterative testing come from product teams and describe a process. This one stays on A/B tests. Below are 2 real test lineages from our archive. In each, a 1st attempt lost or came back flat, and the next test used what it taught. You will also see the rules we use to decide what to test next, and how iterative testing differs from iterative design.

An A/B test splits traffic between the current page and a changed version, then measures which one earns more conversions. On its own, a test answers 1 question: did this change, built this way, beat the original?

Iterative testing links tests into a chain. The answer from the 1st test shapes the question for the 2nd. Over a few rounds, you learn more than whether an idea works. You learn where it works, in what form, and for which shoppers.

The loop has 4 steps:

  1. Test. Run a change against the original for at least 2 full weeks.
  2. Read the signal. Look at how far the numbers moved, on conversion rate and on revenue per visitor.
  3. Refine. Keep the part of the idea that showed a signal. Change where it sits, how it looks, or who sees it.
  4. Test again. Run the refined version against the original, and repeat.

You may also see this called evolutionary testing, as opposed to revolutionary testing. The terms help set scope. An evolutionary test changes 1 part of a page that already works. A revolutionary test replaces the whole page with a redesign. Iterative testing is the evolutionary kind, done on purpose and in order.

1 A/B test answers 1 question at a time

A single test result is easy to over-read. Say you add star ratings to a page and conversion drops. Does that mean shoppers do not care about ratings? Or do they care, but the ratings were in the wrong place, the wrong size, or competing with something else on the page?

1 test cannot tell those apart. A 2nd test, designed around the 1st result, often can. That is the core value of iterative testing. Each round narrows the list of possible reasons until the result makes sense.

It also changes how you treat a loss. In a program of 1-off tests, a losing test is a dead end and the idea gets dropped. In an iterative program, a losing test that moved the number a lot is a strong lead. Shoppers reacted. The job now is to find the version they react to well.

Star ratings on every product card moved conversion 15.6% without a clear result

Lineage A started on 1 store’s collection pages, where shoppers browse a category and pick products to look at. The idea was a common one: show each product’s star rating in the listing, so shoppers can see what other buyers liked before they click.

We tested star ratings in the product listings against the original product cards, which had no stars.

Evidence · Collection page · Not significant

Do star ratings on collection listings lift conversion?

15 days · responsive · We added star ratings under the first 16 products on collections.

Control
Control wireframe
Var A: Star ratings
Var A: Star ratings wireframe

-15.6% conversion, -27.0% revenue per visitor

Read the full test

Shoppers leaned toward the original cards without stars. The version with a rating on every card trailed by 15.6%, but the gap was not large enough to call a result, so the test was not significant.

A result like that is easy to file as “ratings do not work here.” That would have been the wrong lesson. A 15.6% move is large. Something about the ratings mattered to these shoppers. What the test could not say was whether they wanted fewer ratings, a different kind of rating, or the rating in a different place.

1 possible reason: a rating on every card adds a lot to a page. Each product tile gets another line to read, and shoppers can start comparing stars instead of products.

1 star rating at the top of the collection page lifted conversion 11.3%

The next test on the same store kept the idea of ratings and changed where and how they appeared. Instead of a rating on every product, it showed 1 rating at the top of the collection page. It was the same rating shoppers had already seen in the homepage hero, so the page repeated a trust signal they knew.

The test also tried a 2nd way to build trust at the top of the page. Variation A showed a hand-picked customer review with the buyer’s photo. Variation B showed the star rating.

Evidence · Collection page · Won

Does a star rating at the top lift collection page conversion?

33 days · responsive · We added social proof to the top of 5 busy collection pages.

Control
Control wireframe
Var A: Customer review with photo
Var A: Customer review with photo wireframe
Var B: Site star rating
Var B: Site star rating wireframe

+11.3% conversion, -7.5% revenue per visitor

Read the full test

The star rating won. Conversion rose 11.3% over 33 days, across all devices. 1 quick rating, placed early, built more trust than a rating on every card, and more than a single review that shoppers had to stop and read.

Put the 2 tests side by side and the lesson gets clearer. These shoppers did care about ratings. They responded to 1 clear rating near the top that matched what they had already seen. The 1st test found the signal. The 2nd test found the form that worked.

1 of 2 variations was paused partway to speed up the star rating test

Iteration does not only happen between tests. It can happen inside 1.

In the star rating test, the customer review (Variation A) showed little effect as the data came in. Rather than keep splitting traffic 3 ways, we paused it partway through. Its share of the traffic went to the star rating (Variation B), so that variation could reach a clear result sooner.

This works when 1 variation is clearly going nowhere and the others are still in the race. It is not a reason to cut a test short. The star rating test still ran for 33 days, well past our 2-week minimum, and the call was made on full data.

The habit worth copying is to read each variation as its own idea. A test with 2 variations is really 2 small experiments running at once. When 1 has clearly stalled, give its traffic to the 1 that still has something to prove.

Mobile shoppers preferred a 2-across grid by 9.4%

Lineage B started with a layout idea on mobile collection pages. The pages showed products in a grid, 2 across. We tested 1 product per row at full width, so each product card would be bigger and easier to see.

Shoppers preferred the grid. In the single product rows test, the 2-across grid out-converted single rows by 9.4% over 14 days and earned 7.2% more revenue per visitor. Seeing more products for the same amount of scrolling helped shoppers browse. We paused the test and kept the grid.

In a program of 1-off tests, that is where the story ends. But a 9.4% move is a strong signal. Shoppers clearly noticed how the product cards were laid out. The next question was whether a better-built full width card could keep what the grid did well. Our guide What is ecommerce conversion rate optimization? covers this test alongside 3 others on product, collection and checkout pages.

A refined full width product card closed the gap to +0.5%

The 2nd attempt in Lineage B took what the grid test taught and tried a refined version of the full width product card.

Evidence · Collection page · Not significant

Does a full width product card lift mobile collection page sales?

16 days · mobile · We showed 1 product per row in wide cards on the mobile collection page.

Control
Control wireframe
Var A: Single column
Var A: Single column wireframe

+0.5% conversion, +4.4% revenue per visitor

Read the full test

This time the result was flat. The refined card came in 0.5% above the original, far too small to call a win.

That still counts as progress. The 1st version trailed the grid by 9.4%. The refined version matched the original. The refinements earned back the full gap, and the store learned that a full width card can hold its own when it is built with browsing in mind. What it did not do was beat the original, so there was no conversion reason to change the live page.

A flat result like this has another use. I keep a change that does no harm when it serves the brand or the business direction more than conversion. If the store wants larger product images for brand reasons, the refined card is now an option backed by data.

A losing test that moved the number is not a dead end. It tells you where to look next.

The star rating win lifted conversion 11.3% and lowered revenue per visitor 7.5%

A win does not end a lineage either. It usually opens the next question.

The star rating test lifted conversion 11.3%. But revenue per visitor fell 7.5%. More shoppers bought, and each one spent less on average. The test does not tell us why.

That gap is where the next round starts. Did the rating bring in more shoppers ready to buy 1 lower-priced product? Did it speed up decisions so shoppers added fewer items to the cart? Each of those is a hypothesis, and each points to a different test.

This is why we judge every test on 2 numbers. If we had read conversion alone, the star rating test would look like a clean win with nothing left to learn. Reading revenue per visitor too turns the win into the start of the next test.

3 rules decide what to test next after a loss, a flat result or a win

Every result points somewhere. Whatever the result, I like 2 weeks of data before calling a test, so the read covers both weekday and weekend shoppers. Then these 3 rules decide where the next test goes.

What we have seen across tests PATTERN, NOT A SINGLE TEST

A big move is a signal, even when the test does not win.

How far a tested element moves the numbers tells you how much it matters to shoppers. In 2 lineages, a 15.6% move with no clear result and a 9.4% loss each led to a 2nd test: 1 won 11.3%, and the other closed the gap to +0.5%.

1. After a big move, take 2 different directions

We look for the signal: how much movement we saw on the tested element. If the movement is high, we have a signal that the element is important to the user. Generally we take 2 different directions for the next test.

Both lineages above follow this. Ratings on every card moved the number 15.6%, so the next test tried ratings in a new form and place. It also tried 2 directions at once: a single customer review in Variation A and a single star rating in Variation B. Single product rows moved the number 9.4%, so the next test tried a refined version of the same idea.

2. After a win, check it up and down the funnel

When we see a winning signal in a funnel, we check whether that signal holds further up and further down the funnel. A trust signal that wins on a collection page may also help on the homepage, the product page or the cart. Often it even shapes the language and images in paid ads.

3. After a peak-season result, test it again

Results from peak weeks get retested when the audience returns to normal. In the week before Black Friday, and sometimes through New Year, shoppers are focused on buying, and different factors may matter more. A test that won or lost then may not hold in a normal month.

If you do not have a lineage yet, start with research. Our guide How to do a CRO audit shows how to find the 1st test worth running. If you want help planning and reading tests as a chain, that is the work of our experimentation and testing service.

Iterative design gained 22% per round in a Nielsen Norman Group case

Iterative testing has a close cousin in product and UX work: iterative design. Nielsen Norman Group (NN/g) defines an iteration as “an intentional repetition of a step in the design process with the goal of improving the design at that stage.” In practice, a team makes a version, tests it with users, fixes the problems they find, and tests again.

The gains add up. NN/g cites a 1993 study where measured usability improved 38% per iteration. In a later website case, the target number rose 233% across 6 iterations, about 22% per round. NN/g recommends at least 2 iterations, and 5 to 10 or more when you test weekly.

External research · Nielsen Norman Group

NN/g reports that measured usability improved 38% per iteration in a 1993 study, and that a website's target number rose 233% across 6 iterations, about 22% per round. It recommends at least 2 iterations, and 5 to 10 or more when testing weekly.

Therese Fessenden, December 3, 2024 · www.nngroup.com ↗

Our A/B test lineages show the same pattern on live traffic: a 2nd, refined test often does what the 1st could not.

The main difference is the method. Iterative design usually runs on usability studies, where you watch people use the design. Iterative A/B testing runs on live traffic and measures sales. The 2 work well together. Analytics shows what visitors did, and usability testing shows why. An A/B test then proves whether the fix earns more conversions.

Frequently asked questions

What is iterative testing?

Iterative testing is running A/B tests in a chain, where each test is designed from what the previous test taught. You test a change, read how far the numbers moved, refine the idea, and test again until you find the version that wins.

What is the difference between iterative testing and A/B testing?

An A/B test is a single experiment that compares 2 or more versions of a page. Iterative testing is how you string A/B tests together: each new test uses the result of the last 1 to ask a sharper question.

How many times should you retest an idea before you drop it?

There is no fixed number. We look at the signal: how much the tested element moved the numbers. If it moved a lot, win or lose, the element matters to shoppers and is worth another round, usually in 2 different directions. In 1 of our lineages, ratings that moved conversion 15.6% without a clear result led to a 2nd test that won 11.3%.

What should you do after an A/B test loses?

Pause the test, then look at how far the number moved. A big move means shoppers reacted to the element, so test it again in a new form or place. In our tests, mobile shoppers preferred a 2-across grid by 9.4%, and a refined full width card in the next test closed that gap.

Is test and learn the same as iterative testing?

They are close. Test and learn is the general habit of running experiments and acting on what they show. Iterative testing is the more specific practice of building each new test on the result of the last 1, so the tests form a chain.

The evidence

The 3 tests this article leans on

Evidence · Collection page · Not significant

Do star ratings on collection listings lift conversion?

-15.6% conversion, -27.0% revenue per visitor

Read the full test

+11.3% conversion, -7.5% revenue per visitor

Read the full test

Evidence · Collection page · Not significant

Does a full width product card lift mobile collection page sales?

+0.5% conversion, +4.4% revenue per visitor

Read the full test

Let us find where your conversion opportunities are.

Request an evaluation