How to Test a Business Idea Before Spending Money
Short answer: design a test that can fail, and decide what failure looks like before you run it. The most useful discipline in pre-launch validation isn’t the test itself. It’s writing down the number that would make you stop, while you still have enough distance to write it honestly.
Then climb the evidence ladder from conversations to commitments to cash, spending nothing until each rung passes. Almost everyone tests. Very few people test in a way that could have produced a no. A test with only one possible outcome isn’t research, it’s a ritual.
Set the threshold first
Before you design anything, finish this sentence in writing:
“I will consider this idea validated if [specific measurable outcome] within [timeframe]. If I get less than [number], I will [stop, change the segment, change the offer, or change the price].”
Two rules make it work.
The threshold has to be a number, not a feeling. “Good response” is not a threshold. “Eight of forty targeted prospects book a call, and three pay a £150 deposit” is.
Write it down and tell someone else. The failure mode here isn’t dishonesty, it’s the entirely human habit of rationalising after the fact. You will get four results instead of eight and find yourself explaining why four is actually encouraging. Having written “eight” in advance, in front of a witness, is what stops that.
Picking a threshold that means something
Work backwards from the business model. If you need 30 customers to break even and your realistic conversion from a warm conversation is 20%, you need 150 conversations a year, which is roughly 13 a month. So the test should ask whether you can generate 13 qualified conversations in a month using a method you can repeat. That’s a threshold tied to survival rather than to comfort.
The evidence ladder
Eight rungs, ordered by how much the evidence is worth and how much it costs to get. Climb in order. Don’t jump to rung seven because it feels more decisive, or you’ll spend money discovering something rung three could have told you for free.
| Rung | Test | Evidence strength | Cost | What it proves |
| 1 | Problem interviews | Low | Time only | The problem exists and is described consistently |
| 2 | Solution interviews and mock-up reactions | Low to medium | Time only | Your approach is comprehensible and relevant |
| 3 | Landing page plus traffic | Medium | £50 to £300 | People will give attention and contact details |
| 4 | Waitlist or booking commitment | Medium | £0 to £200 | People will invest a small amount of effort |
| 5 | Pre-order or deposit | High | Low | People will part with money before delivery |
| 6 | Concierge delivery, manual and unscaled | High | Time | You can actually deliver the outcome |
| 7 | Paid pilot with real customers | Very high | Time plus some cost | The full loop works: sell, deliver, get paid, satisfy |
| 8 | Repeat purchase or renewal | Highest | None | The value was real, not just the pitch |
The gap between rungs four and five is the one that matters. Everything below rung five is opinion. Everything from rung five upwards involves the customer taking a real risk. A waitlist of 400 people has repeatedly turned out to be worth less than five deposits.
The four tests worth knowing how to run
The smoke test, rung 3
Build a single page describing the offer as though it already exists: the specific promise, what’s included, the price, and a clear call to action. Drive 200 to 500 targeted visitors to it using a small ad budget, a community post, or direct outreach.
Measure conversion to the action, not the number of visits.
Benchmarks worth holding to: 2 to 5% is normal for cold traffic on a considered purchase. Below 1% means the offer, the price or the targeting is wrong. Above 10% suggests either exceptional fit or, more often, traffic that was already warm.
On the ethics of it: state the price and describe the offer honestly, but don’t take payment for something that doesn’t exist. Use “join the waitlist”, “book a call” or “reserve a place”. If you do take deposits, say clearly when delivery happens and refund immediately on request. A test that damages your reputation in the exact market you’re about to enter is an expensive test.
The concierge test, rung 6
Deliver the outcome entirely by hand, for a handful of real customers, with no systems, no software and no scale. Planning a meal-prep business? Cook and deliver for six households yourself for four weeks. Planning a reporting tool? Produce the reports manually in a spreadsheet.
This is the highest-information test available and the most consistently skipped, because it feels like it doesn’t count. It does. It tells you the true cost of delivery, the real time per customer, what customers actually ask for as opposed to what they said they wanted, and whether you can tolerate the work.
Plenty of businesses discover at this stage that the thing customers value isn’t the thing they intended to sell. That discovery is worth more than the rest of the plan combined.
The pre-order or deposit test, rung 5
Ask for money before you deliver. A deposit, a founding-member rate, a pre-order, a paid place in a first cohort.
Design it so the customer’s risk is real but bounded. A refundable deposit is far better than nothing and much better than a free signup. Set a target and a deadline. “If I don’t get eight deposits by the 30th, everyone is refunded and I don’t proceed” is a completely honest structure, and it has the advantage of binding you as well as them.
The paid pilot, rung 7
Best suited to B2B and higher-value services. Offer a scoped, time-boxed piece of work at a real price, possibly discounted, to three to five customers.
Two conditions make it useful. It has to be paid, because free pilots generate polite participation and no information. And it has to have a defined end, so you can ask the only question that really matters: would you continue, and at what price?
Designing a test that can genuinely fail
Four patterns produce meaningless positives.
Testing on the wrong people. Friends, family, and your existing network of wellwishers are validating you, not the business. At least half your test subjects should have no relationship with you at all.
Asking about the future. “Would you buy this?” tests imagination. “Here’s the link, it’s £40, the first ten places start Monday” tests behaviour.
Testing too many variables at once. Change the offer, the price and the audience simultaneously and a failure tells you nothing. One change per cycle.
Stopping at the first good result. One enthusiastic buyer is a data point, not a market. Run until you can tell a pattern from a coincidence, which usually means 30 or more prospects for outreach tests, 200 or more visitors for landing pages, and at least five paying customers for delivery tests.
A worked test plan
The idea: per-job profitability reporting for small building firms. The founder has £2,000 and eight weeks.
| Week | Rung | Test | Threshold | If it fails |
| 1 to 2 | 1 | 12 problem interviews with firms running 3 to 15 jobs | 7 or more describe an existing workaround costing 4+ hours a month | Problem isn’t severe enough; change segment |
| 3 | 2 | Show 8 of them a one-page mock report | 5 or more ask “can I have that for my jobs?” unprompted | Wrong output; redesign the deliverable |
| 4 | 3 | Landing page plus £200 of targeted spend | 3% or more book a call | Message or targeting wrong; iterate the copy first |
| 5 | 5 | Offer a paid fourweek pilot at £250 | 4 firms pay | Price or perceived value wrong; test £150 and a narrower scope |
6 to 8 | 6 and 7 | Deliver all four manually in spreadsheets | 3 of 4 say they’d continue at £180 a month | Value isn’t recurring; consider a one-off product instead |
Total cash at risk: about £250. Total elapsed time: eight weeks. At the end of it, this founder either has four paying customers and evidence of recurring value, or they know exactly which link in the chain broke, and they still have £1,750 in the bank.
Compare that with building software for six months first.
Reading the results honestly
A pass means the threshold was met by strangers, with money or meaningful commitment, in a way you could repeat. Proceed, but proceed to the next rung rather than straight to a full launch.
A near-miss is the most common and most confusing outcome.
Hitting half your threshold usually means one component is wrong rather than the whole idea. Diagnose in this order: wrong audience, then wrong message, then wrong offer, then wrong price, then wrong problem. Most near-misses turn out to be audience or message, and both are cheap to change.
A clear fail is good news, and the sooner it arrives the better. You’ve saved the money, the year, and the emotional cost of a slow failure. The right response is to go back to problem discovery, not to run the same test again with better copy.
The result to distrust most is enthusiasm without commitment. Long, warm conversations, plenty of “definitely keep me posted”, and zero deposits. That pattern means the problem is real but not urgent, or that you’re talking to someone who feels the pain but doesn’t hold the budget.
When you can skip testing
Rarely, but it happens. If the capital at risk is genuinely trivial and the test is the launch, because you can sell a service to your existing network next week with no setup cost, then launch and treat the first ten customers as the experiment. Set thresholds anyway.
Testing exists to stop you spending money you can’t afford to lose on a guess. Where there’s no money at risk and no long build, the fastest test is the real thing.
Frequently asked questions
How much should I spend on testing?
Between 2 and 5% of the capital you’re considering committing, as a working guide. If you’re planning to spend £20,000, a £500 test that could prevent the whole outlay is obviously worth it. The rule to hold onto is that testing should cost a fraction of the decision it informs. If your test costs nearly as much as the launch, you’re not testing, you’re launching in stages.
What if I can’t build anything to test with?
You almost never need to. Rungs one to five require no product at all, only a clear description of what someone will receive and by when. A one-page description, a mockup, or a written outline of the service is enough to test whether people will commit money. Building is what you do after someone has paid, not before.
Is it dishonest to sell something that doesn’t exist yet?
Not if you’re clear about it. Pre-orders, deposits and pilots are legitimate as long as the customer knows delivery is in the future, the date is stated, and refunds are immediate on request. What crosses the line is implying something is ready when it isn’t, or holding money for something you’ve already decided not to build. Say what stage you’re at and most buyers are fine with it, particularly if there’s a founding-customer benefit attached.
How long should a validation phase last?
Six to twelve weeks for most small businesses, which is long enough to climb to rung five or six and short enough that the market hasn’t moved. If you’re still testing at month six, ask what result would actually change your decision. When nothing would, you’ve already decided and you’re using testing to postpone the risk rather than reduce it.
My test passed but I’m still not confident. What now?
Look at which rung it passed on. Passing at rung three or four justifies confidence about interest and nothing more, which is often exactly the gap you’re feeling. Climb one more rung, ideally to a paid pilot or a concierge delivery, since discomfort after a landingpage test usually means you haven’t yet proved you can deliver the thing or that anyone will pay for it twice.
Should I test more than one idea at the same time?
Test one idea properly rather than three superficially. Running parallel tests splits your attention at exactly the point where careful observation matters most, and it makes near-misses impossible to diagnose. The exception is early rung-one interviews, where you can reasonably explore two adjacent problems within the same segment before choosing which to take further.
What if competitors see my test and copy the idea?
It’s a smaller risk than it feels, and the far larger risk is spending a year building something nobody wants. Ideas are rarely the scarce input; execution, access to customers and accumulated reputation are. If your entire advantage could be lost because a competitor saw a landing page, the position was fragile to begin with, and it’s better to learn that during a £200 test than after a £20,000 build.
Do I need a large sample for the test to be valid?
Less than people assume, because you’re looking for a clear signal rather than statistical precision. Thirty targeted prospects, two hundred visitors, or five paying customers is usually enough to distinguish real interest from noise.
What matters far more than sample size is whether the people in the sample genuinely match your target segment. Two hundred of the wrong visitors tell you nothing that thirty of the right ones wouldn’t tell you better.
Comments (0)