WhatsApp Group vs App: How Olio Tested a Marketplace With 12 Neighbours
Short answer: Start with a small WhatsApp group when your main uncertainty is whether nearby people will list, request and complete exchanges. Build an app only after at least 10 to 20 target users produce repeated successful matches without the founder manufacturing every interaction. Measure listings, response time, completed exchanges and return participation, not messages or sign-ups.
An app can make a working exchange easier. It cannot make people supply something nobody wants, persuade buyers to trust strangers or create enough local density for matches. Building first often turns a market-risk question into an expensive software project.
Olio's founders tested their food-sharing idea with 12 people in one neighbourhood using a WhatsApp group. The group was initially quiet, then one person offered half a bag of shallots and activity followed. Two weeks later, the test group told the founders to build the app. Olio's company history describes the 12-person test. Co-founder Tessa Clarke says in this founder interview transcript that the app took five months to launch. Separately, this account from Amazon UK reports that street-level recruitment produced an email list of more than 2,000 people by launch.
That story does not prove every marketplace can start in a group chat. It shows how to test the risky human behaviour before automating it.
Use the Neighbourhood Liquidity Ladder
The Neighbourhood Liquidity Ladder has five rungs. Do not climb to the next until the current one works without heroic founder intervention.
| Rung | Evidence required | Failure signal | |---|---|---| | 1. Supply | Target users list genuine items or availability | Founder provides most listings | | 2. Response | Suitable demand appears within the useful window | Listings expire unanswered | | 3. Match | Both sides agree practical terms | Interest does not become an exchange | | 4. Completion | Handover or service actually occurs | Cancellations and no-shows dominate | | 5. Return | Users list, request or refer again | Activity ends after the novelty period |
My view is that a local marketplace should not commission bespoke software before reaching rung four manually. Some founders argue that a polished interface is necessary to inspire trust. Design can help, but it cannot rescue insufficient supply, weak demand or an exchange that is too awkward to complete. First prove that people will tolerate a rough route to the outcome.
Define one exchange, in one place
Choose the smallest useful geography and one type of exchange. "People in London sharing things" is not a test. "Residents within a ten-minute walk offering unopened surplus groceries for same-day collection" is testable.
The boundary matters because marketplace value depends on density. A hundred members spread across a country may create fewer viable matches than 12 neighbours on three streets. Define how far users will travel, how quickly the item loses value, and the minimum information needed to decide.
Recruit both sides deliberately, but do not fake them. It is reasonable for the founder to explain the group, invite participants and prompt an initial listing. It is not evidence if the founder supplies every item, personally finds every recipient and coordinates every collection indefinitely.
Use [idea testing before spending](/business-ideas/test-your-idea-before-spending/) to decide what evidence would change your decision before the group opens.
Instrument the group like a product
A group chat creates messy evidence unless you keep a simple event record. For every listing, note:
- time posted and time of first suitable response
- number of credible responses
- whether a match was agreed
- whether the exchange completed
- reason for cancellation or failure
- whether each participant returned within the test
Do not treat conversation as demand. A message saying "great idea" is not a match. A match is not a completed exchange. Track the full path.
Set the useful response window based on the product. Surplus cooked food may need a response within hours. A borrowed drill may tolerate two days. A local professional service may need a week. The window should reflect the customer's problem, not flatter the test.
Diagnose silence before adding features
Silence has several causes. Participants may not have supply at that moment. They may fear judgement, lack a clear listing format, distrust collection, or see no evidence that anyone else will act. Test one friction at a time.
Olio's first shallot post mattered because it demonstrated acceptable behaviour. A founder can create the same clarity with one genuine example, a fixed post format and explicit collection rules. Avoid paying for ratings, maps or notifications until you know which uncertainty they solve.
If listings appear but nobody wants them, examine the category and audience. If people respond but handovers fail, the issue is coordination or trust. If exchanges complete once but nobody returns, novelty may be doing the work. This is why [an inconclusive idea test](/business-ideas/why-your-business-idea-test-was-inconclusive/) should be diagnosed, not celebrated.
Worked example: NeighbourShare Pantry
NeighbourShare Pantry tests surplus dry-food exchange across two adjoining housing blocks. It recruits 24 households for 21 days. The founder values her time at £25 an hour.
Participants create 36 listings. Twenty-seven receive a suitable response within 12 hours, 22 matches are agreed, and 18 collections complete. Nine households participate a second time.
| Measure | Calculation | Result | |---|---:|---:| | Response rate | 27 ÷ 36 | 75% | | Match rate | 22 ÷ 36 | 61% | | Completion rate | 18 ÷ 36 | 50% | | Repeat-household rate | 9 ÷ 24 | 37.5% |
Set-up and moderation take 14 hours, costing 14 x £25 = £350 of founder time. A proposed app would cost £9,000. The manual test has therefore answered the basic behaviour question for less than 4% of that build cost: £350 ÷ £9,000 = 3.9%.
The completion rate is promising but not enough to commission the full app. Four agreed matches failed, three because collection times were unclear and one because a user changed their mind. NeighbourShare should run a second test with fixed two-hour collection slots. If at least 24 of 36 listings complete and repeat participation rises, it can specify software around observed friction rather than imagined features. These figures are illustrative, not marketplace benchmarks.
Know what the manual test cannot prove
A small group does not prove city-wide acquisition, revenue, fraud control or the economics of moderation. It proves whether the core exchange creates value under a defined set of conditions. Record the limitations.
Food, transport, payments, insurance, product safety, privacy and consumer obligations vary by country and by whether you operate the transaction or merely introduce people. A group-chat test is not exempt. In the UK, check current official guidance and obtain qualified legal advice before facilitating regulated goods, holding money or processing sensitive personal data.
When the manual exchange works, build only what removes measured friction: structured listings, location boundaries, availability, notifications or trust controls. Do not reproduce every familiar marketplace feature.
Run the test in 21 days
In the first three days, define one exchange, one area, the safe participation rules and four success thresholds. Recruit 12 to 30 genuine participants by day seven. Run the group for two weeks, logging every listing through completion. Interview participants who did and did not complete an exchange. At day 21, either fix the largest failure and repeat, stop because liquidity is absent, or write the smallest software specification supported by evidence. Your next expenditure should answer the next uncertainty, not reward the excitement of having an app idea.
Related guides
Frequently asked questions
Is WhatsApp suitable for every marketplace test?
No. It is useful when a small, invited group can safely coordinate a simple local exchange. It is unsuitable when participants must conceal identities, share sensitive information, undergo regulated checks or search a large structured catalogue. In those cases, use a basic compliant web form, email process or existing specialist platform. The principle is not that WhatsApp is universally safe or appropriate. It is that you should use the cheapest lawful method that exposes the risky behaviour. Check current platform terms, privacy duties and sector rules before inviting participants.
How many people should be in the first group?
Start with 12 to 30 people who genuinely fit the two sides of the marketplace and are close enough for the exchange to work. Fewer may produce no opportunity by chance. Hundreds make it hard to observe why matches fail. The right number also depends on supply frequency. If each person can list only once a year, a small group cannot create useful evidence. Estimate expected listings during the test and recruit enough people to generate at least 20 genuine opportunities, while keeping the geography and use case narrow.
Should I charge people during the manual test?
Charge when payment is part of the central risk. A free exchange can prove matching and completion, but it cannot prove that a commission, subscription or listing fee works. Introduce the intended charge once the core exchange is functioning, disclose it clearly and measure whether behaviour changes. Do not take customer money unless you can fulfil, refund and account for it properly. Payment services, consumer rights and tax obligations vary by model and jurisdiction, so use current official information and qualified advice before holding funds.
What completion rate is good enough to build an app?
There is no universal percentage. Set a threshold from the value window and alternatives before testing. A same-day food exchange may need a high completion rate because expired listings create immediate disappointment. A specialist equipment marketplace may tolerate fewer matches if each one is valuable. As a working guide, investigate rather than build when fewer than half of genuine listings complete. More important, identify whether failure is fixable through software. An app is justified only when missing functionality, rather than missing desire or supply, is the repeated constraint.
How do I stop friends being too supportive?
Recruit people because they experience the problem, not because they want to help you. Do not describe the group as a favour or ask whether they like the idea. Give normal instructions, let them choose whether to list or request, and record what they do. Where safe, include participants with no personal relationship to you. A paid or inconvenient action is stronger evidence than praise. If friends dominate the test, label the result as a usability rehearsal and run a second test with strangers before spending on development.
When is it time to move from the group to software?
Move when supply appears without the founder creating it, demand responds inside the useful window, exchanges complete repeatedly, and the same operational friction is limiting otherwise valuable matches. You should be able to name the first software functions and link each to observed failures. Also estimate whether the reachable market and revenue can repay development, maintenance, moderation and acquisition. If the group works only because you personally chase every participant, software may automate messages but not solve the underlying dependence. Repeat the manual test until that distinction is clear.
Comments (0)