Testing how a balm actually feels: panels, scorecards and consumer trials
Build a repeatable scorecard for hardness, payoff, slip and afterfeel, run a small trained panel, and set up a consumer trial that yields usable data and consent.
Every other number in a balm workshop is weighed or read off a probe. Feel is the one property most makers judge by rubbing a bit on the back of a hand and deciding they like it, which is why so many formulas are changed in the wrong direction. This page sets out the smallest apparatus that produces a defensible answer: a scorecard with anchored attributes, a panel large enough to mean something, a discrimination test for reformulation, and an in-use trial that collects consent properly.
Score ten attributes on a 0 to 9 scale with written anchors and a physical reference for each end. Eight to twelve trained assessors settle whether two formulas differ; thirty or more ordinary users settle which one people prefer. Judge pick-up, application, afterfeel at 2 minutes and residue at 30. Confirm hardness with a penetrometer.
What a panel settles, and what it cannot
Sensory testing measures perception. It tells you what a person detects and how strongly, and nothing else. It will not tell you whether a balm is stable, whether the oils have oxidised below the threshold at which anyone notices, or whether the product is safe. Those belong to shelf life testing, accelerated ageing and the safety assessment respectively, and a panel that likes a rancid balm has told you only that the rancidity is below its detection threshold today.
Within that boundary a panel answers three separable questions, and confusing them is the commonest mistake. Is there a difference? is a discrimination question, answered by a triangle test. What is the difference? is a descriptive question, answered by trained assessors scoring attributes. Which is better? is a preference question, and only untrained users representative of your buyers can answer it. Trained assessors are specifically trained not to express preference, so asking them which one they like wastes the training.
The vocabulary this page scores is defined on balm texture science. Read it first if hardness, payoff, slip, drag, tack, cling, cushion and afterfeel are not already distinct terms in your head, because a scorecard whose attributes overlap produces correlated scores that look like data and are not.
The ten attribute scorecard
Ten attributes is the practical ceiling for an untrained-to-lightly-trained panel. More and the later ratings degrade, because each assessment competes for attention with the sample changing under the finger. The set below follows the phase structure used in descriptive skinfeel work: what you perceive taking the product up, what you perceive spreading it, and what remains.
| Attribute | Phase | 0 means | 9 means |
|---|---|---|---|
| Firmness | Pick-up | Yields to a light finger | No mark under firm pressure |
| Payoff | Pick-up | Nothing transfers in one pass | Thick deposit in one pass |
| Melt speed | Pick-up | Still solid after 10 strokes | Liquid on contact |
| Slip | Application | Finger will not move freely | Glides with no resistance |
| Drag | Application | No skin pull at all | Skin bunches under the finger |
| Cushion | Application | Finger feels bare skin | Thick layer between finger and skin |
| Tack | 2 minutes | No pull on lifting the finger | Finger sticks audibly |
| Gloss | 2 minutes | Matt | Wet-looking shine |
| Greasiness | 30 minutes | Feels like untreated skin | Mobile oil transfers to a second finger |
| Residue | 30 minutes | Nothing perceptible | Continuous film, visible edge |
Use a 0 to 9 integer scale rather than 1 to 5. Nine points give assessors room to separate two close formulas without inventing decimals, and the zero point is genuinely useful because "none of this attribute" is a real observation. Print the anchor wording on the card itself. An assessor who has to remember what 7 meant last week is calibrating against memory, and memory drifts.
Anchoring the scale with real reference samples
Written anchors alone are not enough. Two people reading "glides with no resistance" will place the same balm two points apart until they have both felt something that is definitionally a 9. There is no published reference set for anhydrous balms, so you build one and keep it.
Practical references, coded and stored in identical pots: neat castor oil as tack 9, since its high viscosity and hydroxyl content make it the tackiest common balm liquid, described on castor oil; white petrolatum as greasiness and residue 9, described on petrolatum and mineral oil; squalane as slip 9 and tack 0; untreated skin as residue 0; and a poured puck at 15% carnauba as firmness 9. Present the references at the start of every session, unscored, for one minute. That single habit does more for reproducibility than adding four assessors.
Reference materials age. Petrolatum will not change over a year, but an infused or high-oleic reference oil will, and a shifted reference silently shifts every score on the card. Date the reference pots, replace the oil-based ones every six months, and never top one up.
The timed protocol, pick-up to thirty minutes
Dose and timing vary more between people than the formula difference you are trying to detect, so both get fixed before anyone scores anything.
- Condition the room and the samples. 20 to 22 C and 40 to 60% relative humidity, samples equilibrated in that room for at least two hours, temperature recorded on the score sheet. A balm two degrees warmer is a measurably different balm.
- Prepare the site. Volar forearm, marked into 5 cm squares, washed with water only and left 15 minutes. No other product on the arms for two hours beforehand. One square per sample, never reused in the same session.
- Pick-up, then score firmness, payoff and melt speed. A fixed dose: 0.05 g weighed on a 0.001 g balance, or a single controlled pass of a stick, which is nearer to how the product is used but less repeatable. Whichever you pick, do not change it mid-study. The dose that real users apply is a separate question, covered on how much balm to apply.
- Rub out, then score slip, drag and cushion. Ten circular strokes at roughly one per second under light pressure. Count them aloud for the panel so everyone does ten.
- Wait two minutes, then score tack and gloss. Tack is judged by pressing a clean fingertip and lifting, not by rubbing again.
- Wait to thirty minutes, then score greasiness and residue. This is the interval people skip, and it is the one that predicts complaints. Almost every "too greasy" review is a residue score, not an application score, which is the mechanism set out on greasy, heavy afterfeel.
- Record immediately, with sample code, room temperature, assessor code and time. Then rest at least ten minutes before the next sample, or the previous film contaminates the next score.
Three or four samples per session is the limit. Beyond that the forearm has run out of clean sites and the assessor has run out of attention.
Panel size: eight to twelve trained, thirty or more untrained
A descriptive panel of 8 to 12 assessors, screened and trained, is the standard working size and the one that a small maker can realistically assemble. Screening means checking that a candidate can rank three obviously different firmnesses in the right order and can repeat their own scores within about one point on a blind duplicate. Training means four to six sessions on the references above until the panel's scores for the same sample agree within roughly two points on the 0 to 9 scale. Assessors who cannot repeat themselves are not bad people, they are noise, and a panel of six consistent scorers beats a panel of twelve inconsistent ones.
Preference is a different instrument. Thirty completed responses is the usual floor for a consumer preference test, and it is a floor rather than a target: at n = 30 a 60:40 split is not distinguishable from a coin toss, so if you need to detect a modest preference you need 60 to 100. Recruit people who buy the category, not friends, and never ask them to describe attributes. Ask which they prefer, how much, and why in their own words.
Triangle testing a reformulation
When you substitute an oil, drop a wax by a percent, or change supplier, the first question is not whether the new version is better but whether anyone can tell. That is a triangle test: each assessor gets three coded samples, two identical and one different, and picks the odd one out. Guessing alone gives one correct answer in three, so the whole test is a comparison against that baseline.
| Assessments | Correct needed at 5% | Correct needed at 1% | As a percentage, 5% level |
|---|---|---|---|
| 12 | 8 | 9 | 67 |
| 18 | 10 | 12 | 56 |
| 24 | 13 | 15 | 54 |
| 30 | 15 | 17 | 50 |
| 36 | 18 | 20 | 50 |
| 48 | 22 | 25 | 46 |
| 60 | 27 | 30 | 45 |
Assessments, not people: twelve assessors doing two replicates each gives 24 assessments, although replicates from the same person are not fully independent and a conservative reading treats a marginal result from replicates as unproven. Serve all six possible orders equally often (AAB, ABA, BAA and the three mirrors) so position cannot be read as difference.
A triangle test that fails to reach significance does not prove the two samples are the same. With 24 assessments you have a good chance of missing a difference that a third of people could detect. "Nobody could reliably tell them apart at this panel size" is what you have shown, and it is the sentence to put in the batch record. Claiming identity from a small negative test is the single most common misuse of this method.
Instrument checks that keep the panel honest
Instruments are cheaper than people and do not get bored, but they read one physical property each and no instrument reads afterfeel. Use them to verify that the panel's firmness axis is tracking something real, and to catch drift between batches without convening anyone.
| Measurement | Method | Reads | Checks which attribute |
|---|---|---|---|
| Cone penetration | ASTM D937 cone, 150 g, 5 s, 25 C | Tenths of a millimetre, lower is harder. White petrolatum sits at 100 to 300 dmm under the NF monograph; a lip balm stick reads far lower | Firmness |
| Needle penetration | ASTM D1321, 25 C | The scale used for neat waxes across the wax comparison | Raw material consistency, not finished balm |
| Durometer | ASTM D2240, Shore OO or 000 | 0 to 100 index on a poured puck at least 6 mm thick | Firmness, between batches only |
| Texture analyser | Cylindrical probe, 1 mm/s, 5 mm depth | Peak force in grams, plus work of penetration | Firmness and, loosely, payoff |
| Slip or drop point | Ph. Eur. 2.2.17 or USP 741 | The temperature band where the product loses structure | Melt speed on skin |
| Gravimetric payoff | Weigh stick before and after ten passes | Milligrams deposited per pass | Payoff, directly |
Two cautions. Durometer hardness is defined for rubbers and plastics, so on a fat crystal network it is an index rather than a material property; readings only compare within one geometry, one temperature and one instrument. And every one of these is temperature sensitive to a degree that will swamp your formula change, which is why the full argument for controlling the measurement sits on measuring balm hardness, alongside the underlying structure on rheology and yield stress. Where instruments and panel disagree on firmness, suspect the sample thermal history first: cooling rate changes crystal size without changing the formula at all, as covered on crystallisation and cooling rate.
Controlling bias: codes, order and packaging
Everything a person knows about a sample before touching it moves the score. Four controls remove most of it, and none costs anything.
- Blind three-digit codes drawn from random numbers, never A and B, never 1 and 2, and never the same code for the same sample twice in one study. Someone other than the formulator applies them and keeps the key.
- Randomised or balanced serving order. First sample scores high on almost every attribute simply for being first. Rotate so every sample appears in every position equally often.
- Unbranded, identical packaging. Same pot, same fill weight, same lid. Colour and scent leak identity, so where the change is only textural, colour and fragrance should be matched across the samples or removed from all of them.
- The formulator does not score. You know which one you spent three weeks on. A blind duplicate of the current formula slipped into the set is the cheapest way to see how much expectation is contributing.
Run the same discipline on complaint investigations. When a batch is reported as draggy, code it against a retained sample of the previous batch and let someone else present them, otherwise you will find what you expect. The fault mechanisms themselves are on drag and poor glide and batch to batch inconsistency.
A four week in-use trial with twenty to forty people
A panel scores a single application under controlled conditions. It cannot tell you what happens when someone uses a balm on their own hands, in their own weather, twice a day for a month. That needs an in-use trial, and four weeks is the usual minimum because skin changes slowly and because novelty effects fade within the first week.
- Recruit 25 to 50 to finish with 20 to 40. Budget for 15 to 25% dropout. Below 20 completers you have anecdotes; above 40 the administration outgrows a small business.
- Screen and exclude. No broken skin at the application site, no known allergy to anything in the formula, and a documented decision on whether you include people who are pregnant or under 18. Ask everyone to patch test before day one and to stop at any reaction.
- Fix the regimen. Named site, stated dose, stated frequency, no other product on that site. A regimen people cannot follow produces data about nothing.
- Use a short daily diary. Three to five fixed questions, each on the same scale as the scorecard, plus one free text box. Long diaries get filled in on the last day from memory.
- Weigh the jars back. Issue identical coded jars weighed to 0.01 g and weigh them again at the end. The difference is the only objective compliance measure you will get, and it also gives a real usage rate for your how long a balm lasts figures.
- Analyse before you look for a story. Decide in advance which questions matter and what counts as a result. Trawling twenty diary items for the one that moved is how small trials produce claims that fall apart.
Consent, data and adverse reports
The moment you recruit participants you are processing personal data about their skin, which is health data. Under the UK and EU General Data Protection Regulation, health data is special category data and needs an Article 9 condition, in practice explicit consent for a study like this. Written informed consent should state what the product is, what it contains, what participants will do, that they may withdraw at any time without giving a reason, who holds the data, how long it is kept and when it is destroyed. Store diaries against participant codes and keep the name-to-code key separately.
Record every reported reaction, however minor, with date, description, site and outcome. In the European Union and the United Kingdom, serious undesirable effects must be notified to the competent authority under Article 23 of Regulation (EC) No 1223/2009, and the full obligation set for makers in those markets is on selling balms in the UK and EU. Human test data belongs in the product information file, so keep it in a form your safety assessor can read.
Run the trial on your existing formula before you run it on the new one. A baseline dataset on the product you already sell costs one round of admin and turns every later trial from an absolute measurement into a comparison, which is both easier to interpret and much harder to argue with.
What a small trial will and will not support
The decision rule is short. Changed an ingredient and need to know whether it matters: triangle test, 24 to 36 assessments. Need to know in which direction the texture moved: trained scorecard, 8 to 12 assessors, and expect to detect differences of about two points but not half a point. Need to know which version sells: consumer preference, 30 as an absolute minimum and 60 to 100 if the margin is likely to be narrow. Need to know how it performs over time: in-use trial, four weeks, 20 to 40 completers.
What none of this supports is a performance claim beyond the wording your data literally sustains. In the European Union and the United Kingdom, cosmetic claims must meet the common criteria of Regulation (EU) No 655/2013, of which evidential support and honesty are two. A 30-person unblinded self-assessment can support a properly qualified consumer statement, reported with the number of participants, the duration and the fact that it is self-assessed. It cannot support "clinically proven", it cannot support a barrier repair claim without instrumental measurement of water loss, and it cannot support anything that treats a condition, which is the line drawn on cosmetic versus drug claims. Which claims are lawful in your market, and whether your evidence file satisfies your assessor and your regulator, are decisions that belong to you and to a qualified safety assessor, not to a protocol on a reference site.
One last limit worth stating plainly. Sensory panels measure people, and people are variable, tired, cold and suggestible. A well-run small panel narrows the uncertainty; it does not remove it. If two formulas come out level after all of this, the honest conclusion is that they are level within your ability to measure, and the choice should be made on cost, supply or stability instead.
Frequently asked questions
How many people do I need for a sensory panel?
Eight to twelve trained assessors for descriptive scoring, which is enough to characterise how two formulas differ. Thirty or more untrained users for preference, and 60 to 100 if you expect a narrow margin, because at n = 30 a 60:40 split is not statistically distinguishable from chance. Trained and untrained panels answer different questions and are not interchangeable.
What is a triangle test and how many correct answers do I need?
Each assessor receives three coded samples, two identical and one different, and picks the odd one out. Guessing gives one in three correct, so significance is judged against that baseline. At the 5% level you need 10 correct out of 18 assessments, 13 out of 24, 15 out of 30 or 18 out of 36. Serve all six sample orders equally often.
Can a durometer replace a sensory panel?
No. A durometer reads resistance to indentation, which correlates with perceived firmness and with nothing else on the scorecard. Payoff, slip, tack and afterfeel have no instrument that reads them directly. Shore hardness is also defined for rubbers and plastics, so on a balm it is a comparative index valid only at one temperature, one sample geometry and one instrument.
How long should an in-use trial run?
Four weeks is the usual minimum, because novelty effects fade in the first week and skin condition changes slowly. Recruit 25 to 50 people to finish with 20 to 40 after dropout, use a short daily diary of three to five fixed questions, and weigh the coded jars before and after to get an objective measure of how much was actually used.
Do I need consent forms for a home use test?
Yes. Data about someone's skin is health data, which is special category data under the UK and EU General Data Protection Regulation and needs explicit consent. Written consent should cover the product, the procedure, the right to withdraw at any time, who holds the data, how long it is kept, and how reactions will be recorded and reported.
What claims can I make from a 30 person consumer test?
A properly qualified self-assessment statement, reported with the number of participants, the duration and the fact that it was self-assessed. Not "clinically proven", not a barrier repair claim without instrumental measurement, and nothing that implies treating a condition. In the EU and UK, claims must meet the common criteria in Regulation (EU) No 655/2013, including evidential support.
How do I stop assessors guessing which sample is mine?
Random three-digit codes applied by someone else who keeps the key, identical unbranded pots with the same fill weight, matched or absent colour and fragrance, serving orders balanced so each sample appears in each position equally often, and a blind duplicate of the current formula in the set. The formulator should not score.
Sources and further reading
- International Organization for Standardization, ISO 4120, Sensory analysis: methodology, triangle test, Geneva.
- International Organization for Standardization, ISO 8586, Sensory analysis: selection and training of sensory assessors, Geneva.
- International Organization for Standardization, ISO 13299, Sensory analysis: methodology, general guidance for establishing a sensory profile, Geneva.
- ASTM International, ASTM E1490, Standard practice for descriptive skinfeel analysis of creams and lotions, West Conshohocken PA.
- ASTM International, ASTM D2240, Standard test method for rubber property: durometer hardness, West Conshohocken PA.
- European Commission, Regulation (EU) No 655/2013 laying down common criteria for the justification of claims used in relation to cosmetic products, EUR-Lex.
- European Parliament and Council, Regulation (EU) 2016/679, General Data Protection Regulation, Article 9, EUR-Lex.
Reviewed and updated 6 September 2026. Spotted an error? Tell us and we will fix and log it.