Guides
Measurement and ROI: common questions and clear answers
Measurement and ROI for sponsorship: deciding what to measure before you buy, what each metric can support, and why influencer attribution undercounts.
Every metric used for measurement and ROI in influencer marketing answers a narrower question than its name suggests. Impressions do not mean people. Engagement rate does not mean interest until you know its denominator. Attributed revenue does not mean caused revenue. None of that makes the numbers useless: it makes them instruments with known limits, and the job is to know the limits before you build a decision on top.
The failure this page is written against is the one where a brand reports a return figure to three decimal places, cannot say what it would have earned without the campaign, and quietly stops asking.
What to take away
- Whether the report comes from a creator, an agency, or a platform, a few checks catch most of the trouble.
- A return figure is only as honest as its denominator.
- The temptation is to rank partners on a single efficiency number and cut the bottom.
- Write the plan before the buy, including what a failure would look like.
Decide what you are measuring before you buy
The order matters. A measurement plan written after a campaign is a search for a favorable number; written before, it is a test.
Four questions, answered in writing, before any money is committed.
What do we believe this will do? Not "raise awareness." A specific mechanism: people who have never heard of us will hear a person they trust explain what this is for, and some of them will look it up. The mechanism determines the metric, and the mechanism itself comes from the single aim set in strategy and objectives.
What would we see if it worked? Pick the smallest number of signals that would actually change your mind, and say roughly what movement would count. Vagueness here is what lets any result be declared a success.
What would we see if it did not? This is the question that separates measurement from reporting. If no possible outcome would cause you to stop, you are not running a test.
What confounds this? A product launch, a price change, a seasonal peak, another channel's campaign, a piece of press. Write them down in advance, because after the fact they become excuses rather than controls.
What each metric can and cannot support
| Metric | Genuinely tells you | Does not tell you |
|---|---|---|
| Impressions / views | Roughly how much distribution the content got | Whether anyone watched, or who |
| Reach | An estimate of unique accounts served | Overlap with your existing customers |
| Engagement rate | Relative response, if you know the denominator | Whether the response was positive, or from your audience |
| Watch time / completion | Whether the content held attention | Whether the sponsored section held it |
| Click-through | That some viewers acted immediately | The larger group who acted later, elsewhere |
| Code redemptions | A floor on attributable purchases | Anyone who bought without using the code |
| Attributed revenue | What your model assigned to the channel | What the model missed, or wrongly claimed |
| Brand search volume | A visible spike in intent | Whether it converted, or came from this creator |
| Follower growth | Direct audience transfer | Quality or durability of it |
Two entries deserve elaboration because they mislead the most often.
Before comparing any of them across platforms, establish whose definition each number uses and write it down. A view is not the same event on every service, and even the most basic unit has never settled into one meaning, as the entry on the impression in online media sets out.
Engagement rate is a ratio, and vendors compute it against different denominators (followers, reach, views, or impressions), producing wildly different figures for identical content. Before comparing two creators on it, establish that both numbers were computed the same way. Across different platforms, the comparison is usually meaningless regardless.
Attributed revenue is a modeling output, not an observation. It depends entirely on the attribution window, the model, and what the tracking could see, and the family of models in ordinary use is surveyed under attribution in marketing. Changing the window changes the answer. Reporting it without the method is reporting an opinion as a fact.
Why influencer attribution undercounts
This channel is structurally harder to track than most, and understanding why prevents both over- and under-investment.
Dark social. A viewer sees a recommendation, screenshots it, sends it to a friend, and the friend searches for the brand a week later. That path leaves no trace your analytics can follow. It is also one of the mechanisms the channel works through, which means the tracking is blind to the thing you are buying.
Discovery-to-purchase delay. Someone hears about a product, does nothing, and buys it two months later through a search ad. Last-click credits the search ad. The search only existed because of the video.
Codes and links undercount by design. Many people who are influenced by content do not use the code: they forget it, they buy through a different retailer, they buy a different product from the same brand. A code redemption count is a floor, not a measure.
Platform and privacy constraints. Cross-app tracking is limited, referrers are stripped, and in-app browsers behave inconsistently. The trend has been toward less visibility, and it is reasonable to assume that will continue.
Views on surfaces you did not buy. Content gets reshared, clipped, and embedded. Some of the value shows up in places your reporting has no view of.
The correct response is not to abandon measurement. It is to stop treating tracked conversions as the total, and to use a method that does not depend on following individuals.
Methods that survive the tracking gaps
Holdout and geographic tests. Run the campaign in some markets and not others, chosen to be comparable, and compare the difference. This is the closest most brands can get to a real answer, and it measures the thing that matters, incremental effect, rather than the thing that is easy to see. It requires enough volume to detect a difference and enough discipline to leave the holdout alone.
Before-and-after with a stated baseline. Weaker, but far better than nothing. Establish the baseline over a period long enough to include normal variation, write it down before the campaign, and record everything else that happened in the window. The key discipline is defining the baseline first; a baseline chosen afterward is chosen to flatter.
Post-purchase surveys. Asking customers how they heard about you has known biases (recency, recall, and the fact that people misattribute), but it catches paths your tracking cannot see at all. Useful as a directional cross-check against attributed numbers, not as a replacement.
Brand search and direct traffic. A visible lift in searches for your name, or in direct sessions, around a publication date is a reasonable signal for awareness-led work. Watch for the confounds you listed in advance.
Unique codes and links. Still worth using, for the floor they establish and for the per-creator comparison they enable. Just report them as a floor.
Content performance in isolation. Whether the sponsored segment specifically held attention, where viewers dropped, what the comments said. This tells you about the creative rather than the sale, and it is the input to the next brief.
Use more than one. Where they agree, you can act with some confidence. Where they disagree, the disagreement is the finding, and it is more informative than any single number.
Reading the report you get
Whether the report comes from a creator, an agency, or a platform, a few checks catch most of the trouble.
Ask which numbers are observed and which are modeled. Ask for the definition of every ratio, especially any denominator. Ask what the attribution window was and what happens to the figure if you halve it. Ask whether the period is like-for-like or was chosen to include a peak. Ask what the comparison is: against what baseline, against which alternative use of the money.
Look for aggregation that hides variance. A campaign average across twelve creators can be carried by one, and the average is then a description of nothing. Per-creator numbers change decisions; the average rarely does.
Look for metrics that changed between the plan and the report. If the plan named one measure and the report leads with another, ask what happened to the first.
Where a claim is made about performance, check whether it survives the same questions you would put to a vendor's published case, which apply just as much to your own internal reporting.
The costs people leave out
A return figure is only as honest as its denominator. The fee is the visible part, and what that fee actually bought, rights, exclusivity and amplification included, is set out in rates and negotiation.
Content production costs, if you paid for any. Product supplied, at cost. Shipping. Agency or platform fees, including any percentage taken from the creator that you are effectively funding. Paid amplification spend behind the content. The license, if it was priced separately. Discounts given through the code: a redemption at a discount is not a full-price sale. Returns, which in some categories materially change the picture and arrive after the report was written. And staff time, which is real even though nobody invoices for it.
Leaving these out does not make the campaign look better; it makes the number unusable for the decision it exists to inform, which is whether to spend the money here or somewhere else.
Comparing across creators fairly
The temptation is to rank partners on a single efficiency number and cut the bottom. It is usually wrong, for three reasons.
Different creators were bought for different jobs. One booked for reach and one booked for conversion should not be compared on conversions.
Sample sizes differ enormously. A creator whose content underperformed once may have underperformed once. Distribution is volatile, and a single post is a small sample of an uncertain process.
Efficiency and scale trade off. The most efficient partner is frequently one who cannot absorb more budget. Ranking on efficiency alone leads to a plan you cannot execute.
A more useful frame: separate the decision to keep working with someone from the decision about how much to spend with them, and give a partner more than one outing before concluding anything, if the budget allows. A weak result is also weak evidence about the partner specifically, because the creative, the offer and the timing are tangled into it; what the selection was actually based on is in influencer discovery, and what the deliverable promised is in campaign briefs.
Bottom line
Write the plan before the buy, including what a failure would look like. Treat tracked conversions as a floor rather than a total, because this channel's main mechanism is invisible to tracking. Use a holdout or a stated baseline for the real question, cross-check with a second method, put every cost in the denominator, and report the method alongside the number.
Common questions
What is a good return figure for this channel?
There is no figure worth quoting, because it depends on margin, category, attribution method, and what the alternative use of the money was. A number without its method attached is not comparable to yours.
Should we use last-click attribution?
Know that it systematically undercredits discovery channels. Use it if it is what you have, report it as a floor, and pair it with a method that does not depend on tracking individuals.
How long should the measurement window be?
Long enough to cover the purchase cycle for your product, decided before the campaign. Extending a window after seeing the results is how a disappointing campaign becomes a successful one on paper.
Can we measure a single post reliably?
Rarely with much confidence. Single posts are noisy. Patterns across several pieces, or a geographic test, give a firmer answer.
What should we ask a vendor promising better attribution?
What they can observe directly, what they model, what happens to their number when the window changes, and whether you can export your raw data if you leave.