AI Marketing OSAI Marketing OS
METHODOLOGY

How We Test and Review AI Tools

Last updated:

The AI marketing tool market is full of confident launch pages and thin comparisons that all say the same things. This page explains exactly how a tool gets evaluated here, what our scores mean, and what we are not claiming. If you are going to trust a recommendation, you should be able to audit the method behind it.

How a tool ends up on this site

Three routes. First, it comes up repeatedly in real marketing workflows — in job postings, in community threads, in the tool lists teams actually paste to each other — and we go check it. Second, a vendor submits it, which gives us a demo or trial account but no editorial commitment; submitting a tool does not guarantee coverage, a score, or a mention. Third, a reader tells us we are missing something obvious.

Being listed is not an endorsement and carries no cost. Coverage decisions are editorial. If a tool is in a category we cover and enough marketers rely on it, it belongs in the comparison whether or not it has an affiliate programme — and several tools here genuinely do not.

What we evaluate

Every review is built against the same six dimensions, weighted by what tends to break in practice rather than by what demos well.

Core capability

Does the tool actually do the job it claims, for a marketer, without a data team standing behind it?

We check the specific marketing task — technical audit, keyword cluster, ad-copy variant generation, deliverability path, attribution join — and whether the output is usable after light editing or needs to be rebuilt from scratch.

Accuracy and output quality

When the tool produces something, is it right, and is it reliably right?

For generation tools we look at factual grounding, hallucination rate on domain-specific prompts, and whether output changes character across runs. For analytics tools we check how the numbers are defined and whether the definition is documented.

Pricing reality

What does it cost at the seat count and volume a real team runs, not at the free tier?

We record the entry plan, the plan that unlocks the features we are actually reviewing, and the overage or credit mechanics. A generous free tier with an unaffordable second tier is a finding, and we say so.

Integration and workflow fit

Does it connect to the rest of the stack, or does it become another silo?

We look for native integrations with the platforms marketers actually run — CRM, ad platforms, GA4, CMS, ESP, warehouse — and for an API or webhook path when there is no native connector.

Data handling and trust

What happens to customer data, and can you defend the tool in a security or compliance review?

We read the privacy policy, data-processing terms, sub-processor list, and any published model-training posture. We note when a tool uses customer data for training by default and whether that can be switched off.

Support, transparency, and momentum

Can you get help, and is the product still moving?

Documentation quality, changelog cadence, status page, and whether support is reachable on the plan we evaluate. A tool with no public changelog in a year is a different risk than one shipping weekly.

Where our information comes from

We prioritise primary sources: the vendor\'s own product documentation, pricing page, changelog, status page, privacy policy, and data-processing terms. Where we have access — a trial, a demo environment, or a vendor-provided account for review purposes — we work inside the product and describe what we saw, including screenshots where they add information.

We also read credible third-party material — independent reviews, practitioner write-ups, community threads, and public benchmarks — to pressure-test our own impression. We do not copy competitors\' descriptions, and we do not restate vendor marketing copy as fact. Where a claim is vendor-reported rather than something we verified, the page says so.

What our ratings mean

A score on this site is an editorial rating assigned by our editorial team against the criteria above. It is not a user-review average, it is not aggregated from reviews we did not collect, and we do not publish review counts we cannot substantiate. Where a page shows a score, it reflects our judgement of the tool for the marketing use case named on that page — not a universal verdict on the product.

Scores are relative within a category, not across categories. A 4.5 in email deliverability tooling and a 4.5 in AI ad-creative generation mean different things because the failure modes differ. Read the written assessment, not the number.

What we do not claim

  • Prices and feature sets change without notice. Every review is dated, and you should confirm current terms on the vendor's own site before you buy.
  • A tool we rate highly may still not fit your stack, your compliance posture, or your team's skill level. Ratings are a starting point for your shortlist, not a decision.
  • We cannot test every plan tier, every integration, and every edge case. Where our assessment rests on vendor documentation rather than our own trial, we say so on the page.
  • For brand-new products we may publish a first-look summary based purely on public material, clearly labelled as such, and revisit it once there is enough substance for a full review.

Money, independence, and corrections

The site is funded by affiliate commissions and display advertising. Neither buys coverage, placement, a score, or silence about a weakness. Our full statement is on the Affiliate Disclosure page.

Reviews are dated and revisited as products change. If something on a page is wrong — pricing, a discontinued feature, a claim we got backwards — email [email protected] and we will check it and correct the page where warranted. Corrections are made whether or not the tool in question pays us anything.