Influencer brand safety

Influencer brand safety, proven before you sign.

One unsafe partnership can cost millions. CreatorScore scans every post, transcript, frame, and comment a creator has published — across 12 platforms and 200+ signals — and turns it into one score your team can act on before the contract, then watches it 24/7 after.

12 platforms·200+ signals·reproducible accuracy benchmark
@maya.fit
412K · fitness
Analyzing
Transcript

Reading 412 posts across 3 platforms…

CreatorScore
91/ 100
Low risk
Scanning content, comments & the open web…

What we screen for

Every post, transcript, and frame — checked

A clean caption over a risky voiceover is exactly what slips past manual review. CreatorScore reads the whole footprint — text, audio, on-screen text, and imagery — with NLP, computer vision, OCR, and speech-to-text.

Hate speech & extremism

Slurs, extremist ideology, and discriminatory language across 35+ patterns — in captions, transcripts, and on-screen text.

NSFW & explicit content

Computer vision on thumbnails, video frames, and images for nudity and sexually explicit material.

Profanity & vulgar language

Niche-aware — tells casual comedy language apart from genuinely hostile communication.

Misinformation & false claims

Health misinformation, conspiracy, and misleading claims that draw regulatory scrutiny or backlash.

Controversial topics

Divisive political and social content that can alienate segments of your audience.

Violence & graphic content

Visual and textual analysis for violent imagery and glorification of harm — including live streams.

@maya.fit

512 posts412K followers

Fitness · wellness · daily workouts 💪

Scanning 512 posts, transcripts & comments0 risks

How scoring works

The Content Risk Agent, in the open

Content risk is the most heavily weighted part of every CreatorScore — because one incident can do lasting damage. Nine signals combine into a 0–100 risk score, and every one is visible and auditable.

Beyond the creator's own content

Open-web reputation is corroborated across multiple independent sources — never a single unverified thread.

Partnership & disclosure history

FTC disclosure rates and past collaboration outcomes factor in — measured from confirmed partnerships, not caption guesses.

Traceable to the post

Every penalty links back to the exact post, comment, transcript line, or frame that caused it, with SHAP explainability.

Content Risk Agent

9 weighted signals

20% of score
  • Hate signals18.4%
  • Deceptive / scam content18.4%
  • NSFW detection13.8%
  • Flag severity9.2%
  • Web controversy9.2%
  • Visual risk analysis9.2%
  • Misinformation9.2%
  • Transcript controversy8%
  • Profanity4.6%

When a signal doesn't apply, its weight redistributes across measured signals — creators are never penalized for missing data.

Non-negotiable thresholds

Some risks cap the score — no matter what else is good

A few risks are severe enough that no amount of positive signal should override them. Each cap is precise, evidence-gated, and designed so a single ambiguous frame never sinks a creator.

Bot commenter ratio > 60%

Capped at 20/100

More than half the audience is artificial. Any spend reaches bots, not consumers — the most severe cap short of an auto-fail.

Engagement pods > 80%

Capped at 30/100

Overwhelming evidence of coordinated fake engagement. Metrics are inflated and don't reflect genuine interest.

Produced hate speech (confirmed)

Capped at 35/100

3+ posts read ≥95% hate AND an LLM confirms the creator produced it. Commentary that covers or condemns hate is never counted — a classifier hit alone can't fire this.

Sustained explicit content

Capped at 35/100

3+ vision-confirmed explicit posts across ≥20% of the history. Maternity, fitness, medical, art, and cosplay are excluded, so one ambiguous frame never caps a creator.

FTC disclosure < 10%

Capped at 35/100

Measured across 3+ verified brand partnerships. Caption keyword guesses never fire this — only confirmed partnership records do.

Documented harassment

Capped at 65/100

Fires only on external evidence — a cease-and-desist, a lawsuit, a doxxing incident. An AI reading of tone alone can never cap a score.

Knockouts apply after all seven agents score — non-negotiable thresholds that override the weighted average when triggered.

Live monitoring24/7
  • Hate signal detectednow

    New Reel — slur in on-screen text (OCR)

    @ryder.games

Real-time monitoring

Brand safety isn't a one-time check

A creator who scored 88 at signing can score 57 a month later. New posts, stories, and live streams are analyzed as they appear — you hear about a problem from CreatorScore, not from your CMO.

Instant alerts

Email, webhook, and dashboard notifications the moment a monitored creator crosses your thresholds.

Continuous scanning

24/7 across every connected platform — no waiting for a weekly review.

Score trending

Track how a creator's risk profile shifts over time, with a full historical audit trail.

Manual review vs CreatorScore

Why teams stop screening by hand

How AI-powered brand safety screening compares to a manual review process, across the metrics that matter.

MetricManual review CreatorScore
Content analyzed per creator 10–20 recent posts All posts, comments, transcripts
Time per creator 2–5 hours Under 15 minutes
Visual content screening Manual spot-check AI frame-by-frame analysis
Video transcript analysis Rarely done — too slow Automatic transcription + NLP
Consistency across reviews Varies by reviewer One fixed, reproducible model
Ongoing monitoring Periodic manual checks Continuous 24/7 scanning
Historical content review Limited by time Full-history analysis
Cost per creator $50–200+ in staff time From $0.50 / creator

FAQ

Influencer brand safety, answered

What is influencer brand safety?

+
Influencer brand safety is the practice of evaluating creators for content risk before partnering with them in marketing campaigns. It means screening a creator's history — posts, captions, video transcripts, on-screen text, imagery, comments, and web reputation — for hate speech, NSFW content, misinformation, controversy, and audience fraud that could damage a brand if associated with the creator. Effective brand safety goes beyond a one-time content check; it requires continuous monitoring across every platform where the creator publishes.

How does AI detect unsafe influencer content?

+
CreatorScore combines natural language processing for text, computer vision for images and video frames, OCR for on-screen text, and speech-to-text for video transcripts. The Content Risk Agent weighs nine signals — hate signals (18.4%), deceptive/scam content (18.4%), NSFW (13.8%), flag severity (9.2%), web controversy (9.2%), visual risk (9.2%), misinformation (9.2%), transcript controversy (8%), and profanity (4.6%) — into one 0–100 risk score. When a signal doesn't apply, its weight redistributes across the signals that were measured, so nobody is penalized for missing data.

What triggers a brand safety alert?

+
Alerts fire when a creator's content crosses your risk thresholds. Critical triggers include a confirmed pattern of produced hate speech (which caps the CreatorScore at 35/100 — 'confirmed' meaning the classifier flags it AND an LLM verifies the creator produced it rather than covering or condemning it), sustained explicit content across the history (also 35/100), and corroborated open-web controversy. Lower-severity alerts flag elevated profanity, borderline visuals, or emerging reputational concerns for human review before you proceed.

How often are creators rescanned?

+
Creators in active monitoring are rescanned continuously as new content publishes, with full rescoring at least weekly or whenever significant new content appears. Monitoring runs across all connected platforms at once, so a problematic post on any platform triggers an immediate score update and an alert to your team.

Can I set custom brand safety thresholds?

+
Yes. CreatorScore ships recommended thresholds based on industry best practice, but brands can tune sensitivity per risk category. A children's brand might set zero tolerance for profanity while a gaming brand accepts moderate language. Custom thresholds adjust both the scoring sensitivity and the alert triggers to match your risk tolerance.

Protect your brand before the next partnership

Don't let a preventable brand-safety incident derail your program. Screen creators in minutes, not days — then let monitoring keep watch.

Influencer Brand Safety — AI Content Risk Screening & Monitoring | CreatorScore