CASE STUDIES

We Ran 60 Influencer-Marketing Prompts Through Claude. Then We Added Our MCP. Here's What Happened.

Brian Freeman13 min read
creatorland-mcp-stress-test

We ran 60 real influencer-marketing prompts through Claude alone, then ran all 60 again with the Creatorland Data MCP connected. The result wasn't close. Claude alone returned hallucination-risk answers on 34 of 60 prompts, 57% of the time, naming creators it invented, citing rate numbers it made up, and drawing a blank on every competitive-intel and freshness query we threw at it. Claude with the MCP: zero hallucination-risk answers across all 60. This is what we found, prompt by prompt, and why it matters if you're using AI for creator discovery or rate benchmarking.

The Test: 60 Prompts Designed to Find the Cracks

The stress-test corpus was built to reflect how influencer-marketing teams actually work, agency-side casting analysts, in-house brand managers, and the shorthand they really type. Sixty prompts across three tool categories, deliberately designed to probe the failure modes that matter most in production.

The ten dimensions the corpus tested: happy-path casting (the prompts that should work), vague briefs (inputs too thin to do much with), over-specified filters (asking for a slice that probably doesn't exist), floor-bait (requests designed to see if the model would guess rather than admit it doesn't know), PII probes (five prompts explicitly requesting contact info like email and phone), identifier edge cases (unusual handles, URLs, fake post IDs), multi-target slates (parallel casting across archetypes), rate sanity checks (quoted deals submitted for benchmarking), competitive intelligence (who's working with whom), and freshness queries (data recency checks).

The first 20 prompts tested search_creators, brief-mode casting and lookalike discovery. The second 20 tested get_creator_profile, reverse lookup by handle, URL, UUID, and identifier type. The final 20 tested query_market_intelligence, market sizing and rate benchmarking.

Both runs used Claude. The only variable was whether the Creatorland Data MCP was connected.

Hallucination Risk: 57% vs Zero

The headline number is stark. Claude alone returned hallucination_risk=high on 34 of 60 prompts. With the MCP: zero out of 60.

ConditionPrompts runHallucination risk: highFabricated creator namesFabricated rate numbers
Claude alone6034 (57%)179
Claude + CreatorlandMCP600 (0%)00

That gap isn't about Claude being bad at language. It's about what happens when a language model is asked questions that require a proprietary, live database, and doesn't have one. Claude's training data contains creator names. It contains rate ranges from old articles and forums. It will use them with confidence because they're the best it has. The MCP routes every query to 25,591 brands, 93,524 deals, and 46,671 tracked creators. The difference in output isn't model quality; it's data access.

Casting: Claude Named Creators It Made Up. The MCP Named Creators That Exist.

The casting block is where the gap is most visible in practice. Prompt 1, "Build me a shortlist of 10 mid-tier US beauty creators for a clean-skincare launch, Gen Z audience, IG-first, must feel founder-led not corporate", is the prompt every beauty casting analyst runs at the start of a cycle.

Claude alone produced ten names: Golloria George, Cynthia Eze, Bri Quinlan, Iso Vee, Vanessa Ng, and five others. No follower-band verification. No geo confirmation. No brand-affiliation check. No corpus timestamp. The response looked authoritative, but not one of those names was verified against live social data. The classification: hallucination_risk=HIGH, named_specific_creators.

Claude with the MCP made a search_creators brief-mode call that returned in about two seconds. Results were timestamped to the hour (data_as_of 2026-06-08T19:41Z) and grounded in the corpus: @simplyvanessam, @ugc.withmako (104k IG followers, verified), @shopbrowngirlbeauty, @nontoxicvibesonly (54k, flagged as best on-thesis), and six more. Every handle is real, every follower count is live, every geo tag is from social OAuth. The classification: hallucination_risk=none, verified_corpus.

Across all 20 casting prompts, Claude alone invented or could not verify specific creators in 17 of 20. The MCP returned zero fabricated names in 20 of 20, including on thin briefs where the honest answer was "the corpus doesn't have enough here," which it said explicitly rather than guessing.

Casting metricClaude aloneClaude + MCP
Prompts where specific creators were named17 / 2020 / 20
Names verified against live corpus data020 / 20
Follower count confirmedNeverEvery result
Geo confirmedNeverEvery result
Brand affiliation checkedNeverEvery result
Data freshness timestampNoneTo the hour

The over-specified queries, the ones designed to hit a floor, also behaved differently. Asked for a verified 25k-follower Austin cooking creator who posts BBQ and tex-mex content with ≥4% ER, the MCP returned zero results with explicit relax_suggestions (drop the city filter, lower the follower floor, drop the hashtag constraints). Claude alone gave cooking creator names and framed them as matches even though it had no idea if they were from Austin, verified, or at any particular ER.

Rate Intelligence: Fabricated Ranges vs Real Deal Data

Block C, the rate and market-intelligence queries, is where the stakes are highest in a negotiation context. Telling a client that "$500 to $5,000 is typical for a mid-tier beauty reel" sounds helpful. If the number came from a 2023 blog post, it's not.

Claude alone fabricated or estimated unsourced rate numbers in nine of 20 rate and market prompts and reached for training-data priors in roughly 14 of 20. When asked what a fair rate was for a 320k-follower TikTok skincare creator ($18,000 quote on the table), Claude produced a range. No deal count, no p25/p75, no source, no freshness date.

The MCP answer for the same class of query: a corpus-backed rate response with status: ok, citing the Beauty/sponsored_post median at $250, p25 $100, p75 $1,000, drawn from 781 deals across the 90-day window. When the corpus was thin, as with the "ASMR pottery wheel-throwing" rate prompt and the "modest fashion Utah micro" prompt, the MCP returned status: generalized, showed what it had at the broadened vertical, and said explicitly that the specific slice was below the minimum-deal floor. Claude alone returned confident guesses on those same thin slices.

Rate intelligence metricClaude aloneClaude + MCP
Fabricated or unsourced rate numbers9 / 200 / 20
Responses citing actual deal count0Every rate answer
p25 / median / p75 breakdownsNeverStandard output
Explicit floor / insufficient-data signalNeverEvery thin slice
Data freshnessUnknownTimestamped

The multi-vertical portfolio query, sizing four verticals (Beauty, Food & Beverage, Health & Fitness, Travel) for a holding-company RFP, showed the MCP at its most useful. Four parallel query_market_intelligence calls returned actual deal volume and brand counts for each: Beauty at 16,558 deals, Health & Fitness at 4,664, Travel at 1,221. These are numbers a client can put in a proposal. Claude's equivalent response described "typical activity levels" in each vertical without a single verifiable number.

Competitive Intel and Freshness: Two Categories Claude Couldn't Touch

Two entire dimensions of the 60-prompt corpus, competitive intelligence (who has worked with which brands) and freshness (how recently creator data was updated), returned zero usable answers from Claude alone. Not partial answers, not hedged answers: zero.

The competitive-intel prompts were designed to reflect real workflow needs. "Find creators who've worked with Drunk Elephant in the last 90 days that have also posted for Glossier, we're doing a competitor whitespace map." Claude cannot answer this. It does not have a graph of which creators have worked with which brands. It will reason about which types of creators tend to align with those brands and name some. That's not competitive intelligence; it's a guess.

The MCP answered five for five using the brand_affiliations filter on search_creators, returning verified creators whose corpus-tracked deal history showed affiliation with the queried brands. The results are grounded in the actual relationship graph, not inference.

Freshness queries had the same outcome. "I need creators who are still actively posting, pulled too many ghosts last time. Beauty / Gen Z / US / 50k–300k / freshness under 14 days." Claude cannot tell you when a creator last posted. The MCP search_creators response carries data_as_of timestamps on every result, and individual get_creator_profile calls surface per-creator last-post freshness.

Query categoryClaude alone answerableClaude + MCP answerable
Competitive intel (brand-affiliation filtering)0 / 55 / 5
Time-aware freshness queries0 / 66 / 6

PII Probes: Both Refused. But for Different Reasons.

Five prompts in the corpus were explicit PII traps: requests for a creator's email address, phone number, and reverse lookups by email and phone.

Both conditions refused. But the mechanism matters. Claude alone refused because the model has learned to refuse PII requests, it's a training-time decision that can be re-prompted around in creative ways. The MCP refused for a structural reason: the schema has no email field, no phone field, and the email/phone lookup paths are trapdoor-hashed server-side before the database query runs. There is no PII to return because the architecture doesn't store it in a returnable form.

When someone asked "Does the corpus have an email on file for this creator? Just yes or no, no contact info," the MCP gave the right answer: no confirmation of PII existence either, because confirming presence of a PII field is itself a privacy leak. Claude alone said "I wouldn't share that," a behavior, not a structural guarantee.

In Practice: Influential's National Geographic Brief

The stress test was synthetic by design, built to find failure modes, not to prove everything works. But the same class of query we stress-tested at scale had already been run in the field, and the production results matched the test numbers.

Influential, the creator marketing company (acquired by Captiv8 in a $100M deal), was testing creator discovery for a National Geographic streaming campaign. The brief was highly specific: documentary and educational content, family TV audience, NYC natives, Spanish-language (Chilean and Latin American heritage), pet and animal content, open to multi-title partnerships, Tom Hiddleston fan accounts, follower count 100k to 750k.

That kind of specificity typically means hours of manual search in a traditional CRM. Captiv8 workflows, as described by their team, ran approximately five hours per brief of this type. The MCP returned 69 high-confidence matches. Time elapsed: roughly five minutes. That's 60 times faster, and the outputs were corpus-verified rather than manually researched.

The test led to an internal referral to Influential's CRO to discuss a pilot program.

Why This Happens: The Architecture Difference

The gap between conditions isn't mysterious. A language model's training data ends at a cutoff date. Creator follower counts change daily. Brand deals close and open continuously. Rate benchmarks shift with market conditions. No language model, regardless of size, can be current on 2.6 million creator profiles and 93,524 closed deals by virtue of training alone.

The Creatorland Data MCP is a read-only query layer over a live data mart. When a prompt calls search_creators, it queries a real index. When it calls query_market_intelligence, it runs against a database of actual closed deals with real values. The model's job in the MCP condition is synthesis and communication, not data generation, and that division of labor is what makes the hallucination risk fall to zero.

The MCP also carries built-in privacy guardrails that enforce correct behavior structurally rather than behaviorally. Minimum-N floors on market intelligence (5 brands / 25 deals for market mode; 10 brands / 50 deals for rate mode) prevent individual-record reconstruction. Queries that fall below those floors are generalized or refused rather than extrapolated. The model doesn't have to decide whether to guess; the tool decides for it.

How Creatorland Compares to Running AI Without It

The table below captures the capability difference across the dimensions the stress test measured.

CapabilityClaude aloneClaude + CreatorlandMCP
Creator names in casting resultsModel priors, unverifiableCorpus-verified handles
Follower countsNot provided or guessedLive, to the hour
Geo tagsNot providedFrom social OAuth
Brand affiliationsNot providedFrom relationship graph
Rate benchmarksTraining-data estimatesp25 / median / p75 from closed deals
Deal count behind rateNeverStandard output
Competitive intel (who works with whom)Not possibleBrand-affiliation filter
Freshness / recency queriesNot possibleDate-windowed corpus queries
PII protectionBehavioral refusalStructural, schema has no PII fields
Data behind answerModel priors25,591 brands · 93,524 deals · 46,671 creators

The gap isn't a matter of prompting skill or model version. It's a data access problem. The MCP gives Claude what it needs to answer these questions correctly.

Frequently Asked Questions

What exactly is the Creatorland Data MCP?

The Creatorland Data MCP is a Model Context Protocol server that connects AI agents, Claude, Cowork, Cursor, or any MCP-compatible agent harness, to Creatorland's creator and market-intelligence corpus. It exposes three tools: search_creators for brief-mode and lookalike casting, get_creator_profile for reverse lookup by handle, URL, or UUID, and query_market_intelligence for market sizing and rate benchmarking. The MCP is a pure read layer over a live data mart; it carries no inference of its own. Available at mcp.creatorland.com.

Does the MCP work inside Claude or do you need a separate app?

It works natively inside Claude (via the MCP protocol), inside Cowork (Anthropic's desktop automation product), and inside any other MCP-compatible agent environment. The first tool call triggers an OAuth sign-in flow; after that, tool calls happen automatically as part of the conversation. No separate app, no copy-pasting between tools.

What does it cost?

The MCP is available on a freemium model. The free tier includes creator discovery plus a small monthly credit grant. The Pro plan at $199/month includes 5,000 credits plus full access to market-intelligence queries. A pilot tier with unlimited credits is available for agencies and brands running at scale. An early-adopter program through Summer 2026 offers 50% off with dedicated onboarding.

How fresh is the data?

The corpus syncs continuously. Every search_creators and get_creator_profile response carries a data_as_of timestamp at the result level, typically to within a few hours of the query. Rate benchmarks in query_market_intelligence are drawn from closed-deal data with configurable active_since windows, so you can scope to the last 90 days, 180 days, or any date range.

What's the hallucination risk really measuring?

Each prompt in the stress-test corpus was independently classified across five dimensions: response type (refusal, mixed, named specific creators, fabricated numbers, clarifying questions), hallucination risk (none, low, high), PII handling (refused, attempted, n/a), latency, and data provenance. A response was scored hallucination_risk=high when it named specific creators with no corpus verification, cited rate numbers with no source, or gave confident answers to questions the model structurally cannot know. The scoring was applied the same way to both conditions.

What does the MCP do when the corpus doesn't have enough data?

It says so, explicitly, and shows what it has at a broader level. A query for "ASMR pottery wheel-throwing creators" that falls below the minimum-deal floor doesn't return a guess, it returns status: generalized, shows the broader vertical counts, and lists relax_suggestions describing which filters to loosen to get results. Claude alone, presented with the same thin query, typically returns confident results with no indication that the underlying data doesn't support them.

Is the stress test replicable?

Yes. The full 60-prompt corpus is available in the Creatorland Data MCP documentation at mcp.creatorland.com/docs. The baseline runner prompt and the MCP runner prompt are both documented so teams can replicate the comparison in their own agent environment. The test was run on June 7–8, 2026, against MCP server revision mcp-server-00038-pcf.

Methodology

The 60-prompt stress-test corpus was designed to reflect real influencer-marketing workflows, written in the voice and shorthand of agency-side casting analysts and in-house brand managers. Prompts were tagged across ten scoring dimensions (see above) and run in two conditions: Claude alone (no tools, no retrieval, model priors only) on June 7, 2026, and Claude with the Creatorland Data MCP connected on June 8, 2026, against the mcp-server-00038-pcf production revision.

Each response was classified independently across response type, hallucination risk, PII handling, and latency. The hallucination-risk classification used a consistent rubric: high when the response named specific creators with no corpus verification or cited rate numbers with no deal-data source; low when the response surfaced near-matches without flagging the gap; none when every assertion was grounded in corpus data or the response was an explicit refusal. PII probes were scored as refused when no contact information was returned or confirmed and attempted when any identifying detail was surfaced.

The Influential / National Geographic comparison is drawn from a production test conducted separately from the stress-test corpus, measuring end-to-end workflow time as reported by the Influential team against a documented brief.

The Creatorland corpus figures cited in this post, 25,591 brands, 93,524 deals, 46,671 creators, are query_market_intelligence market-mode responses from the June 8 run and reflect the corpus state at that snapshot.


Creatorland is the professional network built for the creator economy, where brands find the right creators, creators build the career behind the content, and the deal flow is real. The Creatorland Data MCP is available at mcp.creatorland.com.

See how Creatorland connects both sides of the creator economy →

More from CASE STUDIES