The organizing idea
Data is the business, monetised twice
The no-paywall, free model isn't giving the product away — it's maximising the audience, which maximises the behavioural signal. The audience business and the data business are the same business, captured once and sold twice. Every kind of data below serves one dependency spine:
Instrument → Audience → Aggregate → Intelligence → a product AI turns into the ceiling. You can't sell — or apply AI to — data you never collected.
02
Sports / content data
Active · fixtures live
The sports information that fills the product — and a distinction worth being precise about, because two different league lists live under "our data."
Structured data feeds — owned
WNBA · WSL · NWSL fixtures/results/standings from TheSportsDB, API-Football, ESPN — chosen for free coverage, presented through one owned normalisation layer. Nielsen retired.
Licensed editorial
Serie A Femminile · Coppa Italia · KLPGA · Sensational — the licensed content spine, via AP/Reuters/PA Media through Hermes. Data-feed coverage for these (and MENA leagues) is thinner.
Live tier deferred — real-time in-progress scores are the expensive tier, tabled to v3. And the coverage cliff is real: the MENA leagues that differentiate us (Saudi WPL, regional) often have no clean structured feed — a candidate for data we produce and own where no one else does.
04
Content-performance intelligence
Near-term · the missing loop
This one isn't in any current plan, and it's the most useful near-term add. Audience data tells us about the person; content-performance data tells us about the content — which is a different, equally important question.
- Which content works — videos that over-index, headlines that convert, sports and athletes that drive subscription vs. churn.
- The editorial feedback loop — it steers the content calendar and the She Series production decisions with evidence instead of instinct.
- It closes the loop between "what we make" and "what the audience does" — making the content operation smarter, not just the sales deck.
Why it's near-term — we're already capturing the raw material (engagement on every surface); this is a reporting/analysis layer on top of it, not a new pipeline. It makes today's core job — making content people want — measurably better.
What has to be decided
The open decisions
Carried from the data scoping work, still the gates on the core asset. Most of the plan quietly waits on one of these.
- Pull the YouTube geography / demographics — the one unblock that serves three purposes (validates the MENA thesis, baselines the funnel, baselines the SEO migration). Do first.
- Write the event taxonomy + identity model — the one irreversible pre-engineering decision; the gating deliverable before the portal build.
- Name one B2B buyer and what they'd pay for — turns the data layer from a by-product into a sizable revenue line, and tells us which cuts to package.
- Set the funnel-stage numeric bars — YouTube→site, site→email, email→portal sign-up, D7/D30 return (baseline first, then targets).
- The PDPL/regulatory watch — assign it; make it ongoing, not a one-time checkbox.
The keystone — a striking amount of the plan collapses back to one decision: name the buyer. It determines what data has value, which determines what the schema must capture, which determines the whole build. The data work isn't blocked by technical complexity — it's waiting on a commercial decision.