πŸ“œ IP Infrastructure / Royalty Processing

AI Training Data Royalty Compliance SaaS: The ASCAP for Machine Learning

ASCAP and BMI collectively processed $3.4 billion in music performance royalties in 2025 through infrastructure whose basic architecture dates to 1914. The dataset licensing market for AI training hit $4.8 billion the same year, and the royalty processing layer that should sit between content owners and AI companies does not exist.

Abstract visualization of data streams flowing through a crystalline processing structure with copyright symbols embedded in the flow

The Problem

In 1914, Irving Berlin and a handful of other composers founded ASCAP because they had a problem no individual songwriter could solve alone: when a restaurant played your song, you had no way to know it happened, no way to bill the restaurant, and no leverage to collect. The solution was collective licensing. Pool the rights, negotiate blanket licenses, monitor usage, distribute royalties. A century later, ASCAP reported $1.945 billion in revenue for 2025, distributing $1.759 billion to its members. BMI processed another $1.573 billion in its last reported fiscal year. Between them, these two organizations move more than $3.4 billion annually through a system that lets a coffee shop in Phoenix play Taylor Swift without calling her lawyer.

The content licensing market for AI training is now larger than the entire U.S. music performance royalty market was in 2010. According to DataIntelo's 2025 market analysis, spending specifically on dataset licensing for AI training reached $4.8 billion in 2025, with proprietary and custom-negotiated licenses accounting for more than 55% of total market revenue. The broader AI training dataset market, which includes data creation, labeling, and curation alongside licensing, totaled $8.74 billion in 2025 and is projected to reach $49.82 billion by 2031, a 33.14% compound annual growth rate (ResearchAndMarkets, July 2026). The licensing segment is the fastest-growing slice of that broader market. Average enterprise-grade proprietary dataset license deal sizes grew 34% between 2023 and 2025 according to DataIntelo, reaching $1.2 million per contract for large-scale NLP and computer vision applications.

The deals are getting done:

LicensorLicenseeAnnual Value (est.)Source
News CorpOpenAI~$50M/yrCB Insights
News CorpMeta$50M/yrCB Insights
RedditGoogle~$60M/yrEverything PR
NYTAmazon$20–25M/yrPublic reporting
Dotdash MeredithOpenAIβ‰₯$16M/yrAdweek
Financial TimesOpenAI$5–10M/yrPublic reporting

CondΓ© Nast, Hearst, Vox Media, The Atlantic, Associated Press, Stack Overflow, and Shutterstock have all signed deals as well. The undisclosed agreements likely outnumber the public ones by a wide margin.

Here is the part that should make a startup founder's ears perk up. None of these deals have standardized compliance infrastructure behind them. Every contract is bilateral. Every usage audit is manual or nonexistent. No publisher can independently verify how many times their content was used in a training run, which model versions incorporated it, or whether the AI company's usage stayed within the contractual scope. The music industry solved this problem (imperfectly, expensively, but functionally) a hundred and twelve years ago. The AI content licensing market, processing billions of dollars in annual deal value, has no equivalent system.

Market Size

Original TAM calculation: The addressable market has two layers. The first is compliance SaaS for content owners (publishers, stock media libraries, academic publishers, music labels, software repositories) who have signed or are negotiating AI training data licenses. Based on the CB Insights tracker and public deal announcements, at least 80 named publishers have signed AI content licensing agreements as of mid-2026. The long tail of smaller publishers, independent creators with collective representation, stock media platforms, and academic publishers likely brings the total to 2,000–5,000 entities with active or pending AI training data licenses within 36 months, given the acceleration curve of deal-making and the regulatory pressure from AB 2013 and the EU AI Act.

At an average SaaS subscription of $2,400/month for mid-tier publishers (compliance documentation, usage tracking, audit trail) and $8,500/month for enterprise publishers with multi-platform licensing (News Corp, CondΓ© Nast, Shutterstock), with an estimated 70/30 split between tiers, the blended ARPU is $4,230/month. At 3,500 addressable publishers, the supply-side SaaS TAM is $177.7 million in annual recurring revenue.

The second layer is royalty processing. ASCAP takes roughly 10% of collections as operating overhead; BMI operates at a similar ratio. If the dataset licensing segment grows at its historical pace from a 2025 base of $4.8 billion, it reaches roughly $20–25 billion by 2034. A royalty processing organization handling even 15% of that transaction volume at a 5–8% processing fee generates $150–300 million in annual processing revenue. More conservatively, capturing $2 billion in managed licensing volume within five years at a 6% processing fee yields $120 million. Combined TAM across both layers: $298–449 million. Year 3 realistic target: 400 publisher subscribers at blended $4,230/month plus $500 million in managed licensing volume at 6% = $50.3 million ARR.

The Product

A compliance and royalty processing platform purpose-built for the AI training data licensing market, serving both sides: content owners who need to track and monetize their IP, and AI developers who need to document and verify their data provenance for regulatory compliance. Four core modules:

A critical design decision: the platform must serve individual creators, not just large publishers. ASCAP's history is instructive and cautionary in equal measure. The DOJ placed ASCAP under a consent decree in 1941 because it was using its monopoly position to set licensing terms unilaterally. That decree is still in effect 85 years later, and BMI operates under a similar one. For decades under that regime, independent songwriters received a fraction of what major publisher affiliates earned, even when their per-play rates should have been comparable. Any new royalty processing intermediary faces the same structural risk: becoming the gatekeeper it claims to displace. The product must include a self-serve tier for independent content creators, journalists, photographers, and bloggers whose individual works are swept into training datasets, and the distribution formulas must be transparent from day one. Their licensing power is negligible alone, but collectively they represent millions of copyrighted works. The platform can aggregate their rights into blanket licenses, the same way ASCAP lets a coffee shop license a million songs at once instead of negotiating with each songwriter.

A related risk: the content fingerprinting and usage audit system this startup proposes is, at its core, a surveillance apparatus. It tracks what content was ingested, by which company, into which model, at what volume. That is precisely the point, and precisely the danger. A registry of billions of fingerprinted text passages and images, cross-referenced with usage records across major AI companies, is a powerful dataset in its own right. The startup must build privacy-preserving audit mechanisms (zero-knowledge proofs, differential privacy for usage aggregates) and resist the temptation to monetize the metadata. Content creators who register their work for royalty tracking should not find that the platform knows more about the distribution and influence of their corpus than they do, or that it sells that intelligence to the AI companies they are trying to audit.

Unit Economics

MetricValue
Monthly subscription (Standard: compliance docs + usage tracking)$2,400/publisher
Monthly subscription (Enterprise: multi-platform licensing + audit)$8,500/publisher
Blended ARPU$4,230/month
Royalty processing fee5–8% of managed volume
Infrastructure cost per subscriber/month$180
Content fingerprinting compute cost/month$95
Customer acquisition cost$18,500
Expected LTV (36-month avg retention, 93% gross margin)$141,552
LTV:CAC ratio7.7:1
Gross margin93%
Startup cost (24-month runway)$6.2M
Break-even22 months

Methodology note: The 36-month retention assumption is based on the stickiness of compliance infrastructure. Once a publisher's licensing agreements reference the platform's audit reports, switching costs are contractual, not just operational. CAC of $18,500 reflects enterprise B2B sales to media companies and publishing groups, where deals close through direct sales at industry events (the News Media Alliance's Nexgen conference, the SIIA's annual conference, Digital Content Next's member meetings) and require multiple stakeholder sign-offs. The $6.2 million startup cost covers 24 months of a 14-person team: 6 engineers (content fingerprinting, API integrations, compliance engine), 2 legal/regulatory specialists, 3 enterprise sales, 1 customer success, 1 product, 1 founder/CEO. Gross margin of 93% reflects high software margins offset by non-trivial compute costs for content fingerprinting at scale. Perceptual hashing of millions of images and locality-sensitive hashing of billions of text passages requires meaningful GPU/CPU time, but amortizes well across the subscriber base. LTV calculation: $4,230 x 36 months x 93% gross margin = $141,552. Payback period: $18,500 / ($4,230 Γ— 0.93) = 4.7 months.

Go-to-Market

Phase 1 (months 1–9): Start with AI companies, not publishers. This is counterintuitive but correct. AB 2013 creates an immediate, enforceable compliance obligation for every generative AI developer serving California, which is effectively all of them. The product wedge is a compliance documentation tool that auto-generates the 12 required categories of training data disclosure from structured inputs, audit logs, and dataset metadata. Sign 5–8 AI companies on the compliance module. These are not the OpenAIs of the world (they will build internally); these are the Series A through Series C AI startups building domain-specific models who lack dedicated compliance teams and face the same regulatory deadline. Once AI companies are reporting their training data provenance through your platform, publishers have a reason to join: independent verification of what their licensees are actually doing with their content.

Phase 2 (months 10–18): Bring publishers onto the demand side. With AI companies already generating provenance reports through the platform, approach mid-tier publishers who have signed licensing deals but cannot verify compliance. The initial target list comes from public deal announcements: Dotdash Meredith, Vox Media, The Atlantic, CondΓ© Nast, Hearst, TIME, and the Associated Press all signed OpenAI deals in 2024. Offer them the content fingerprinting and audit verification modules. Build the content fingerprinting library from participating publishers' archives. Integrate with training pipeline orchestration tools (Weights & Biases, MLflow, Databricks) to automate usage reporting on the AI company side.

Phase 3 (months 19–30): Introduce the royalty processing layer. By this point, the registry contains enough content and the usage data is granular enough to calculate variable royalties. Position as the neutral intermediary, owned by neither publishers nor AI companies, that both sides trust for accurate accounting. Expand internationally to serve EU publishers and AI companies subject to the AI Act's Article 53 requirements. Target the academic publishing segment (Elsevier, Springer Nature, Wiley), which controls massive proprietary datasets and is just beginning to negotiate AI licensing terms.

Competitive Field

CompanyWhat It DoesRoyalty Processing?Compliance Docs?
Spawning AIOpt-out registry (HaveIBeenTrained), Source.Plus datasetNo: tracks what creators don't want used, not what was licensedNo
Fairly TrainedLicensed Model (L) certification for AI companiesNo: binary certification, not ongoing compliance trackingNo
Data & Trust Alliance / OASISCross-industry data provenance standard (in development)No: standard-setting body, not a productPartially: provenance spec, not AB 2013-specific
Scale AI / Hugging FaceData marketplaces and labeling servicesNo: sell and host data, not compliance for licensed dataNo
Law firms (bilateral)Draft individual licensing agreementsNo: create contracts, not the infrastructure to enforce themManual, per-client
This startupRegistry + usage audit + compliance docs + royalty processingCore product: the ASCAP model for AI training dataAutomated AB 2013 + EU AI Act

The competitive gap mirrors what existed in music before ASCAP: plenty of people writing contracts (law firms), some advocacy organizations (Fairly Trained is closer to a trade association than a performing rights organization), and emerging opt-out tools. Spawning is defensive infrastructure ("don't use my stuff"), not monetization infrastructure ("use my stuff and pay me fairly"). Nobody is building the operational backbone for ongoing compliance monitoring, usage verification, and royalty calculation across a fragmented market of thousands of content owners and dozens of AI companies. Spawning raised $3 million in seed funding to build opt-out infrastructure. Necessary first step. The monetization layer that comes after opt-in ("yes, you can use my data, here are the terms, here is the invoice") is where the real revenue lives.

Why Now

Three forces are converging in a window that will not stay open long.

First, regulatory mandates have crossed the threshold from theoretical to enforceable. California's AB 2013 took effect January 1, 2026, requiring every generative AI developer making systems available to Californians to publish detailed documentation of their training data: sources, ownership, licensing status, copyright status, processing history, and 8 other categories of disclosure. Since virtually every major AI company serves California, this is effectively a national mandate. The EU AI Act's Article 53 imposes parallel requirements on general-purpose AI providers in Europe. These are not aspirational guidelines. They are laws with compliance deadlines that have already passed. Every AI company needs tooling to generate and maintain this documentation, and every publisher with a licensing deal needs tooling to verify it.

Second, the volume and complexity of licensing deals is overwhelming bilateral management. When OpenAI had three publisher deals in late 2023, each one could be managed by a BD team with a spreadsheet. By mid-2026, OpenAI alone has deals with News Corp, Axel Springer, Financial Times, AP, Vox Media, The Atlantic, Le Monde, Prisa Media, Dotdash Meredith, CondΓ© Nast, Hearst, Reddit, Shutterstock, and Stack Overflow. At minimum. Add Google, Microsoft, Meta, Amazon, and Anthropic, each assembling their own portfolio of content licenses, and the matrix of bilateral relationships becomes unmanageable. This is exactly the scaling problem that created ASCAP: when a restaurant played songs from 50 publishers, no songwriter could manage 50 separate licensing relationships. Collective infrastructure replaced bilateral chaos.

Third, the shift from flat-fee to variable-rate licensing is beginning, and variable rates require usage tracking infrastructure that does not exist. News Corp's OpenAI deal includes both fixed and variable components. Dotdash Meredith's deal has a fixed component of $16 million and a variable component to be "calculated in the future." As more publishers demand variable compensation tied to actual model usage (per-training-run fees, revenue-share on API calls that surface licensed content, performance-based bonuses), the need for metered, verifiable usage data becomes non-negotiable. ASCAP could not distribute per-play royalties to a million songwriters without automated play-count infrastructure. AI content licensing cannot scale variable royalties without equivalent automation.

Original Contribution: The Licensing Overhead Tax

A calculation nobody has published: What does it actually cost a publisher to manage an AI content licensing deal today, and how much of the deal value gets consumed by compliance and administration?

Consider a mid-tier publisher, say one of the Dotdash Meredith brands individually, with an estimated $3 million annual AI licensing deal spread across two AI company partners. Based on the operational requirements visible in AB 2013's 12 documentation categories and standard enterprise licensing audit provisions:

Compliance Cost CategoryAnnual Cost% of $3M Deal
FTE: documentation & audit coordination$145,0004.8%
Outside legal counsel$80,000–120,0002.7–4.0%
Technical staff: content cataloging & hashing$60,0002.0%
Executive time: negotiations & renewals$40,0001.3%
Total compliance overhead$325,000–365,00010.8–12.2%

Salary benchmarks from the Bureau of Labor Statistics for legal and technical roles; outside counsel rates from ALM's legal billing surveys for media companies. These are estimates, not audited publisher data, and the actual burden may be lower if publishers assign compliance as a fractional responsibility to existing legal staff.

That means the publisher is spending an estimated 10.8–12.2% of its $3 million deal value on compliance overhead. For a mid-sized publisher running thin margins after years of digital advertising decline, that is punishing. A $2,400/month SaaS subscription ($28,800/year) that automates 70–80% of that compliance burden saves the publisher $199,000–227,000 annually, a 6.9:1 to 7.9:1 return on the subscription cost. The math gets even more favorable for publishers with multiple AI company deals, because compliance costs scale linearly with the number of bilateral relationships while the SaaS subscription provides a single, multi-partner dashboard.

Scale this across the market: if 2,000 publishers are managing AI licensing deals by 2028 with an average compliance overhead of $300,000 per publisher, the total annual compliance spend across the industry is $600 million, a dead-weight cost that creates no value for either publishers or AI companies. A SaaS platform that reduces that to $150 million unlocks $450 million in recaptured value. That is the market this startup actually addresses. Not just selling subscriptions, but eliminating a structural inefficiency that grows more expensive with every new licensing deal signed.

Limitations

This analysis rests on several assumptions that could prove wrong.

First, the compliance cost calculation extrapolates from AB 2013's requirements to estimate total publisher burden, but we do not have audited data on what publishers actually spend managing AI licensing compliance. The $325,000–365,000 estimate is constructed from publicly available salary benchmarks (Bureau of Labor Statistics for legal and technical roles) and typical outside counsel rates for media companies (reported in ALM's legal billing surveys). If publishers are managing compliance with fewer resources than projected, perhaps by assigning it as a fractional responsibility to existing legal staff, the cost savings from the SaaS product shrink proportionally.

Second, the market size projection depends on deal velocity continuing to accelerate. If the major AI companies consolidate their content sourcing around a small number of large publishers and stop signing new deals, the addressable market of 3,500 publishers could plateau at a few hundred. The trend since 2023 suggests broadening, not narrowing, but a major fair use ruling in the New York Times v. OpenAI case could reduce publishers' leverage and slow deal-making.

Third, the ASCAP analogy has a structural limitation: ASCAP works because music performance rights are well-defined by statute (the Copyright Act's public performance right, Sections 106(4) and 106(6)). AI training data rights are legally contested. If courts broadly rule that model training constitutes fair use, the legal foundation for licensing could erode, and with it the need for compliance infrastructure. The Delhi High Court denied ANI's injunction request against OpenAI in July 2026, finding insufficient grounds to block training under India's fair dealing provision (Section 52(1)(a) of the Indian Copyright Act). That was an interim order, not a final ruling on the merits, and it carries no precedential weight in the U.S. or EU. But it signals how some jurisdictions may lean. The U.S. cases remain pending.

Fourth, the startup itself could become the gatekeeper it claims to displace. A dominant royalty processing intermediary would control which content gets registered, which usage counts as legitimate, and how payments flow. ASCAP required DOJ consent decrees to prevent exactly this kind of abuse. If this startup succeeds, it will face the same regulatory scrutiny, and its governance structure (investor-owned vs. cooperative, board composition, fee-setting transparency) will determine whether it serves creators or extracts from them.

Strongest Counterargument

The best case against this startup is that the AI training data licensing market may not evolve toward standardized collective licensing at all, and may instead remain a bilateral, bespoke deal market where intermediary infrastructure adds cost without sufficient value.

Consider the structural differences between music performance and AI training data. Music has a well-defined "performance event": a song plays on the radio, in a store, on Spotify. It can be monitored and counted. AI training has no equivalent discrete event. A publisher's corpus gets ingested during a training run that might last weeks, and the resulting model weights are an opaque, undifferentiated amalgam of all training data. You cannot point to a specific parameter in GPT-5 and say "that came from the Financial Times." This makes usage metering fundamentally harder than counting song plays, and potentially impossible in a cryptographically verifiable way.

The largest content licensors (News Corp at $250 million, Reddit at $60 million per year) have enough leverage and internal resources to manage bilateral relationships directly. They do not need a middleman. The publishers most likely to benefit from collective infrastructure are the smallest ones, with the lowest deal values, making the unit economics of serving them unattractive. ASCAP works partly because Taylor Swift and a coffeehouse songwriter both benefit from blanket licensing. In AI training data, the equivalent of Taylor Swift (News Corp) can negotiate directly, and the equivalent of the coffeehouse songwriter (an independent blogger) may have content too small to license at all.

Finally, the technology companies on the other side of these deals have every incentive to resist standardized compliance infrastructure that increases their audit exposure. OpenAI, Google, and Meta are unlikely to voluntarily integrate with a platform that lets publishers independently verify training data usage. If the platform cannot get both sides of the market to participate, it becomes a one-sided compliance tool: useful, but not the ASCAP-scale intermediary the market sizing assumes.

The Bottom Line

The AI content licensing market is following the same trajectory the music industry followed a century ago: a transition from informal, bilateral, unmonitored usage toward standardized licensing with compliance infrastructure and royalty distribution. The deals are being signed at a pace that will make bilateral management unsustainable within 24 months. The regulatory mandates are live. The variable-rate compensation clauses in existing deals will require usage metering that no current tool provides. What is missing is the operational layer (the registry, the audit engine, the royalty calculator) that turns a collection of bilateral contracts into a functioning market.

ASCAP took 20 years to become ubiquitous. This market is moving faster because the regulatory pressure (AB 2013, EU AI Act) is front-loaded and the deal values are orders of magnitude larger from day one. The window for building the neutral intermediary platform is now, before one of the large AI companies builds a proprietary compliance system and positions itself as the de facto standard β€” which would be the equivalent of Spotify building its own ASCAP and telling songwriters to trust the company paying them to also count the plays.

What You Can Do

If you are a publisher who has signed an AI content licensing deal: request an audit clause in your next contract renewal that specifies machine-readable usage reporting. Do not accept "we will provide annual summaries" as a compliance mechanism. Insist on quarterly or monthly automated reports with content-level granularity. Your contract is only as valuable as your ability to verify it. Second, start building a cryptographic content registry now: generate perceptual hashes of your image library and locality-sensitive hashes of your text corpus. The cost is negligible (open-source tools like pHash and MinHash do this), and having your content fingerprinted before the compliance infrastructure market matures gives you leverage in negotiations. Third, if you are managing AB 2013 verification manually, calculate your actual compliance cost per deal (FTE time, outside counsel, technical staff) and compare it to the deal value. If compliance overhead exceeds 8% of deal value, you are a customer for this product.

If you are a founder considering this space: the wedge is AB 2013 compliance for AI companies, not publisher-side royalty tracking. AI companies have an immediate, enforceable obligation to document their training data. A tool that automates that documentation and gives them an auditable, regulator-ready output is a painkiller, not a vitamin. Sign up AI companies first, then use their participation to bring publishers onto the platform. Publishers will join when AI companies are already reporting through your system, because that is when the data has independent verification value.

Related

πŸ“° Consumer AI Memory Ownership & Portability: the user-side data rights question that parallels the creator-side licensing question

πŸ“° Enterprise Model Sovereignty & Proprietary Context: the enterprise version of the same IP protection problem, applied to corporate data in AI systems

πŸ“° Billboard Yield Intelligence SaaS: another rate intelligence play in a market with opacity and fragmented pricing