Mario Ruiz Díaz
← All case studies

Case study · Distributed systems

Frequency capping for Pluto TV ads: rules, quotas and 24 hours of impressions

A rules-driven system, shipped to production, that stops viewers from seeing the same ad over and over across ad breaks, built on Kafka for roughly 6k impression events per second with 10x headroom.

Visit pluto.tv
Product
pluto.tv ↗
Client
Pluto TV
Industry
Free ad-supported streaming (FAST)
Role
Tech Lead · Solution designer
Period
2021 – 2024
Live
in production at Pluto TV
24 h
of impressions per device
~6k/s
impression events
10x
headroom over worst case
Frequency capping for Pluto TV ads: rules, quotas and 24 hours of impressionsVIEWER APPSREAD PATHSTORAGEWRITE PATHViewer appsmobile · web · TVStitcherVOD · linear streamsadPod serviceFreeWheel pods · dedupe · filteradSelectionremaining quotas at break timeRulesgeneric · specific · prevalenceQuota counterssliding window per ruleImpression historyRedis · 24 h · ~120 GBad-beacon APIwatched-ad events · ~6k/sKafkaNEW_IMPRESSION · ~12k/sHistory writerinsert per deviceQuotas managerupdate counters per ruleClean-up jobsexpire old windowsstreamwatched adfilter by quotaentry inserted

Scroll sideways to see the full diagram →

Pluto TV frequency capping. Viewer apps report watched ads through a Kafka-based write path that maintains impression history and quota counters; the ad delivery path asks adSelection for remaining quotas at ad-break time.Download diagram (PNG) ↓

Context

Seeing the same ad again and again in the same ad break or the same hour is a bad viewer experience: people tune out when they feel they are being beaten over the head with a message. Pluto TV could already detect duplicates inside a single ad pod, but not across pods.

The challenge

Extend deduplication across ad pods with a Pluto frequency cap (the same ad at most X times per Y minutes per viewer), configurable by business rules, decided at ad-break time on the delivery path, and fed by impression events from every viewer device at production scale.

The solution

Two decoupled paths around shared storage. The write path ingests watched-ad events through Kafka, keeps 24 hours of impressions per device and updates quota counters per rule. The read path, where the ad pod service fills breaks from FreeWheel, asks an adSelection service for each device's remaining quotas, computing them from the rules on the fly when no counters exist.

  • Rules with prevalence

    Generic rules (creative, brand and industry repetition) and specific ones per brand or creative, including OFF. Specific rules override generic ones, and creative beats brand, which beats industry.

  • 24-hour impression history

    One entry per watched ad per device, with fingerprint, brand and industry, kept in Redis for 24 hours.

  • Sliding-window quota counters

    Each impression opens a window per matching rule and dimension; only impressions still inside their window count against the cap.

  • adSelection service

    Encapsulates the decision: returns remaining quotas for a device at the ad break's absolute time, so the ad pod service can filter creatives.

  • Event-driven write path

    An ad-beacon API publishes NEW_IMPRESSION events to Kafka; independent consumers insert history, update quotas and clean up expired windows.

Key decisions

  1. 01

    Decide at the ad break's absolute time

    Whether an ad is allowed depends on which impressions are still inside their window when the break plays, not when it was requested. The same history can allow one ad at 9:00 and block it at 8:30.

  2. 02

    Rules as data, with explicit prevalence

    Ad operations needed to tune repetition without deploys. Modeling rules by type (generic, specific) and dimension (creative, brand, industry) with a clear override order kept behavior predictable.

    Alternatives considered Hard-coded caps per campaign.

  3. 03

    Precompute quotas, fall back to rules

    The delivery path reads precomputed counters to stay fast; when a device has no counters yet, adSelection evaluates the rules on the fly instead of blocking the ad break.

  4. 04

    Loosely coupled through Kafka

    Services only know the event stream, not each other. Ingestion, history, quotas and clean-up scale and fail independently, and the design stays cloud-agnostic.

  5. 05

    Sized for 10x

    Capacity was planned from six months of worst-case production metrics: roughly 6k impression requests per second at the beacon, 12k messages per second on Kafka and about 120 GB of impression history, with the implementation required to support ten times those values.

What I did

  • Designed the solution: rule model, prevalence, sliding-window quotas and the decision at ad-break time.
  • Designed the architecture: decoupled read and write paths, storage and Kafka events.
  • Sized the system from production metrics, with 10x headroom.
  • Led the implementation as Tech Lead, from design review to production.

Outcomes

  • Shipped to production: frequency caps enforced across every ad break a viewer sees.
  • Deduplication extended from a single ad pod to 24 hours of impressions per device.
  • A write path built for roughly 6k impression requests and 12k Kafka messages per second.
  • Capacity planned for ten times the worst-case production load.
  • Repetition controlled by business rules that ad operations can change without deploys.

Sources & press

Contact

Bring me a real problem.

A scaling issue, a risky migration, an AI workflow that must become reliable, or a team that needs sharper technical direction.

Founder-level conversation · Practical feedback · NDA available on request