Case study · Distributed systems
Frequency capping for Pluto TV ads: rules, quotas and 24 hours of impressions
A rules-driven system, shipped to production, that stops viewers from seeing the same ad over and over across ad breaks, built on Kafka for roughly 6k impression events per second with 10x headroom.
- Product
- pluto.tv ↗
- Client
- Pluto TV
- Industry
- Free ad-supported streaming (FAST)
- Role
- Tech Lead · Solution designer
- Period
- 2021 – 2024
- Live
- in production at Pluto TV
- 24 h
- of impressions per device
- ~6k/s
- impression events
- 10x
- headroom over worst case
Scroll sideways to see the full diagram →
Context
Seeing the same ad again and again in the same ad break or the same hour is a bad viewer experience: people tune out when they feel they are being beaten over the head with a message. Pluto TV could already detect duplicates inside a single ad pod, but not across pods.
The challenge
Extend deduplication across ad pods with a Pluto frequency cap (the same ad at most X times per Y minutes per viewer), configurable by business rules, decided at ad-break time on the delivery path, and fed by impression events from every viewer device at production scale.
The solution
Two decoupled paths around shared storage. The write path ingests watched-ad events through Kafka, keeps 24 hours of impressions per device and updates quota counters per rule. The read path, where the ad pod service fills breaks from FreeWheel, asks an adSelection service for each device's remaining quotas, computing them from the rules on the fly when no counters exist.
Rules with prevalence
Generic rules (creative, brand and industry repetition) and specific ones per brand or creative, including OFF. Specific rules override generic ones, and creative beats brand, which beats industry.
24-hour impression history
One entry per watched ad per device, with fingerprint, brand and industry, kept in Redis for 24 hours.
Sliding-window quota counters
Each impression opens a window per matching rule and dimension; only impressions still inside their window count against the cap.
adSelection service
Encapsulates the decision: returns remaining quotas for a device at the ad break's absolute time, so the ad pod service can filter creatives.
Event-driven write path
An ad-beacon API publishes NEW_IMPRESSION events to Kafka; independent consumers insert history, update quotas and clean up expired windows.
Key decisions
- 01
Decide at the ad break's absolute time
Whether an ad is allowed depends on which impressions are still inside their window when the break plays, not when it was requested. The same history can allow one ad at 9:00 and block it at 8:30.
- 02
Rules as data, with explicit prevalence
Ad operations needed to tune repetition without deploys. Modeling rules by type (generic, specific) and dimension (creative, brand, industry) with a clear override order kept behavior predictable.
Alternatives considered Hard-coded caps per campaign.
- 03
Precompute quotas, fall back to rules
The delivery path reads precomputed counters to stay fast; when a device has no counters yet, adSelection evaluates the rules on the fly instead of blocking the ad break.
- 04
Loosely coupled through Kafka
Services only know the event stream, not each other. Ingestion, history, quotas and clean-up scale and fail independently, and the design stays cloud-agnostic.
- 05
Sized for 10x
Capacity was planned from six months of worst-case production metrics: roughly 6k impression requests per second at the beacon, 12k messages per second on Kafka and about 120 GB of impression history, with the implementation required to support ten times those values.
What I did
- Designed the solution: rule model, prevalence, sliding-window quotas and the decision at ad-break time.
- Designed the architecture: decoupled read and write paths, storage and Kafka events.
- Sized the system from production metrics, with 10x headroom.
- Led the implementation as Tech Lead, from design review to production.
Outcomes
- Shipped to production: frequency caps enforced across every ad break a viewer sees.
- Deduplication extended from a single ad pod to 24 hours of impressions per device.
- A write path built for roughly 6k impression requests and 12k Kafka messages per second.
- Capacity planned for ten times the worst-case production load.
- Repetition controlled by business rules that ad operations can change without deploys.