System Design: News Aggregator (100K Sources, Dedup, Personalized Ranking)
Goal: Build a news aggregation platform that:
- Crawls 100K sources (RSS/Atom feeds, news APIs).
- Ingests 5 million articles per day.
- Deduplicates near-identical stories.
- Ranks articles by relevance and freshness.
- Serves personalized feeds to 50 million daily active users.
- Supports breaking news detection with alerts, category classification, and multi-language content.
1. The Three Problems
A news aggregator that checks 100,000 RSS feeds every five minutes makes 333 feed fetches every second, day and night. Most of those fetches find nothing new, because the source has not published since the last check.
News aggregation looks simple until you put numbers on it. Three problems shape every decision in this design.
Problem 1: Efficient source polling at scale. 100K RSS/Atom feeds need to be checked regularly. Fixed-interval polling at 5 minutes means 100K / 300s = 333 feed fetches per second.
Each fetch is an HTTP request that may return zero to fifty new articles. Most fetches return nothing new because the source has not published since the last check.
Polling too frequently wastes bandwidth and risks getting the crawler IP-banned. Polling too infrequently means missing breaking news by hours.
The New York Times publishes 200+ articles per day. A niche tech blog publishes 2 per week. Polling both every 5 minutes means 288 daily polls to the NYT (productive, finds new articles on most checks). It also means 288 daily polls to the blog (wasteful, finds something new once per 288 * 7 = 2,016 polls).
Problem 2: Deduplication is not exact matching. When a major event occurs, 500 news outlets publish stories about it within an hour. These stories cover the same event but have different headlines, different wording, and different details.
Simple exact-match dedup (matching on title or URL) misses 90% of duplicates. "Trump wins election" and "Donald Trump elected president" describe the same event.
The system needs fuzzy, near-duplicate detection that groups articles covering the same story into clusters.
Problem 3: Ranking without engagement signals. Unlike social media where likes, shares, and comments provide ranking signals, news articles arrive cold. There is no user engagement data for a brand-new article.
Ranking must rely on source authority, freshness, topic importance, and content quality. Engagement data only becomes useful after the article has been served to enough users, creating a cold-start problem for every single article.
Scale numbers:
- 100K RSS/Atom sources
- 5M raw articles/day ingested
- ~2M unique stories/day (after dedup, ~60% are near-duplicates)
- 50M DAU
- 10K feed requests/sec at peak (personalized feed reads)
- Average article record: ~1.2KB (metadata only, not full text)
2. Requirements
Functional Requirements
| ID | Requirement | Priority |
|---|---|---|
| FR-01 | Crawl 100K RSS/Atom feeds on adaptive schedules | P0 |
| FR-02 | Ingest and parse articles (title, summary, author, published date, category) | P0 |
| FR-03 | Deduplicate near-identical articles (same story, different sources) | P0 |
| FR-04 | Rank articles by freshness, source authority, topic relevance | P0 |
| FR-05 | Personalized feed per user based on reading history and interests | P1 |
| FR-06 | Breaking news detection and push notification | P1 |
| FR-07 | Category classification (politics, sports, tech, business, etc.) | P0 |
| FR-08 | Multi-language support (English, Spanish, French, German, Japanese) | P1 |
| FR-09 | Search across articles by keyword | P1 |
| FR-10 | Trending topics (most-covered stories in the last hour) | P1 |
| FR-11 | Source management (add, remove, configure sources) | P1 |
| FR-12 | Read history tracking and "read" markers | P1 |
| FR-13 | Topic following (user subscribes to specific topics) | P2 |
| FR-14 | Bookmarking/saving articles for later | P2 |
Non-Functional Requirements
| ID | Requirement | Target |
|---|---|---|
| NFR-01 | Feed generation latency (p50 / p99) | < 50ms / < 200ms |
| NFR-02 | Article ingestion-to-available | < 15 minutes (normal), < 5 minutes (breaking news) |
| NFR-03 | Throughput (feed reads) | 10K personalized feeds/sec |
| NFR-04 | Deduplication accuracy | > 90% precision, > 85% recall |
| NFR-05 | Availability | 99.99% |
| NFR-06 | Crawler politeness | Respect robots.txt, min 1s delay between requests to same domain |
| NFR-07 | Article retention | 30 days (full), 1 year (metadata only) |
| NFR-08 | Breaking news detection latency | < 5 minutes from event to notification |
3. Scale Estimation
🔒 Premium section
4. Why Naive Approaches Fail
🔒 Premium section
5. Architecture Overview
🔒 Premium section
6. API Design
🔒 Premium section
7. Data Model
🔒 Premium section
8. Partitioning Strategy
🔒 Premium section
9. Caching Strategy
🔒 Premium section
10. Consistency Model
🔒 Premium section
11. Technology Selection
🔒 Premium section
12. Queue and Stream Capacity Planning
🔒 Premium section
13. System Flows
🔒 Premium section
14. Adaptive Polling
🔒 Premium section
15. The Ingestion Pipeline
🔒 Premium section
16. Near-Duplicate Detection with MinHash + LSH
🔒 Premium section
17. Topic Classification
🔒 Premium section
18. Article Ranking
🔒 Premium section
19. Breaking News Detection
🔒 Premium section
20. Personalized Feed Assembly
🔒 Premium section
21. Background Reconciliation
🔒 Premium section
22. Bottlenecks and Mitigations
🔒 Premium section
23. Component-Level Failure Modeling
🔒 Premium section
24. Deployment and Operations
🔒 Premium section
25. Observability
🔒 Premium section
26. Security
🔒 Premium section
27. Testing and Validation
🔒 Premium section
28. Cost and Capacity
🔒 Premium section
29. Multi-Region Considerations
🔒 Premium section
30. Three Things That Break at 10x
🔒 Premium section
31. Explore the Technologies
🔒 Premium section