How every score is calculated
Every number on SoonViral comes from a named open source and a timestamp. Scores are plain formulas, the weights below are the ones the code uses (scoring version 2026-10-1), and each topic page shows its own breakdown. We say a topic is rising and show the evidence; we never claim it will go viral.
1. Growth per source
For each topic and each source we compare the recent window with the topic's own baseline, using robust statistics so a single spike day does not distort it.
spike = mean(last 3 days) / max(median(days 4–31), 1)
robust z = (mean(last 3 days) − median(baseline)) / max(1.4826 × MAD(baseline), 0.1 × median, 1)
acceleration = ln((sum of last 7 days + 1) / (sum of previous 7 days + 1))
A source shows a spike when spike ≥ 1.5× and robust z ≥ 2. Spikes are capped at 10× and mapped to 0–1 on a log scale (1× = 0, 10× = 1). Relative change beats raw volume everywhere: a small topic doubling counts more than a big topic growing 5%.
2. Trend score (0–100)
trend = 100 × ( 0.3 × attentionGrowth
+ 0.25 × coverageGrowth
+ 0.2 × crossSourceBreadth
+ 0.15 × acceleration
+ 0.1 × freshness
− 0.25 × saturation )- attentionGrowth: normalised spike in wikipedia, google trends (Google Trends only once access is approved).
- coverageGrowth: mean of the strongest three normalised spikes among gdelt, arxiv, pubmed, hacker news, podcast index, rss, github, hugging face, reddit.
- crossSourceBreadth: independent sources spiking ÷ 3, so three sources at once earns full marks and one source alone at most a third.
- acceleration: mean week-over-week acceleration across sources ÷ ln 2 (doubling earns full marks), clamped to 0–1.
- freshness: 1 on the day a topic is first detected, falling linearly to 0 after 45 days.
- saturation: 0.5 × volume percentile within the niche + 0.3 × days elevated (full at 28) + 0.2 × “in Wikipedia's overall top articles”. A day counts as elevated above 1.3× the topic's quiet level.
If a source is unavailable, it is left out and the weights of the remaining components are scaled up so the score stays on the same scale. The topic page then says which source was missing.
3. Lifecycle stage
Rules are checked in this order; the first match wins.
- DecliningMost sources have a 3-day moving average that fell for 5+ consecutive days, and acceleration is negative.
- ExplodingSpiking in at least 2 sources, a peak spike of 3× or more, and acceleration ≥ 0.35 (about +42% week over week).
- SaturatedVolume in the niche's top 40%, elevated for 21+ days, and flat or falling (acceleration ≤ 0.05).
- Volume in the niche's top 15%, at least 2× its baseline, with acceleration near zero (|acceleration| ≤ 0.25).
- EmergingA spike of 1.5× or more, first detected within 21 days, volume below the niche's 60% percentile.
- SteadyAnything else: no clear rise or fall right now.
4. Opportunity score (0–100)
opportunity = trend × (1 − saturation) × stageWeight × nicheRelevance
Stage weights reward getting in early:
Emerging 1 · Exploding 0.9 · Viral 0.55 · Saturated 0.25 · Declining 0.15 · Steady 0.4
Niche relevance (0–1) combines keyword overlap (0.5), Wikipedia category membership (0.3) and a cached AI classification (0.2). Topics below 0.45 are dropped from the niche.
5. Confidence and labels
confidence = 0.45 × min(sources / 4, 1)
+ 0.25 × data completeness
+ 0.3 × min(history days / 31, 1)Below 45% confidence a topic is labelled Early signal. With fewer than 14 days of history it is labelled New topic, limited history. Topic pages are only indexed by search engines once they have at least 2 sources and 14 days of data.
Topics appear on a radar from a trend score of 20. AI summaries and content angles are written from a trend score of 35, and rewritten only when the score moves by more than 10 points or the stage changes.
6. Guardrails and integrity
- Topics about deaths, disasters, violence, tragedies or elections are marked sensitive: they may appear with a neutral label, but get no opportunity score, no content ideas and no sponsors or affiliate links nearby.
- Health topics: rely on evidence from reputable sources, and do not present a trend as a treatment. Generated text is checked for medical claims and, in finance niches, for investment advice.
- AI writes only from a structured evidence file for each topic or issue. A validator checks every number in the draft against that evidence and rejects mismatches. Output separates what was observed from our interpretation.
- Sponsors and affiliates never influence which topics appear, their scores, rankings or AI-written text. Sponsor blocks are labelled “Sponsored” and affiliate links are disclosed.
- For news and articles we store and show only headlines (up to 140 characters), publisher, date and link. We never republish article or paper text, and never hotlink publishers' images.
7. Data sources
Only official APIs and open datasets whose terms permit commercial use. Every request carries a descriptive User-Agent with contact details and respects the published rate limits.
Analytics API data is CC0 1.0. We send a descriptive User-Agent with contact details, request human traffic (agent=user), send one request at a time and stay under the 200 requests/minute limit for identified clients.
Attribution: Attention data: Wikipedia pageviews via the Wikimedia Analytics API · Licence: CC0 (pageview data)
GDELT data is free for any use including commercial, provided every use cites the GDELT Project and links to gdeltproject.org. We use the open 2.0 files and DOC API, never GDELT Cloud.
Attribution: News data: GDELT Project · Licence: Free for commercial use with citation
arXiv descriptive metadata (titles, abstracts, authors, categories) is CC0. We use titles and links only and never the paper content. One request every 3 seconds, one connection.
Attribution: Research data: arXiv (thank you to arXiv for use of its open access interoperability) · Licence: Metadata CC0
Uses only the official Firebase API. We store story titles, points, comment counts and links, never comment text. Commercial use is not explicitly addressed by the API docs: see docs/DATA_SOURCES.md.
Attribution: Community data: Hacker News (Y Combinator) official API · Licence: Public API; titles, scores and links only
We show headlines (truncated), publisher names, dates and links only, never article text or images. Feeds whose terms forbid reuse must not be configured.
Attribution: Headlines: niche publications via their public feeds (headline, date and link only) · Licence: Headline and link only
We use record counts and article titles with links only. Abstracts may carry publisher copyright and are never stored or shown.
Attribution: Research data: PubMed, courtesy of the U.S. National Library of Medicine · Licence: NLM terms; metadata only
Episode titles, podcast names, dates and links only. Enable after confirming the current API terms (docs/DATA_SOURCES.md).
Attribution: Podcast data: Podcast Index · Licence: Podcast Index API terms; titles and links only
Not used until the official API application is approved. No unofficial libraries (e.g. pytrends) or scraping.
Attribution: Search interest: Google Trends · Licence: Google Trends API terms (by approval)
Repository counts, names and links only; no user or contributor personal information.
Attribution: Developer activity: GitHub API · Licence: GitHub API terms; repository metadata only
Model ids, counts and links only; never user profiles.
Attribution: Model activity: Hugging Face Hub API · Licence: Hub API terms; model metadata only
Official OAuth Data API only, never scraping. Off until Reddit approves commercial use (FEATURE_REDDIT + credentials). Stores post titles, permalinks and daily counts only; no bodies, comments or usernames; removed posts are purged; no model training.
Attribution: Community data: Reddit official Data API · Licence: Reddit Data API Terms, under a commercial agreement; titles and links only
Never used: YouTube, Instagram, TikTok, X, Reddit and Twitch data, scraping of Google or any site whose terms forbid it, unofficial libraries such as pytrends, the Algolia Hacker News search API, and GDELT Cloud.
8. Limitations
- Open sources measure attention and coverage, not views or sales on any social platform. A topic can rise on TikTok without showing up here, and the reverse.
- Wikipedia pageviews arrive about a day late, so the most recent day is yesterday.
- Merging synonyms (“GLP-1” vs “semaglutide”) is automated and occasionally wrong. Use “Report an issue” on any topic page.
- New topics have little history; the labels above say so. Public pages show the last 30 days.
- A high score means attention is rising now. It is not advice, and it is not a forecast.
Questions or corrections: contact us.