Every transcript CNN published at transcripts.cnn.com from 2000-01-01 to 2025-03-15: 336,902 show segments, each with program, subhead, air date and time, and the full text. This repository holds the scraper that produced the corpus and the record of how each era of CNN's site was parsed. The corpus itself lives on Harvard Dataverse at doi:10.7910/DVN/ISDPJU; for copyright reasons access is restricted to research use.
The Dataverse dataset holds eight CSV files, one per scraping run:
| File | Aired | Segments |
|---|---|---|
cnn-1.csv |
2000-01-01 to 2000-04-20 | 7,017 |
cnn-2.csv |
2000-04-21 to 2001-04-03 | 21,381 |
cnn-3.csv |
2001-04-04 to 2002-08-06 | 35,269 |
cnn-4.csv |
2002-08-07 to 2002-09-16 | 2,343 |
cnn-5.csv |
2002-09-17 to 2012-05-18 | 101,336 |
cnn-6.csv |
2012-05-19 to 2014-06-17 | 23,536 |
cnn-7.csv |
2014-06-18 to 2022-02-05 | 102,458 |
cnn-8.csv |
2022-02-01 to 2025-03-15 | 43,562 |
cnn-7 and cnn-8 overlap by five days; cnn-transcripts to-parquet (below) drops the repeats by URL. Two segments from July 2024 were lost to connection errors in the 2025 run; the scraper now retries and reports such URLs so they can be re-run.
Columns. The CSVs use the original scraper's columns. The Parquet build renames them and drops two that carried no information:
| CSV | Parquet | Notes |
|---|---|---|
url |
url |
As fetched. Old form .../TRANSCRIPTS/0507/12/lkl.01.html, new form .../show/acd/date/2025-03-14/segment/01. |
program.name |
program |
Show name as printed, e.g. CNN LARRY KING LIVE. |
subhead |
subhead |
Segment title. From 2024 CNN appends the slot ("Aired 8-9p ET"); the scraper strips it. |
year, month, date |
aired_date |
date32. |
time |
aired_time |
HH:MM, 24-hour, as printed. |
timezone |
timezone |
Almost always ET. |
uid |
uid |
<show code>.<segment>, e.g. lkl.01, acd.01. |
path |
path |
URL path after the host, e.g. 0507/12/lkl.01.html. The CSVs kept only the last two segments. |
wordcount |
wordcount |
int32, recomputed from text where the CSV left it blank. |
text |
text |
Paragraphs joined by newlines. The "THIS IS A RUSH TRANSCRIPT" and "TO ORDER A VIDEO" notices are removed. |
channel.name |
dropped | Constant. |
duration |
dropped | Never populated. |
source |
Input file the row came from. |
CNN began posting transcripts online around 1999-10-01, and its own index for late 1999 still exists at edition.cnn.com/TRANSCRIPTS/1999.10.01.html, but every transcript it links to returns "Page not found", so the corpus starts on 2000-01-01.
The Internet Archive does not extend this. Its CDX index holds only two distinct CNN transcript URLs captured before 2000 (/TRANSCRIPTS/9812/16/blair/ and /TRANSCRIPTS/9904/28/pin.00.html), and the eight 1999 daily index pages it captured are mostly redirects. Archive Team's CNN Transcript Collection 2000-2012 (1 GB) overlaps the range already covered here and could serve as a completeness check, not an extension. Pre-2000 CNN transcripts exist in full text only in Nexis Uni, which requires institutional access.
The scraper visits the daily index https://transcripts.cnn.com/TRANSCRIPTS/YYYY.MM.DD.html, follows every transcript link, and parses the page. CNN changed the markup twice, so the parser dispatches on what it finds:
| Aired | Index page | Transcript page |
|---|---|---|
| 2000-01-01 to 2001-04-03 | <LI><A HREF> list |
<h2> program, <h3> subhead, bare "Aired ..." text, <p> paragraphs |
| 2001-04-04 to 2002-09-16 | <LI><A HREF> list |
<h2> program, <h4> subhead, body in a following <table> split by <br> |
| 2002-09-17 onward | div.cnnTransDate + div.cnnSectBulletItems |
p.cnnTransStoryHead, p.cnnTransSubHead, p.cnnBodyText (first is the "Aired" line) |
From 2024 the links changed to /show/<code>/date/<YYYY-MM-DD>/segment/<NN>; the page markup did not. Each layout has a fixture in tests/fixtures/ taken from the Wayback Machine (see tests/fixtures/SOURCES.md) and a test asserting the parsed fields.
The scripts that produced the eight CSVs are preserved at commit 26aa2e1. They were rewritten in 2026 into the package here, fixing bugs that affected the CSVs at the margins: a two-digit-hour regex that dropped the air time on single-digit hours, variables that leaked a previous page's date into a page with no "Aired" line, and body extraction that kept only the last paragraph in some layouts.
uv sync --group dev
# Scrape a date range into an append-only JSONL checkpoint. Re-running the same
# command skips URLs already in the file, so a crash costs nothing.
uv run cnn-transcripts scrape --start 2025-03-16 --end 2025-03-31 --out data/cnn.jsonl
# Combine legacy CSVs and JSONL into one typed Parquet file, deduplicated by URL.
uv run cnn-transcripts to-parquet data/cnn-*.csv data/cnn.jsonl --out data/cnn_transcripts.parquet
# Add a file to the Dataverse dataset (needs DATAVERSE_API_TOKEN in the environment).
uv run cnn-transcripts upload data/cnn_transcripts.parquetscrape defaults to 60 requests per minute, three retries, and a 30-second timeout. It exits non-zero and lists the failed URLs if any transcript could not be fetched.
Development: uv run ruff check . && uv run ruff format --check . && uv run pytest.
See CITATION.cff. Please cite the Dataverse DOI for the data.
Code is MIT. The transcripts are CNN's; the Dataverse copy is restricted to research use.
- notnews/fox_news_transcripts — Fox News Transcripts 2003--2025
- notnews/msnbc_transcripts — MSNBC Transcripts: 2008–2022
- notnews/archive_news_cc — Closed captions from Internet Archive TV News, 2009–2023
- notnews/nbc_transcripts — NBC-hosted MSNBC transcripts 2008–2014
- notnews/stanford_tv_news — Stanford Cable TV News Dataset
✨ Powered by Adjacent 🚀
This is a point-in-time data collection; see the shared maintenance policy. Run the affected parser tests when code changes and the relevant data validators when inputs or outputs change. Full-data checks and publication are explicit operations. Routine edits do not require hosted CI, Docker, a Python-version matrix, Preen or pre-commit.