Clean Clipper Add to Chrome (free)

Who it is for

Archive a source before it is changed

Pages are edited and taken down. A clip holds the full text, the source URL and the publication date the page itself declared, in a plain file you keep: not in a service that can close, change its terms or quietly lose the record.

The page that was there yesterday

A company posts a claim, you cite it, and by the time an editor reads the draft the sentence has been quietly reworded – no correction note, no timestamp, nothing to point at. Or the page has gone entirely, and the public archives never caught it because it lived for four hours on a Sunday afternoon.

Even when a page survives, the version you read may not be the version anyone else was served. Consent walls, regional variants and A/B tests mean a URL is not a stable reference to a text. What you can hold onto is what was on your screen when you read it, saved with the address and the date the page declared for itself.

Which leaves the practical problem – capture has to be instant or it does not happen. A story touches forty pages, most of them read at speed and half of them at an hour when nobody is going to open a second tool, paste a URL and wait for a service to respond. A clip takes tens of milliseconds and the measured time is shown in the corner of the window; at that cost you capture everything you read, including the thirty-nine pages that turn out not to matter and the one that does.

What the record contains

A news page as it is filedthe frontmatter of a clipped news article
---
title: "Regulator opens inquiry into data-sharing arrangement"
source: "https://apnews.com/article/…"
author: "Staff reporter"
date: "2026-03-04T09:12:00.000Z"
extraction: "dom"
---

The inquiry will examine agreements signed between 2023 and 2025, according to
a notice published on Wednesday.

Setting it up for a story

The whole point is that capture costs nothing at the moment you read. Everything below is chosen to remove a decision from that moment and move it into the settings.

  1. Open Options from the extension icon and point the destination at a folder for the story – stories/data-sharing/sources. One folder per story, decided once, so nothing has to be filed later.
  2. Set what the icon click does to “save to folder”. Unlike a research workflow, a news workflow should not open a window – the check comes after the capture, because the page may not be there when you get back to it.
  3. Set the filename template to {date}-{domain}-{title}. The capture day and the outlet are the two things you sort by when a folder holds forty pages from a fortnight.
  4. Turn on all frontmatter fields. date is the publication date the page declared; source is the exact URL in the address bar, including the query string that may identify a regional variant.
  5. Confirm Alt+Shift+M is bound at chrome://extensions/shortcuts. A keystroke gets used at eleven at night; a menu does not.
  6. Add a per-site rule for the discussion sites you work from, reddit.com into its own subfolder, so leads and published sources do not sit in the same pile.
  7. Clip a page before you contact anyone about it. A request for comment is the most reliable way to make a page change, and the capture you take afterwards is the wrong one.

Settings for capturing at speed

Every value here trades a little verification at capture time for the certainty of having captured. You can check the file at leisure; you cannot re-read a page that has gone.

SettingValueWhy this value here
Icon clickSave to folderNo window, no dialogue, no decision at the moment the page is still up
DestinationOne folder per storySourcing is per story, and a folder is what an editor or a lawyer can be shown
Filename template`{date}-{domain}-{title}`Capture day and outlet are the two axes you sort forty sources by
FrontmatterAll fields on`source` keeps the exact URL, including a query string that marks a regional variant
ImagesKeep as linksThe image URL is often the only trace of which photograph was used before it was swapped
Per-site rule`reddit.com` → subfolder `leads`Unverified leads and published sources should not be one pile
Shortcut`Alt+Shift+M`Capture that needs a menu is capture that does not happen out of hours
Two captures of one URL, three weeks apartdiff of the same page clipped twice
--- 2026-03-04-regulator-opens-inquiry.md
+++ 2026-03-25-regulator-opens-inquiry.md
@@
-date: "2026-03-04T09:12:00.000Z"
+date: "2026-03-04T09:12:00.000Z"
@@
-The inquiry will examine agreements signed between 2023 and 2025.
+The inquiry will examine agreements signed in 2024.
@@
+Updated 25 March: an earlier version misstated the period under review.

Three captures

The page that lived for four hours

A company publishes a statement on a Sunday afternoon. You read it, press the shortcut, and the file is in the story folder before you have decided whether it matters. By Monday the page returns a 404 and the public archives never saw it, because nobody submitted the URL while it was up.

What you hold is the text, the address and the date the page declared for itself: not a proof, but a record with enough in it to put the question to the company – this was published at this address, and here is what it said.

Two captures and a diff

You clip a wire story the morning it lands, and again three weeks later when a colleague mentions the framing has shifted. The second file takes a numeric suffix rather than replacing the first, so both are on disk.

A diff between them shows the period under review narrowed from “2023 and 2025” to “2024”, and an update note appended at the foot. Neither change is visible from the URL, and neither would have been recoverable from a bookmark.

A thread as a lead, kept as a thread

A local forum thread names three people who worked at a site during the period you are asking about. You clip it: the nesting comes across as nested blockquotes and each comment keeps its score, so the reply everybody agreed with is still distinguishable from the one nobody did.

It goes into the leads subfolder rather than the story folder, by rule rather than by decision. Six months later that distinction is the difference between sourcing you can point at and material you merely read.

Against the other ways to keep a page

These are not alternatives so much as different guarantees. Read the third column carefully. The strongest guarantee is not the one this extension offers.

How it is done nowWhat you getWhat it costs
A public archiving serviceAn independent, timestamped copy anyone can checkNeeds the network and the service; some sites block it, and it is a second tool at the moment you are reading
Screenshot the pageThe layout as it appearedNo searchable text, no URL inside the image, no publication date
Print to PDFA page-shaped copy on your diskConsent banner included, not diffable, not searchable across a folder
Save the page as HTMLA complete local copy with assetsHeavy, awkward to read, and a diff between two captures is unreadable
Copy into a notes documentThe words, fastThe furniture comes too, and the URL and date do not
Clean ClipperText, address and declared date, in secondsNo layout, no assets, and no third-party timestamp. It is your copy, not proof

When the capture is not what you expected

Some sites ship no article at all until consent is answered, so there is nothing in the page for any extension to read and the clip reports “no article”. Answer the dialogue, let the page render, then clip. The consent bar itself is removed from the output either way. This is also why the version you capture is specifically the version served to you, in your region, in that session.

It refuses a live blog or a section front

Both are mostly link labels, and the extension declines any page where more than about a quarter of the extracted characters sit inside links. On a live blog, clip the permalink of the individual entry where the site provides one. On a section front, clip the story rather than the list. The list is the thing that would have arrived as three hundred links.

The comments are missing

Comments are treated as page furniture and cut, with one exception – on Reddit the thread is kept, with each comment’s score. Elsewhere a comment section is indistinguishable from the promotional blocks around it, and keeping it would mean keeping those too.

The publication date is not the date I expected

The field holds the date the page declares in its own markup: JSON-LD datePublished, article:published_time, time[datetime]. A site that rewrites that value when it edits an article will report the newer date, and the extension repeats what the page says rather than second-guessing it. That is one more reason the capture date lives separately, in the filename.

What this is not

It is not a web archive. It saves the text and its metadata, not the page as rendered – no layout, no screenshots, no assets, and no third-party timestamp proving when you captured it. For a visual, independent record use an archiving service; this is faster, stays on your disk, and answers a different question. Comments are treated as page furniture and cut, except on Reddit, where the thread is kept with each comment’s score.

Add to Chrome (free)Free in full. No account, no sign-up, no limits.

Questions

Is this a substitute for a web archive?
No. It saves the text, not the page as rendered: no layout, no screenshots, no assets. For a visual record use an archiving service; for the text and its metadata this is faster and stays on your own disk.
Would a clip stand up as evidence?
On its own, no. It is a file on your computer with no third-party timestamp, so it records what you saved rather than proving what the world could see. Pair it with an independent archive when the point may be contested.
Does it capture comments?
On Reddit, yes: as a compact thread with the score of each comment. Elsewhere comments are treated as page furniture and cut.
Can it record when I clipped it?
The filename template supports {date}, which is the clip date. The publication date stays a separate frontmatter field, so the two are never confused.
Does the extension send my clips anywhere?
No. Extraction and conversion happen inside your browser; the clip is never uploaded and there is no account behind it.
Can I clip a social media post?
Sometimes. A feed is refused as “no article” because it is almost entirely link labels. A permalink page for a single post often has enough body text to extract, and on Reddit the thread comes across with its scores. Where a platform renders posts only inside a scrolling feed, there is nothing stable to capture.
Which URL ends up in the file?
The one in the address bar at the moment you clip, query string included. That matters when a site serves regional or experimental variants from parameterised URLs – the parameter is part of what you captured, and it stays in the record.
Can a colleague or an editor open the files?
Yes. They are plain text with a YAML header. No extension, no account and no particular application is needed to read them, and they can be attached to an email or dropped into a shared folder like any other document.
How quickly does a clip happen?
Tens of milliseconds for an ordinary article, a few hundred for a very long page with hundreds of reference links. The measured time appears in the corner of the clip window, and with the icon click set to “save to folder” no window opens at all.
Does it work in an incognito window?
Only if you allow the extension in incognito in Chrome’s own extension settings. Even then, folder grants are not remembered there, so the clipboard or a download is the reliable destination for a private session.