Clean Clipper Add to Chrome (free)

Who it is for

Keep reference material you can still quote

Research collected as screenshots and bookmarks stops being usable at the moment you need to attribute it. A Markdown file keeps the title, the author, the date and the address in the same file as the text, where a fact-check can reach them.

The quote you cannot attribute

The passage is in your draft, in quotation marks, and you cannot remember where it came from. Somewhere in four hundred bookmarks, or in a screenshot with no URL in it, or on a page that has since been rewritten. Reconstructing the source takes longer than the paragraph took to write, and sometimes it cannot be done at all, so the line comes out of the piece.

Pasting from the browser has its own tax. The site’s furniture arrives with the text – a byline squeezed between two share buttons, a newsletter box in the middle of a paragraph, a “read next” list where the article ended. Cleaning that up every time is why most research folders quietly become a graveyard of half-formatted fragments nobody opens again.

Then somebody else has to check it. An editor or a fact-checker cannot open your bookmarks, cannot see the subscriptions you read through and will not take a screenshot as sourcing, so the check turns into a conversation where you reconstruct from memory where each line came from. A folder of text files, each with the address and the date inside it, is something you can hand over – and handing it over is usually the difference between a query taking five minutes and taking an afternoon.

What the research folder becomes

The author travels with the passagegutenberg.org: Moby-Dick, chapter 1
---
title: "Moby-Dick; or, The Whale: Chapter 1: Loomings"
source: "https://www.gutenberg.org/files/2701/2701-h/2701-h.htm"
author: "Herman Melville"
extraction: "dom"
---

Call me Ishmael. Some years ago–never mind how long precisely–having little
or no money in my purse, and nothing particular to interest me on shore, I
thought I would sail about a little and see the watery part of the world.

Setting it up for a commission

One folder per piece, and every clip in it carrying its own attribution. The settings that matter are the ones that make the folder handable to somebody else.

  1. Make a research folder for the piece, pieces/baltic-charts/sources, before you start reading. Sourcing collected into a folder that already exists is sourcing; collected afterwards it is archaeology.
  2. Open Options from the extension icon, set the destination to your research directory, and put the piece’s name in the subfolder field. Changing that one field is how you switch commissions later.
  3. Set the filename template to {date}-{title}. That date is when you captured it, which is the fact an editor asks about when a page has changed since.
  4. Turn on every frontmatter field. author and date are the two that make a clip quotable, and source is the one a fact-checker will actually click.
  5. Leave the icon click as the preview window. It takes two seconds to see whether the byline came through, and a wrong byline is the error that survives all the way into print.
  6. When only a passage matters, select it on the page and then clip – the window opens on the selection, and the switch at the top confirms which you are looking at.
  7. Clip the page the day you read it, not the day you write. Pages are edited quietly, and the capture you did not take is the one the editor asks for.

Settings for research you can hand over

The test for every value below is whether a stranger with the folder could verify a quotation without asking you anything.

SettingValueWhy this value here
DestinationOne subfolder per pieceSourcing is per commission, and a shared folder is what a fact-checker can be given
Filename template`{date}-{title}`Records when you captured it – the question that arises when a page has since changed
FrontmatterAll fields on`author`, `date` and `source` are the attribution; the rest is context
Icon clickPreview windowTwo seconds to confirm the byline, before a wrong one gets into the draft
ImagesKeep as linksA photo credit is often only findable from the image URL
SelectionSelect the passage, then clipA long feature quoted for one paragraph does not need to be kept whole
Per-site ruleSites you use constantly → their own subfolderBackground reading and primary sources are different piles at the fact-check stage
A pull-quote stays a quote, and the byline is the bylinea long-form magazine feature
---
title: "The last cartographers of the Baltic"
source: "https://example-magazine.com/features/baltic-charts"
author: "Ingrid Salo"
date: "2025-11-18T06:00:00.000Z"
extraction: "dom"
---

The survey vessel still carries paper charts, folded into a drawer nobody has
opened since the second officer retired.

> We stopped correcting them in 2019. Nobody decided to stop; the corrections
> simply stopped arriving.

The hydrographic office confirmed that paper corrections ended that year.

Three moments in a piece

The quote whose page has been rewritten

You clipped a company statement in November and quote it in a February draft. The page still exists and no longer says that – the wording was softened, with no correction note and no timestamp on the change. Your file has the original sentence, the URL and the date the page declared for itself.

Clip the page again and any diff tool shows the two versions side by side. That is a different claim from “the page said this” – it is “here is what I captured, and here is what is there now” – and it is a claim you can put in front of an editor.

Handing the sourcing to somebody else

The fact-checker gets the folder. Each file opens in any editor with the passage, the byline, the publication date and the address in the first lines, so the checks that would have been questions to you are lookups they can do alone.

This is where screenshots fail hardest. A picture of a page carries no URL, no date and no text a checker can search, so every screenshot in a research folder becomes a message asking where it came from.

Background before an interview

Six pieces of background reading, clipped into the piece’s folder the evening before. During the call you search the folder rather than eight open tabs, and the search covers the words in all six at once, including the two behind subscriptions.

Because the files are text, that search costs nothing and works with the wifi off in a building with no signal. The tabs would have gone stale, been reloaded, asked for the consent banner again, or logged you out.

Against the usual research habits

Nothing here is a bad habit – each of these is somebody’s working method. What differs is what survives to the fact-check.

How it is done nowWhat you getWhat it costs
Bookmark the pageA pointer, instantlyIt points at whatever the page says today, not what it said when you read it
Screenshot the passageThe words, visuallyNo URL, no date, no searchable text, and nothing a checker can verify alone
Paste into a research documentThe text, in one placeSite furniture comes with it and the attribution does not
A read-later serviceA clean copy, syncedDepends on the service continuing to exist and to let you export
Print the page to PDFA fixed record of the layoutConsent banner included, text not searchable across a folder, no metadata header
Clean ClipperText, byline, date and address in one fileNo layout, no screenshots, and no third-party proof of when you captured it

When a clip is not quotable

The author field is empty

The page states no byline the extension recognises, or the byline sits in markup indistinguishable from a section heading. Author extraction reads structured data and byline markup and then filters the known false positives (headings, publication names, button labels) and when nothing survives that filter the field is left empty rather than filled with the first bold string on the page. Fill it in yourself; it is a text file.

Only the first three paragraphs came through

That is usually a metered paywall – the page genuinely rendered a teaser and loaded the rest only for a session it did not think you had. It can also be lazy loading, where the body arrives as you scroll. Scroll to the end of the article, confirm the last paragraph is on screen, and clip again. The extension captures what the browser has rendered at that moment.

The article never appeared behind the consent dialogue

Some sites render nothing at all until consent is given, so there is no article in the DOM to extract and the extension will say so. Answer the dialogue however you normally would, let the page render, then clip. The consent bar itself is removed from the clip either way.

The filename date and the frontmatter date do not match

They are not the same date and are not meant to be. {date} in the filename is the day you captured the page; the date field in the frontmatter is the publication date the page declared for itself. Keeping them apart is what lets you say both when the piece was published and when you saw it – and an empty frontmatter date means the page declared none.

What it does not do

It does not archive the page as it looked: no screenshots, no layout, no assets. It does not summarise or paraphrase. The article text is converted, not edited. Images are kept as links to the original site, so they break when that site does. And when a page carries no author or no date, the field is left empty rather than filled with a plausible guess.

Add to Chrome (free)Free in full. No account, no sign-up, no limits.

Questions

Does it keep the author?
When the page states one. Clean Clipper reads structured data and byline markup and filters the false positives (section headings, publication names, button labels) rather than putting the first plausible string into the field.
What about articles behind a subscription I pay for?
They clip normally. The extension reads the page your browser rendered for you, so anything you can see on screen can be saved.
Can I keep the images?
Yes, as links by default. You can switch them off globally or for one site.
Does it change the wording?
No. The structure changes, HTML becomes Markdown, but no words inside the article body are added, removed or reordered.
Is it free?
Yes, all of it. There is no paid tier, no account and no limit on how much you clip.
Do italics and links inside a quotation survive?
Yes. Bold, italic, links, inline code and lists map to their Markdown equivalents, and a pull-quote or blockquote stays a blockquote – so a passage that was marked as somebody else’s words still looks like somebody else’s words in your file.
Does a clip prove that the page said this?
No, and it should not be presented as proof. It is a file on your own computer with no third-party timestamp – it records what you saved, which is enough for your own working notes and not enough for a contested claim. Where a point may be disputed, pair it with an independent archive.
Can I clip a newsletter or an email?
Only where it is published as a web page – many newsletters have a web archive, and that page clips like any other. An email in a mail client is not a web page the extension can reach.
May I republish what I clip?
That is a copyright question about the source, not about the tool. Clipping is a reading and note-taking act; what you may quote, and how much, is governed by the licence and the law exactly as it would be if you had copied the passage by hand.
Can I clip the same page repeatedly as a story develops?
Yes, and each clip is a separate file – a name collision takes a numeric suffix instead of overwriting. Two captures of a developing page, compared in any diff tool, are the record of what was added and what was quietly removed.