Measurement
What was measured, on what, and what came out
Claims about clean output are worth nothing without a corpus and a script. Both are published.
Run one: 109 pages, four engines
The corpus is the hundred most-visited sites in the world and the hundred most-visited in Russia. Non-content domains (messengers, banking, streaming, government portals) were excluded in advance. Pages were captured with a real browser, resolving a content link from each home page where one existed.
Four engines ran over the same captured DOM: Defuddle (the core of the official Obsidian clipper), Mozilla Readability, the real shipped code of a third Markdown clipper unpacked from its store build, and Clean Clipper.
| Metric | Engine A | Engine B | Engine C | Clean Clipper |
|---|---|---|---|---|
| Pages | 109 | 107 | 107 | 109 |
| Failures | 0 | 2 | 2 | 0 |
| Leftover HTML | 532 | 718 | 2 | 0 |
| Duplicated menu lines | 491 | 282 | 478 | 102 |
| Short-line ratio | 0.278 | 0.262 | 0.263 | 0.220 |
Where Clean Clipper is behind: it found 20 authors against 28 for the best engine, and extracted 4 tables against 11, though 41 of the 56 source tables sit on a single page and are layout, not content, which no engine should extract. Its share of “useful text” is lower because it refuses feed pages that other engines happily return.
Run two: 403 pages in twelve European languages
The top fifty sites in each of twelve countries, captured with the browser set to that country’s language, because otherwise half of them serve an English page and the measurement is of something else.
| Language | Pages | Failures | Leftover HTML | Link share |
|---|---|---|---|---|
| German | 40 | 0 | 0 | 0.058 |
| French | 39 | 0 | 0 | 0.050 |
| Spanish | 40 | 0 | 0 | 0.060 |
| Italian | 40 | 0 | 0 | 0.036 |
| Polish | 30 | 0 | 0 | 0.064 |
| Portuguese | 36 | 0 | 0 | 0.060 |
| Dutch | 37 | 0 | 0 | 0.049 |
| Turkish | 31 | 0 | 0 | 0.042 |
| Ukrainian | 13 | 0 | 0 | 0.054 |
| Czech | 28 | 0 | 0 | 0.039 |
| Swedish | 28 | 0 | 0 | 0.036 |
| Romanian | 41 | 0 | 0 | 0.039 |
What the numbers mean
- Leftover HTML tags – how much raw markup survived into the Markdown. Anything above zero is a defect.
- Duplicated menu lines – lines longer than twelve characters that appear more than once. This is navigation repeated in every section.
- Short-line ratio – the share of lines under twenty-five characters. High means clipped link lists rather than prose.
- Link share – the share of characters sitting inside link labels. High means a menu came along.
Honest limits of this measurement
- The Ukrainian sample is 13 pages, not 40 – most of the corpus did not respond from the capture location. Treat that row as indicative only.
- Turkish and Czech top-fifty lists skew toward shops and timetables, which carry no publication date at all. That is a property of the corpus rather than of the extraction.
- Competing engines were run only on the first corpus. The European run measures Clean Clipper alone.
- Pages where every engine returned under 500 characters were excluded – that is a captcha or a block, not extraction quality.