{"id":4036,"date":"2026-07-25T08:03:56","date_gmt":"2026-07-25T07:03:56","guid":{"rendered":"https:\/\/21percent.org\/?p=4036"},"modified":"2026-07-25T10:41:14","modified_gmt":"2026-07-25T09:41:14","slug":"the-tail-tells-the-story-extreme-textual-overlap","status":"publish","type":"post","link":"https:\/\/21percent.org\/?p=4036","title":{"rendered":"The Tail Tells the Story: Extreme Textual Overlap"},"content":{"rendered":"\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"644\" src=\"https:\/\/21percent.org\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-25-at-07.38.22-1024x644.png\" alt=\"\" class=\"wp-image-4037\" srcset=\"https:\/\/21percent.org\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-25-at-07.38.22-1024x644.png 1024w, https:\/\/21percent.org\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-25-at-07.38.22-300x189.png 300w, https:\/\/21percent.org\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-25-at-07.38.22-768x483.png 768w, https:\/\/21percent.org\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-25-at-07.38.22-1536x965.png 1536w, https:\/\/21percent.org\/wp-content\/uploads\/2026\/07\/Screenshot-2026-07-25-at-07.38.22-2048x1287.png 2048w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<hr class=\"wp-block-separator has-alpha-channel-opacity\"\/>\n\n\n\n<p>There is now a lot of press coverage of in <a href=\"https:\/\/www.thetimes.com\/uk\/education\/article\/jason-arday-plagiarism-cambridge-nathan-cofnas-q8q6f8lpx\" title=\"The Times\">The Times<\/a>, <a href=\"https:\/\/www.telegraph.co.uk\/news\/2026\/07\/24\/cambridges-diversity-poster-boy-in-plagiarism-row\/\" title=\"the Telegraph \">the Telegraph <\/a>and <a href=\"https:\/\/www.dailymail.com\/news\/article-16003587\/Plagiarism-jealousy-youngest-black-professor-Cambridge-life-story.html\" title=\"the Mail\">the Mail<\/a>.<\/p>\n\n\n\n<p>From a mathematician&#8217;s or programmer&#8217;s perspective, we highlight an outstanding piece of work by Alexey Titorenko, available on GitHub<a href=\"https:\/\/titorenko.github.io\/analysethat\/posts\/thesis-overlap-analysis\/\" title=\" here\"> here<\/a>. It presents a fully reproducible, embedding-based analysis of the PhD theses by Arday (2015) and Zwozdiak-Myers (2009).<\/p>\n\n\n\n<p>Wonderfully, the complete pipeline is included end-to-end: extraction, segmentation, embedding, control construction, Poisson testing, provenance classification, and extent estimation. The code automatically retrieves all source documents from their respective repositories. Equally remarkably, the entire process runs on a standard CPU in approximately twenty minutes.<\/p>\n\n\n\n<p><strong>The two PhD theses have considerable overlap.<\/strong><\/p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>&#8220;The signal is in the tail, and it survives controls. Against two independent education doctorates matched on discipline and partially on topic, the focal pair shows:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>9 near-identical sentence pairs<\/strong>&nbsp;(cosine \u2265 0.999) in 11.09M comparisons, against&nbsp;<strong>1<\/strong>&nbsp;in 34.29M control comparisons &#8211; and that single control match is the standard thesis declaration (\u201cThis work has not previously been presented for an award at this, or any other, university\u201d).<\/li>\n\n\n\n<li><strong>179 sentence pairs sharing 12 or more consecutive identical words<\/strong>&nbsp;(107 spatially independent blocks), against an expected 0.32 under the null &#8211; an exact count over all 11.09M pairs, no embedding model involved. Longest shared run: 63 words.<\/li>\n\n\n\n<li>Conditional on high semantic similarity, cross-document pairs share a median run of&nbsp;<strong>11.5 identical consecutive words<\/strong>, versus 4.0 and 7.5 for sentences&nbsp;<em>within<\/em>&nbsp;each thesis &#8211; that is, among semantically matched sentences, the two documents phrase shared ideas alike more often than either thesis does when restating itself.<\/li>\n<\/ul>\n\n\n\n<p>Under the most conservative assumptions available, the probability of observing this pattern under independent composition is&nbsp;<strong>5.7 \u00d7 10\u207b\u2074<\/strong>; on the central estimate of the null rate,&nbsp;<strong>8 \u00d7 10\u207b\u00b9\u00b9<\/strong>. &#8220;<\/p>\n\n\n\n<p><em>[From Titorekno&#8217;s GitHub page]<\/em><\/p>\n<\/blockquote>\n\n\n\n<p>It is very unlikely this happened independently (specifically one in&nbsp;a&nbsp;<strong>one-in-100-billion chance<\/strong>&nbsp;of happening).<\/p>\n\n\n\n<p>The origin of the extreme textual overlap perhaps remains unclear &#8212; but both Liverpool John Moores University and Cambridge University have got to come up with better explanations than they have done so far.<\/p>\n\n\n\n<p>And for Cambridge University&#8217;s senior management, there are serious questions as to why this is playing out so damagingly in public, since they have known about it for a number of years.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>There is now a lot of press coverage of in The Times, the Telegraph and the Mail. From a mathematician&#8217;s or programmer&#8217;s perspective, we highlight an outstanding piece of work by Alexey Titorenko, available on GitHub here. It presents a fully reproducible, embedding-based analysis of the PhD theses by Arday [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"om_disable_all_campaigns":false,"_monsterinsights_skip_tracking":false,"_uf_show_specific_survey":0,"_uf_disable_surveys":false,"footnotes":""},"categories":[1],"tags":[],"class_list":["post-4036","post","type-post","status-publish","format-standard","hentry","category-blog"],"aioseo_notices":[],"_links":{"self":[{"href":"https:\/\/21percent.org\/index.php?rest_route=\/wp\/v2\/posts\/4036","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/21percent.org\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/21percent.org\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/21percent.org\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/21percent.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=4036"}],"version-history":[{"count":5,"href":"https:\/\/21percent.org\/index.php?rest_route=\/wp\/v2\/posts\/4036\/revisions"}],"predecessor-version":[{"id":4043,"href":"https:\/\/21percent.org\/index.php?rest_route=\/wp\/v2\/posts\/4036\/revisions\/4043"}],"wp:attachment":[{"href":"https:\/\/21percent.org\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=4036"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/21percent.org\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=4036"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/21percent.org\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=4036"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}