Nila Kartika Wati is a reporter for Tech Ledgers covering Blockchain & Cryptocurrency. She/He is based in Indonesia.
23 August 2026 • 8 min read
Nearly two-thirds of recently published religious and belief-focused titles sampled on Amazon were likely authored by artificial intelligence, according to a report released Wednesday by AI detection firm Originality.ai. The study analyzed 2,034 recently published books across 14 belief and religious categories, flagging 1,272 titles—or 63%—as likely written by AI.
Executive Overview
The findings highlight a significant prevalence of synthetic text within specialized, self-help, and spiritual publishing categories on Amazon. The proportion of flagged titles varied markedly depending on the specific subject matter. Books categorized under Wicca, Witchcraft & Paganism exhibited the highest rate of AI content, with 78% classified as likely AI-generated. Hinduism followed closely at 76%, while Taoism registered at 74%. Other categories—including Mormonism (42%), Atheism (40%), and Satanism (22%)—demonstrated comparatively lower rates, though researchers noted that sample sizes varied substantially among categories.
Michael Fraiman, the report’s author, emphasized that Originality.ai’s evaluation framework determines statistical likelihood rather than absolute definitive proof. The process involved evaluating book descriptions, author biographies, and available text samples using an algorithmic model where scores of 50 or higher were classified as "Likely AI." Fraiman noted that manual spot-checks conducted by the research team reinforced the automated classifications.
The report emerges alongside broader debate concerning the reliability of AI detection technologies and platform governance on digital storefronts. While Amazon maintains an author self-reporting policy regarding the use of AI tools, it does not mandate verified disclosures or publicize compliance metrics. Simultaneously, the broader AI ecosystem continues to face scrutiny over the physical acquisition, scanning, and destruction of printed books used by developers to compile large-scale training datasets for synthetic text models.
Detailed Chronology
The timeline of key developments surrounding AI content generation, detection reliability, and literary data collection highlights the recent evolution of these issues:
August: An investigation tracked a shipment of rare books directly to an Amazon processing facility situated in Las Vegas. Investigative findings revealed that facility workers systematically cut off the physical bindings of rare books, scanned their pages to convert the text into training datasets for AI models, and subsequently discarded the physical remnants. Amazon confirmed that it purchases physical books through standard commercial channels to support this processing operation, reflecting broader industry practices where developers purchase millions of physical titles, strip them apart, scan them into training corpora, and destroy the originals.
October 2024: A controlled benchmark test evaluated four leading AI detection software platforms using the historical text of the U.S. Declaration of Independence. The experiment demonstrated stark inconsistencies among detection algorithms: ZeroGPT evaluated the foundational historical document as 97.93% AI-generated, whereas QuillBot identified it as 100% human-written. Meanwhile, GPTZero assigned the text an 89% probability of human authorship.
March: A legal filing presented to Colombia’s Supreme Court was formally rejected after automated AI detection tools flagged the legal document as computer-generated text. In response, an attorney submitted the Supreme Court’s own official written rejection ruling back through the identical detection software. The tool flagged the Colombian Supreme Court’s judicial ruling as 93% AI-generated.
Wednesday: Originality.ai officially published its research report examining newly published religious and belief titles on Amazon. The study established that out of 2,034 sampled books across 14 sub-genres, 1,272 titles (63%) met or exceeded the threshold for classification as "Likely AI."
Supporting Context & Metrics
Research Methodology and Scoring Framework
Originality.ai’s investigation evaluated a sample of 2,034 newly published titles available on Amazon. To determine whether a work was generated by artificial intelligence, researchers analyzed three distinct textual elements for each listing:
Book Descriptions: The marketing prose and narrative summaries listed on the product detail page.
Author Biographies: The background information and promotional profiles provided for listed authors.
Sample Text: The preliminary excerpt text accessible to prospective buyers.
Textual samples scoring 50 or higher under Originality.ai’s detection model were flagged as "Likely AI." Report author Michael Fraiman clarified the underlying scope of these metrics, noting that the software evaluates statistical probability rather than establishing absolute certainty.
"We don’t make claims to certainty about whether books are definitively AI-written," Fraiman stated. "Our model determines the likelihood that something was AI-written to a degree of certainty."
To validate the automated scoring system, the research team conducted secondary spot-checks on flagged books. Fraiman expressed confidence in the team’s capacity to verify the results manually, stating:
"By this point, our research team can tell the difference between likely AI-written content and human-written content."
Category Breakdown and Metric Distribution
The report evaluated books across 14 discrete belief, faith, and philosophical classifications. Across the total sample of 2,034 books, 1,272 (63%) met the threshold for AI classification. The findings demonstrated clear variance across categories:
Note: Originality.ai researchers explicitly cautioned that total sample sizes varied significantly across the 14 individual categories.
Subject Matter Analysis: Witchcraft and Alternative Healing
The highest concentration of AI-generated content was observed in Amazon’s Wicca, Witchcraft & Paganism category, where nearly four out of five sampled books (78%) were flagged as likely AI-written.
An examination of the flagged titles within this category revealed a consistent focus on specific topics:
Application and properties of healing crystals.
Formulation and usage of herbal remedies.
Routines designed for cleansing "negative energies."
Fraiman highlighted the intersection between the target audience for these titles and the unregulated nature of non-traditional wellness literature, providing insight into why this category may be particularly vulnerable to automated content generation:
“One can picture the ideal customer as someone looking for solutions to heal themselves, improve their mental health, or ‘detox’ from commonplace chemicals and drugs,” Fraiman told Decrypt. “These Wiccan books can address these concerns without being held to any scientific standard.”
Technical Challenges and Detection Flaws
While Originality.ai stood behind its methodology, the broader context of AI detection technologies remains marked by known technical limitations, including false positives and contradictory outputs.
The structural reliability of AI detection algorithms faces ongoing scrutiny, as evidenced by documented edge cases:
Historical Text Variance (October 2024): Analysis of the U.S. Declaration of Independence yielded opposing conclusions across platforms, spanning from ZeroGPT’s 97.93% synthetic determination to QuillBot’s 100% human-written assessment, with GPTZero rating it 89% human-written.
Judicial Feedback Loop (March): The Colombian Supreme Court’s rejection of an allegedly AI-generated legal filing was countered when the court’s own official document was evaluated by the same detector and assigned a 93% AI score.
Official Statements & Platform Policies
Originality.ai Findings on Platform Disclosures
Michael Fraiman detailed the operational dynamics surrounding self-publishing platforms, noting that Amazon relies primarily on voluntary author disclosures regarding artificial intelligence utilization.
"Amazon is aware of how common AI-written content is on their platform, and uses the honour system for authors to self-identify their work as AI-written," Fraiman said.
Fraiman confirmed that Originality.ai did not share its specific research dataset or report findings directly with Amazon prior to publication. He added that there is a total lack of public visibility regarding author compliance with platform guidelines:
"We have no data on how many authors actually disclose that."
Amazon Status and Physical Data Acquisition
Amazon did not immediately respond to a request for comment regarding the Originality.ai study findings.
However, regarding the broader practices involved in sourcing text for AI models, Amazon previously confirmed its participation in physical book acquisition and digitization efforts. Following the August investigation into its Las Vegas facility—where rare physical books had their bindings removed and pages scanned prior to disposal—Amazon confirmed that it purchases physical books through standard commercial channels to supply datasets for AI training operations.
Future Outlook
The findings presented in the Originality.ai report illustrate an evolving landscape in commercial digital publishing, where automated text production directly intersects with online retail marketplaces.
Self-Regulation and the Honor System Model
The reliance on an honor system for author disclosures presents ongoing challenges for online storefronts. Without mandatory, verified technical checks, retail platforms face difficulty monitoring the influx of synthetic manuscripts. Because authors are trusted to self-report AI usage voluntarily, marketplace operators possess minimal verified data regarding the actual proportion of human-authored versus machine-generated books sold across diverse categories.
Detection Inconsistencies and Enforcement Dilemmas
The implementation of automated detection enforcement remains complicated by tool variance and high-profile false positives. As demonstrated by conflicting assessments of historical documents and court rulings, digital storefronts that consider automated scanning risk mistakenly penalizing genuine human authors. Consequently, platforms remain in a precarious position between unmonitored self-reporting and unreliable automated enforcement.
Data Pipelines and Synthetic Content Loops
The relationship between physical print media and synthetic digital publishing represents a dual dynamic in contemporary AI development:
Ingestion: Developers acquire, unbind, scan, and destroy physical books—including rare titles—to assemble vast training datasets.
Output: Generative models process these datasets to produce rapid, low-cost digital manuscripts that are listed back onto retail platforms like Amazon.
As synthetic works proliferate across specialized categories—particularly those unconstrained by rigorous scientific standards, such as spiritual self-help and alternative wellness—consumers and platforms face an increasingly complex literary marketplace where machine-generated content accounts for a major share of new releases.