Methods and definitions
How We Produce the Editorial Insights
The reports are generated from the same repository files used to publish this website. This page defines what is counted, what is excluded, and what the results cannot establish.
Executive Summary
- The unit of analysis is a published web page. Story files and blog files are counted separately.
- Metrics are regenerated from source files. Counts, shares, word lengths, and CSV rows are not manually maintained.
- The reports measure editorial supply and source-link visibility. They do not measure medical prevalence, truth, outcomes, audience beliefs, or book contents.
Story corpus
The story index reads every Markdown file in src/data/stories. Category, specialty, country, publication date, provenance, and tags come from each file's front matter. Word count is calculated from the Markdown body after common formatting characters and HTML tags are removed. Shares equal group count divided by total story count, rounded to one decimal place.
Provenance definition
Editorial composite means a web narrative created for thematic education and reflection. It is not presented as a verified interview, patient record, clinical case report, or story from the published book. If verified submissions or book-derived cases are published later, they must use a different provenance label backed by permission and documentation.
Blog source audit
The audit reads every Markdown file in src/data/blog and detects linked URLs in standard Markdown link syntax. Amazon and amzn.to URLs are excluded as commercial links. A post is considered "with sources" when at least one other HTTP or HTTPS link is present. DOI counts are links containing doi.org/.
Claim-to-source map
The map creates one row for every qualifying contextual link and records its article, nearest preceding H2 section, visible Markdown passage, citation label, URL, domain, source class, caution-language signal, and medical-review status. Lines shorter than 35 plain-text characters are excluded so a standalone source-list entry is not mislabeled as claim context. Source class is assigned from the domain. The caution signal is a case-insensitive text match for a disclosed list of terms including association, cannot, limited, may, small, and uncertain.
Known limitations
- Link presence does not establish source quality or claim support.
- The map detects cited passages, not every factual claim in an article.
- Source class and caution language are reproducible heuristics, not expert judgments.
- Plain-text citations without a link are not detected.
- Word count is a consistent website measure, not a publishing-industry count.
- Editorial classifications may change as the library is reviewed.
- The dataset is a snapshot that updates when site content changes.
Reproducibility
The application calculates all reports through one shared analysis module. An automated quality check independently recounts the underlying files and confirms that report routes, CSV routes, disclosures, and structured data remain present.
Open the story index, open the evidence report, or inspect the claim-source map.
