How the Dawson commentaries project integrates public posts

This page explains how the Michelle Dawson posts analysis project integrates public posts from @autismcrisis, reconstructs cited links where possible, and identifies scientific references or study registration records mentioned through those links.

Scope of the Project

The project focuses on Michelle Dawson’s public posts (@autismcrisis) relating to autism science, scientific publications, and related debates. It covers two public sources: Twitter/X as a historical archive, from 10 October 2009 to 12 October 2023, and Bluesky for ongoing monitoring from 13 October 2023 onwards.

These two sources are handled separately because their technical and legal conditions differ. Bluesky provides an open public API that supports regular synchronization. Twitter/X is a closed platform with a paid API, so AutiHub treats Twitter/X content as a historical archive rather than as a continuously synchronized source.

Bluesky

For Bluesky, AutiHub synchronizes public posts from @autismcrisis.bsky.social and verifies whether Michelle Dawson has added direct replies to her own posts. These replies are integrated with the original post when they function as continuations or additional notes.

During synchronization, AutiHub also verifies recent posts already stored locally. If a post has been deleted on Bluesky, it is removed from AutiHub during a subsequent synchronization so that empty embeds do not remain visible. For new Bluesky posts, cited links are inspected so that scientific resources can be associated with the post whenever possible.

Twitter/X

For Twitter/X, the project follows an archival approach. Public posts from @autismcrisis were integrated through the X API for the period before the move to Bluesky. The archive also includes direct replies posted by the same account when they extend or continue one of the integrated posts.

The raw text of Twitter/X posts is used internally for search and analysis. Public display relies on links or official embeds whenever possible, rather than presenting the archive as a downloadable republication of the original posts.

Reconstructing Cited Links

Many older Twitter/X posts contain shortened links, including Twitter/X internal links such as t.co and public URL-shortening services such as bit.ly or j.mp. AutiHub preserves the original URL detected in the post and then attempts to reconstruct the destination URL where this is still possible.

This reconstruction process is conservative. Some older URL-shortening services are no longer reliable, some publisher pages now redirect to generic cookie or error pages, and some websites block automated access. In such cases, AutiHub may retain a fallback destination, mark the link as inactive, or leave it unresolved rather than invent a destination.

Finding DOI and Registry Identifiers

When a reconstructed destination appears to correspond to a scientific publication or study record, AutiHub attempts to extract a stable identifier. The most useful case is a DOI. The project also detects identifiers from study registries and research databases, including ClinicalTrials.gov, CRD/PROSPERO, and NIH RePORTER.

Depending on the source, several identification strategies may be used: DOI patterns visible directly in URLs, publisher-specific URL structures, metadata embedded in publisher pages, registry identifiers contained in URLs, and manual correction workflows for more complex cases, such as pages where a DOI is visible to a human reader but is not reliably exposed through a simple automated request.

Structured Metadata

Once a DOI or registry identifier has been identified, AutiHub retrieves structured information from the appropriate metadata sources. For DOI-based publications, this usually means Crossref, with fallback sources when necessary. For registry records, AutiHub retrieves available information such as titles, dates, summaries, institutions, and identifiers from the relevant registry whenever possible.

This structured layer makes it possible to list the scientific references and study registration records cited in Michelle Dawson’s posts, sort them by publication date or post date, and link each resource back to the post in which it was cited.

Search and Public Display

The text of public posts is stored in AutiHub so that the project can provide search, filtering, and structured analysis. This internal storage is used to make the project searchable; it is not intended to create a separate downloadable republication of the raw archive.

For public reading, AutiHub favors official embeds and links to the original platforms. This keeps the displayed content connected to Bluesky or Twitter/X whenever possible, while allowing AutiHub to provide research-oriented search features and contextual information.

Limits

The structured reference list is not a list of every URL cited by Michelle Dawson. It only includes studies and study registration records for which the destination of a cited link contains sufficient structured information, such as a DOI or a registry identifier.

Older links may be inactive, redirected to generic pages, blocked by anti-bot systems, or may no longer point to their original destination. The goal is to make the source material more searchable and easier to understand while maintaining a clear distinction between original public posts, imported metadata, reconstructed links, and structured scientific references.