Methodology
The analyzed corpus includes 533 citations of scientific studies, collected over the period from March 12, 2025 to June 13, 2026. These citations come from content published under the name “Thérapie Autisme”. It includes posts from the Instagram account “therapietsa” up to June 9, 2026, articles from the “Thérapie Autisme” blog up to June 13, 2026, bibliographic documents associated with tools marketed by “Thérapie Autisme”, and the manuscript titled “Variabilité neuro-architecturale dans l’autisme adulte : vers une cartographie morpho-fonctionnelle”, published on the Skool.com platform “Thérapie TSA - Autisme Adultes”.
Process overview
- 1. Citation. A cited reference must contain usable author information, a year and a title. Journal, volume, pages and DOI are useful but optional.
- 2. Extraction. Visible bibliographic elements are separated without reformatting the citation into a standard academic style.
- 3. Method 1. Crossref, OpenAlex and PubMed are searched from author, year and title.
- 4. Method 2. Google Scholar is queried through SerpApi when this source may provide a better match.
- 5. Final score. The final candidate is selected, each bibliographic element is compared, an overall score is calculated, and the title Jaccard index is recorded.
- 6. Records. Each record presents the citation, the identified reference, verification links and calculation details.
Corpus construction
The analysis includes citations containing at least one author, a publication year and a sufficiently explicit title. Citations without a title, and citations that appear to be approximate French paraphrases of an English scientific title rather than identifiable bibliographic titles, are excluded.
- Author(s). Must be present in a usable form: single author, multiple authors or “et al.”.
- Year. Must be present to situate the reference and limit homonymy.
- Title. Must be present as an identifiable bibliographic title, rather than only a reformulation or summary.
- Journal, DOI, volume, pages. Useful for assessing citation fidelity, but not required for inclusion in the corpus.
Candidate reference search
- M1: Crossref, OpenAlex, PubMed. Searches are based on author, year and title. The cited DOI is not used to drive the search, because an incorrect DOI could lead directly to the wrong reference. This method may fail to retrieve some titles or may return incomplete metadata.
- M2: Google Scholar via SerpApi. Used as a supplementary source when the citation is incomplete, imprecise or phrased differently. Its metadata are less structured, and information such as journal or volume may be missing.
- Final: M1 or M2. This is the reference used in the charts and citation records. Ambiguous cases still require human review.
Final candidate selection
When multiple candidates are available, the selected reference is the one whose title shows the closest correspondence with the cited title according to the Jaccard index. In the event of a tie, the general score is used to distinguish between candidates. This assessment considers title, authors, year, journal and DOI.
- M1 has the best title correspondence. The M1 reference is retained.
- M2 has the best title correspondence. The Google Scholar reference is retained.
- M1 and M2 are tied on title correspondence. The general score is used to distinguish between candidates.
- M1 and M2 have the same DOI. Journal, volume and page information may be taken from M1 because those metadata are more structured.
General score
The general score is a weighted average based only on the elements actually available in the citation. Missing elements do not reduce the score: their weight is removed from the denominator.
- Title: 20%. Compared using significant words and controlled tolerances.
- Authors: 20%. Positional comparison of cited authors in the order provided.
- Year: 20%. Exact match or near match, depending on the case.
- Journal: 20%. Journal name, including volume or issue when cited.
- DOI: 20%. Exact match if the DOI is cited.
Missing DOI example. If title, authors, year and journal are assessable but DOI is absent, the denominator is reduced to 80%. A citation scoring 80 for title and 100 for the other three available elements therefore receives: (80×0.20 + 100×0.20 + 100×0.20 + 100×0.20) / 0.80 = 95/100.
Title fidelity: Jaccard index
The Jaccard index compares the significant words in the cited title with those in the identified title. It is calculated as the number of shared words divided by the total number of significant words appearing in either title.
- Normalization. Case, accents, punctuation, compound words and singular/plural variants are normalized.
- Acronyms. ASD or ASC may correspond to Autism Spectrum Disorder(s) or Condition(s).
- Language variants. Variants such as behavior and behaviour are treated as equivalent.
- Overly generic words. Terms such as autism, spectrum, disorder(s), autistic, ASD, ASC and condition(s) are excluded from the calculation.
- Incomplete titles. A title segment appearing before or after a colon may be tolerated only if all cited words correspond to the remaining segment.
Special rules and edge cases
- Cited but incorrect DOI. The DOI does not drive the search. It is compared afterwards and may penalize the DOI element.
- Abbreviated journal title. Standard acronyms or abbreviations may be accepted, for example JADD for Journal of Autism and Developmental Disorders.
- Volume only. A citation such as 41 may be considered compatible with 41(9) if the issue number was not cited.
- Different volume. A citation such as 41(2) compared with 40(3) is considered divergent.
- Severe false positive. If no authors are shared and the title correspondence is below 50%, the reference is considered invalid and the final score is set to 0.
- Missing element. Missing elements are not penalized; their weight is removed from the calculation.
- Approximate citation. The procedure allows certain variations to determine whether a real and coherent reference exists, while still reporting identified divergences.
Verification and corrections
- Verification links. Records include the citation source, the selected bibliographic source and a Google Scholar search generated from the citation.
- Reading records. Detailed comparisons show authors, year, title, journal, DOI and Jaccard index to support manual verification.
- Corrections. Possible errors or useful corrections can be reported through the AutiHub contact page.