Acute rheumatic fever and rheumatic heart disease: where the research is missing

88 papers read, 3,545 claims extracted, every candidate gap attacked by a reader whose job was to destroy it.

21 August 2026, 20:26 · nothing here has been checked by a subject expert

What survived

Two things, and they are the same shape: the association is settled, and the study that would act on it has never been run. That shape matters, because it is the one kind of gap that cannot be manufactured by authors being modest about their own work.

survived

Nobody has tested whether echo screening changes what happens to a child

Every study is a yield or implementation report: how many cases found, how well nurses read the images, what the handheld device costs. No controlled trial has compared a screened population against an unscreened one on death, heart failure, surgery or recurrence.

Why it is credible: this is an absent study design, not absent attention. A trial exists or it does not. And the World Health Organization’s own refusal to recommend population screening rests on this exact absence.

Why nobody has done it: it needs years of follow-up and it means randomising children away from a test people already believe helps.

Beside it: no end-to-end costing of a real running programme either — confirmatory echo, referral, prophylaxis, and the cost of false positives.

survived

The diseased tissue has never been looked at closely enough

No single-cell or spatial profiling of explanted rheumatic valve tissue, with paediatric material scarcest of all. Three readers arrived here independently from different directions.

Why it matters: it is the bottleneck sitting behind several other questions. Every mechanism and drug-target question in this field now needs that resolution, and the candidate pathways that exist cannot be told apart without it.

How much to trust these. Both came through fast triage — two searches each — not the full adversarial attack that killed everything else on this page. Tonight’s record is that everything properly attacked died. Treat these as the best places to point a real attack next, not as findings.

What was tested, and what happened

AttackedCountKilledLeft standing
Contradictions between papers
same quantity, incompatible numbers
17170
Gap candidates, fully attacked
top-ranked, one skeptic each
129 refuted
3 weakened
0
Themes, fast triage
clustered from 871 stated limitations
8770 dead
12 weakened
5*

* two of those five are the same already-refuted claim, revived by a shallow search. 560 literature searches were run in total.

Why everything died

Not one contradiction died because researchers disagree. Every single one died on our own reading of the papers, and that is the most useful thing this exercise produced.

A quote cut before the word ‘but’

A sentence was quoted accurately and stopped one clause early, which reversed its meaning. It passed the verbatim check, because that check confirms a string exists in the paper — not that it still means what the author meant.

A survival curve compared against a headcount

51.9% and 7.9% looked like a six-fold disagreement between two national cohorts. One was a Kaplan-Meier estimate, the other a raw count over uneven follow-up. Like for like, the gap was 2.2x, and the two papers cite each other and explain it.

A denominator that was already filtered

A screening specificity of 19.7% would have sunk the case for task-shifted screening. The table it came from was headed ‘all screen-positive’, so the true population figure is about 95% — higher than the number it supposedly contradicted.

One row of a table read against another

Two reports of the same Cuban programme, 12.2→2.1 and 18.6→2.5. First attacks versus first-plus-recurrent, from the same table. The numbers add up exactly.

Two studies read as one

A review reporting a programme as both a 58% reduction and a non-significant 23% effect was correctly reporting two different evaluations, at 61 schools and at 244.

A comparison that was never made

A meta-analysis credited with finding no sex difference had run no sex comparison at all. Its two figures were each sex’s cases divided by the total sample.

Two of these patterns were found in published reviews, not just in our own pipeline: a ten-fold units error that the review contradicts in its own abstract, and a ‘significant’ result that was a writing group’s own uncorrected recalculation, labelled as such in its own table.

What the method is worth

The first approach was to mine every paper for what its authors admitted they could not answer, then cluster those. It produced 871 candidates from 88 papers and, when tested, nothing.

The reason is structural. “We did not study X” is a paper being modest about itself, written to head off a reviewer. Twelve papers each declining to explain X reads like a hole in the field and is twelve authors each saying their paper was about something else. Clustering makes it worse, because it attaches a paper count to the artefact, and a count looks like evidence.

Three separate times, the answer was inside the candidate’s own sources. One candidate’s four citations were the model it claimed nobody had built.

Kills are trustworthy. Survivals are not. Finding a paper that closes a question is positive evidence. Failing to find one in two searches is evidence of nothing — three triage batches independently revived a claim that a fifteen-search skeptic had already destroyed, using a paper cited inside two of the sources.

The detail

Every contradiction, and how it died17

The starting and ending acute rheumatic fever incidence figures reported for the same 10-year Pinar del Rio, Cuba primary/secondary prevention programme
Both numbers are verbatim from the same table in the primary source (Nordet 2008) -- Abrams quotes the first-attack row (12.2 to 2.1) and Zuhlke quotes the total-attack row, first plus recurrent (18.6 to 2.5); 12.2+6.4=18.6 and 2.1+0.4=2.5 exactly.
died by: dissolves

How much weight the PARF (peptide associated with rheumatic fever) M-protein/collagen-IV mimicry mechanism deserves as a driver of RHD autoimmunity
Both papers say the same thing: PARF-collagen mimicry is one candidate mechanism among several, not the sole one, and Zuhlke's own paper states this in the same paragraph it cites the 1/74 figure from.
died by: dissolves

Direction of circulating IL-2 in ARF/RHD — strongly elevated in acute disease vs. low/deficient in valve disease
IL-2 in ARF/RHD is an activity marker, not a fixed trait: it goes up in active/acute disease and falls back to normal in quiescent/chronic disease, a stage-dependent pattern the field already documented over 30 years before either shelf paper, and carapetis-2025 itself says its acute signature is expected not to match chronic RHD.
died by: already-known

Whether RHD prevalence in children/adolescents is roughly equal between sexes or substantially higher in females, in pooled meta-analyses of population-based echo screening.
Abdu 2024 never compared the sexes: its two 'sex-specific prevalences' are each sex's case count divided by the TOTAL sample size, no risk ratio or test of any kind appears anywhere in the paper, and a design that pools two marginal proportions across studies with I2=96% cannot detect a 1.4:1 sex effect even when one is present in every study.
died by: just-an-error

Population-level RHD prevalence in Ethiopia, Sudan and Uganda -- rigorous pooled meta-analysis vs a narrative review's cited country figures.
aliyu-2024's Ethiopia/Sudan/Tanzania numbers are not a rival estimate, they are the single highest per-study per-1000 prevalence figures from the same underlying literature relabeled as percentages -- a ~10x unit error the review makes repeatedly and even contradicts in its own abstract.
died by: just-an-error

Whether the landmark 1998-2001 South Auckland (NZ) school-based sore-throat clinic RCT (Lennon et al.) found a statistically significant reduction in first-attack rheumatic fever, or no significant reduction.
The trial's own authors, in the actual RCT report, call the result nonsignificant under both the old and revised Jones criteria (P=0.47 and P=0.27) -- the 'significant' RR 0.67 quoted by side_a is a different paper's own uncorrected recalculation, not the trial's finding, and it says so on its own table.
died by: already-resolved

The magnitude and statistical significance of New Zealand's national school-based sore-throat programme's effect on first-presentation ARF, across two primary evaluations covering overlapping years, cited within the same review paper.
The two figures are two different primary studies the review cites separately by name (Lennon 2017, 61 Auckland schools; Jack 2018, 244 schools nationally), and the larger study itself reconciles them: its own high-coverage high-risk region subgroup reproduces the big effect, while its diluted national average does not.
died by: dissolves

Whether task-shared (non-physician-led) programmes for RHD tertiary prevention/care in LMICs have any published evidence at all.
Abdullahi's zero-hit review searched only 'task shifting/task sharing/non-physician/cadre' vocabulary and had its search fixed via a 2017-registered PROSPERO protocol, while Rusingiza's paper -- published online seven-plus months after that protocol -- never once uses any of those search terms, calling itself 'decentralised, integrated' care instead, so the review's own search strategy could not have found it whether or not the timing lined up.
died by: dissolves

What fraction of a first episode of acute rheumatic fever (ARF) actually progresses to rheumatic heart disease (RHD), measured empirically over roughly a decade of follow-up in a comparable high-burden, Indigenous-majority, high-income-country population with an active national ARF/RHD register.
Oliver 2021 cites He 2016 by name, states the Australian register figure, names this exact gap in its Discussion and attributes it to echocardiography outreach, then says in its own Limitations that its hospitalisation-only figure 'markedly underestimates' progression -- so the two papers are in a citing relationship that already resolves the discrepancy, not in conflict.
died by: already-known

Whether 3-weekly benzathine penicillin G is a proven-superior regimen for preventing recurrent ARF compared with 4-weekly, or an underpowered non-significant trend from a single small trial.
Steer's sentence was quoted with the reversing clause amputated -- his very next words say guidelines recommend 4-weekly as standard with 3-weekly reserved for breakthrough recurrence, and the Cochrane review's own Authors' Conclusions say 3-weekly 'appeared to be more effective', so all three papers are on the same side and there is no disagreement to referee.
died by: extraction-error

Whether standard 4-weekly benzathine penicillin G actually maintains protective serum levels for most of the dosing interval (pharmacokinetic evidence) versus whether the standard regimen is highly effective at preventing recurrence in practice (population-level clinical outcome evidence).
The PK-vs-clinical-efficacy discordance for 4-weekly benzathine penicillin is not a tension this shelf uncovered -- Neely et al. 2014 modelled the same subtherapeutic pattern five years before hand-2019, hand-2019's own authors call it 'a major knowledge gap' in their abstract, and the field has been running trials on higher doses, shorter intervals and subcutaneous formulations (Bennett 2025 SCIP trial) specifically because of it, with a 2022 hollow-fibre study directly re-testing whether the 0.02 mg/L threshold is even the right number; on top of that, the entry's own headline number for the Ethiopian cohort is wrong by 6x.
died by: already-known

What counts as 'adherent'/'optimal' varies enough between papers, even within similar hospital-registry settings, to make the headline adherence percentages non-comparable -- a definition conflict, not a data conflict.
The field already handles this by holding the definition constant (bimerew-2024's meta-analysis only pools studies using the same >=80% cutoff) and still gets 33-fold spread (10.7% to 92.9%), which proves definition is not the dominant driver of incomparability that the entry claims it is -- and Edwards-2021's 32% never enters that pool in the first place because it used a different (undisclosed-as-nonstandard) definition, so no one in this literature was ever treating it as comparable to the Ethiopian numbers.
died by: already-resolved

Whether echo-detected RHD prevalence differs by sex in African population-based screening studies, pooled
abdu-2024's African 'near-parity' comes from eyeballing two independently-pooled, overlapping subgroup CIs and never running a formal within-study sex comparison; mutarelli-2025's significant female excess comes from computing the male:female risk ratio inside each study (mostly the SAME African primary studies) and pooling that ratio directly -- a properly-specified, more powerful test, not a different population.
died by: dissolves

Whether rural residence is a significant risk factor for RHD prevalence
Not a real disagreement: noubiap-2019's rural/urban term never even entered its multivariable model (univariable p=0.373, below the 0.20 inclusion bar, in an unadjusted-age global pool), while mutarelli-2025 is a fragile unadjusted RR from 10 pediatric studies (I2=88%, p=0.03) — different scale (global-pooled vs pairwise), different adjustment, different age band, and a third source (Yemen 2026, n=4135, one country) finds rural DOES survive adjustment locally, which is the classic aggregation pattern, not a contradiction.
died by: dissolves

Whether acute rheumatic fever (ARF) incidence itself — as opposed to RHD prevalence — differs by sex
Not a contradiction: a global-review generalization vs. a single, already-known-extreme population (NT Indigenous Australians, studied twice by the same overlapping author team including Carapetis himself), and independent literature explicitly reconciles the two via age-stratification and population heterogeneity.
died by: dissolves

Whether task-shifted handheld/portable echocardiography screening (non-experts, read against standard echo) has high specificity (~85%) or very low specificity (~20%) for detecting latent RHD in a school-age population.
Diniz's 19.7% is not a population specificity at all: Table 4 is explicitly headed 'All Screen-Positive (N = 110)', so the denominator contains only children who had already screened positive, and Diniz's own reported counts (110/1390 screen-positive, 3.2% confirmed prevalence) force a true population specificity of at least 91.8% and about 95% as computed, which sits above Providencia's pooled 0.85 rather than against it.
died by: dissolves

Whether pooled RHD prevalence differs significantly between the 2012 WHF and the WHO echocardiographic diagnostic criteria, across the same general body of population-based screening literature.
The two reviews test different corpora with different rigor (82 all-age studies + formal p<0.0001 subgroup test vs 58 pediatric-endemic studies + a qualitative eyeball call on an 8-study, I2=97% WHO subgroup), and a dedicated 2026 head-to-head empirical study already confirms criteria choice does shift RHD classification -- so this is a documented, expected field problem, not a live contradiction.
died by: dissolves

Every gap candidate, and how it was refuted12

Why RHD affects women roughly twice as often as men — REFUTED
A biological mechanism for the female excess in RHD has already been published in Circulation (Passos et al. 2022): an estrogen-receptor-alpha-linked, prothymosin-alpha-driven CD8+ T-cell autoimmune pathway against valve collagen, and it is even cited inside two of the candidate's own seven source papers (mutarelli-2025, lupieri-2025), so the gap sentence claiming this 'remains unexplored' is contradicted by the candidate's own reading list.

Cost-effectiveness and sustainability of echo screening programmes — REFUTED
The claim that nobody has modelled the cost-effectiveness of RHD echo screening is false: at least seven published cost-effectiveness/cost-utility models exist since 2013 across Australia, Brazil, India and Rwanda, one of them (Ubels/Nascimento 2020, Brazil PROVAR+) explicitly costing handheld devices with task-shifted, telemedicine-read screening, and Nepal has already run a cluster-randomized school-based echo screening trial.

The economic value case for a GAS vaccine — REFUTED
The candidate's own four citations are already a full health-economic model of a Strep A vaccine (trillion-dollar benefit ranges, income-stratified cost-effectiveness thresholds, DALYs and deaths averted by age of vaccination), so "nobody ever built the model" is false on its face.

Economic evaluation of prevention strategies other than vaccines — REFUTED
The model has been built at least five times since 2002, including a 55-country African investment case (2021) and an equity-stratified India analysis (2023) that does exactly what the candidate says nobody has done.

Social determinants of ARF and RHD outcomes are asserted but never measured — REFUTED
Social determinants of ARF/RHD are one of the most heavily measured relationships in this literature - household crowding, individual SES and post-surgical socioeconomic status all have quantified, adjusted effect sizes across multiple settings - the gap sentences are three different papers' single-study caveats, not a field-wide silence.

At what age a GAS vaccine should be given — WEAKENED
The stated obstacle ("nobody ever built the model") is false — the same author group built it at least three times, comparing infant vs age-5 GAS vaccination on cost, cost-effectiveness and DALYs/deaths averted; what actually survives is a narrower causal question those models explicitly admit they didn't solve, and it can't be closed by more modelling alone.

Managing and predicting risk in pregnant women who already have RHD — REFUTED
Both the screening decision and the risk-prediction decision for RHD in pregnancy are already served by a 2024 WHO-guideline systematic review, an RHD-specific validated risk score (DEVI, replicated twice), a 129-patient Australian outcomes study, and a live Australian screening trial — the source paper's own quote that no Australian outcomes study exists is simply wrong, contradicted by a paper published a year before it.

An affordable point-of-care test for GAS pharyngitis — WEAKENED
The exact combination (molecular/RADT test x LMIC unit costs x Markov CEA) hasn't been published, but every piece it would require already has been — an LMIC cost-effectiveness model for GAS diagnosis-and-treat strategies exists (Irlam 2013, South Africa), a full systematic review + economic model of POC GAS tests exists (Fraser 2020, UK HTA), and the candidate's own cited paper (Armitage 2025) already supplied LMIC accuracy data showing the tests currently perform too poorly to be worth costing out — and the real reason WHO tells endemic countries to skip testing is stated explicitly as poor test ACCESS, not missing analysis.

What actually drives adherence to secondary prophylaxis and what level is enough — REFUTED
The 'exact adherence level required' is already quantified with dose-response curves (de Dassel 2018), the candidate's own second source paper states the ≥80% threshold as established fact, and Uganda's own registry team already joined adherence data to RHD-progression outcomes and got a published, if awkward, answer.

School-based sampling versus true community prevalence — WEAKENED
Someone already ran exactly this test at meta-analytic scale in 2014 and found no significant difference between school-based and community-based RHD prevalence surveys, so the premise that this bias is unexamined and unquantified is wrong even though the specific test was underpowered.

Prevalence data simply do not exist for most countries — REFUTED
Country-level RHD burden data already exist for nearly every named 'data-free' region (SE Asia, Pacific, Latin America, North Africa, Central Asia, sub-Saharan Africa) as primary surveys, region-wide meta-analyses, a national multi-country mortality registry, and a Global Burden of Disease model that assigns every country including Malawi a numeric estimate -- and that GBD estimate for Malawi (323,422 cases, 2021) is sitting inside the candidate's own source paper.

GBD model estimates cannot be checked against real data — REFUTED
Joining registry/screening data to GBD RHD estimates to test their validity has already been done and published, including a paper already sitting on our own shelf next to this one.

Themes that were weakened but not killed12

Each overstated an absence. What is written here is the narrower thing left after the overstatement was removed.

Whether decentralised, non-specialist-run RHD service can carry the care and survive the pilot grant ending
Task-shifting echo screening/diagnosis to nurses is already well established as feasible and accurate across multiple LMIC pilots — that half is dead — but a controlled evaluation of decentralised chronic secondary-prophylaxis delivery on adherence/retention, costed, with an explicit post-donor-funding sustainability test, turns up nowhere in search or on our own shelf.

Which biological target a drug developer should aim at for rheumatic valve fibrosis
Candidate-pathway mechanism papers already exist (cytokine imbalance, TGFβ/SMAD3, S1PR1/STAT3 signalling) so 'nobody has a hypothesis' is false, but a human single-cell-resolution valve tissue atlas across disease stages feeding an organoid/ex-vivo drug-testing platform does not appear in the literature at all.

Whether to spend on housing/crowding/skin health instead of clinics and drugs
Crowding-ARF association is one of the most over-replicated findings in the field (repeated case-control studies, all correlational) so that part is dead on arrival, but no housing/environmental-intervention trial with a measured Strep-A-transmission or ARF-incidence endpoint turned up anywhere — the causal test genuinely does not exist. Framed as spending allocation, this is policy dressed as research; framed as 'does fixing housing change transmission,' it is a real unanswered experiment.

Whether to screen pregnant women for RHD and how to manage those found
Pregnancy-with-RHD management and maternal/fetal outcomes is a large, well-covered literature (registries, case series, recommendations), but a controlled trial of routine antenatal echocardiographic screening versus standard care, with screening-strategy and gestational-timing endpoints, was not found — the screening-policy question specifically is open even though the downstream-management question is not.

A drug that halts or reverses RHD once it has started
The premise is overstated: corticosteroids, aspirin and IVIG for acute rheumatic carditis have already been through multiple RCTs (synthesised in a Cochrane-type review with a null result), so it is false that no anti-inflammatory has ever entered a trial; what is actually missing is newer immunomodulator targets (TNF, IL-6, Treg, TGF-beta) being tried specifically in RHD, which is a drug-development pipeline lag, not an unstudied question.

RHD in adults and in people who are not in school is essentially unmeasured
Adult data is not absent, it runs through a different pipe than school screening: hospital-based multi-country registries (REMEDY-type) and adult community studies already exist and an NHLBI workshop already set research priorities across the age range; what's genuinely thinner is nationally-representative all-age community echo screening specifically, not adult RHD data as a category.

What drives progressive valve fibrosis long after the acute episode
A 2025 synthesis already maps the causal chain from immune trigger through inflammatory to mechanical progression, so 'mechanism unresolved' overstates it; what is genuinely missing is the fine-grained part — single-cell/spatial resolution explaining why stenosis versus regurgitation diverges by valve region.

Whether penicillin prophylaxis changes the course of latent or subclinical RHD
The GOAL trial already answers this for definite latent RHD, and a 2023 commentary ('a naive realism?') already argues the case for borderline disease in print — so the debate is live and worked, not silent; a dedicated RCT restricted to borderline latent RHD specifically was not found.

ARF has no diagnostic gold standard and no field-usable algorithm
ARF genuinely has no biological gold standard by definition (this is structural, not a gap a study fills), but simplified field algorithms already exist and are deployed in endemic settings, e.g. Sudan's simplified guidelines.

Whether a GAS vaccine can be made without triggering autoimmunity
Genuinely still unresolved and named as a standing research priority by an NHLBI workshop, but not an untouched gap — candidate vaccines are already in monitored human trials for exactly this risk.

Immune correlates of protection for a GAS vaccine
Real and still open, but not untouched — there is already a paper titled exactly this, and a controlled-human-infection-model program exists to chase it.

Standardised immunoassays and safety surveillance for GAS vaccine trials
Assay-standardization work is already underway across several labs (opsonophagocytic killing assays), but a true cross-lab ring trial with an international reference serum is not yet shown to exist.

Themes killed outright70

already-done — 55

not-a-question — 7

dissolves — 4

not-research (funding/procurement/policy) — 2

not-research (funding/procurement) — 1

not-research — 1

The raw material871 stated limitations

871 limitations mined from 88 papers, each with a quote checked against the source PDF, then clustered three separate ways by three independent readers. This is the material everything above was built from. It is kept because the measurement — that it yielded nothing — is worth more than the material itself.

55 apparent contradictions were also checked and found innocent before the 21 above were reported.

Built from the evidence shelf at 21 August 2026, 20:26. Every claim traces to a PDF on disk. No part of this has been reviewed by a clinician or a subject expert, and the two surviving items have not yet been adversarially attacked.