Where the research is missing

Acute rheumatic fever & rheumatic heart disease · built 2026-08-21 20:02

871open questions
88papers
3ways of grouping
122doable & live
Everything here is a candidate, not a conclusion. Each line is something a paper’s own authors wrote down as unresolved, and the quoted words are checked against the PDF by machine. What has not happened yet: nobody has checked whether another paper already answered it, and no person has read these. Treat a single-paper item as one team’s opinion.

Where two papers cannot both be right

Ten readers searched this shelf for claims that contradict each other: the same quantity, comparable populations, incompatible numbers. Unlike everything further down, this cannot be an artefact of authors being modest about their own studies.

21real conflicts
55looked like conflict, weren’t
72%dissolved on inspection
133searches run
Population-level RHD prevalence in Ethiopia, Sudan and Uganda -- rigorous pooled meta-analysis vs a narrative review's cited country figures.high
One sideCountry-level pooled RHD prevalence from systematic review/meta-analysis of population-based echo studies
Ethiopia 25.83/1000 = 2.58% (95% CI 9.56-42.1/1000, 5 studies); Sudan 18.23/1000 = 1.82% (95% CI 7.83-28.63/1000, 4 studies) · Sub-groups of the 22 African population-based studies using WHF echocardiographic criteria, general/school-age population, per-country pooled estimates
abdu-2024-prevalence-pattern-rheumatic
The other sideRHD prevalence reported by country from selected school-based studies in East Africa
Ethiopia 56.74%, Uganda 56.7%, Tanzania 33.7%, Sudan 61.5% · "Selected school-based studies in East Africa" (narrative review, sourcing/methodology of the underlying primary studies not detailed in the claim)
aliyu-2024-rheumatic-heart-disease
Why they cannot both be rightThese are the same quantity (population/school-based RHD prevalence) in the same countries (Ethiopia, Sudan), but aliyu-2024's figures are 20 to 34 times higher than abdu-2024's rigorously pooled estimates from multiple studies per country with confidence intervals. A true population prevalence of 56-61% would mean more than half of the general population has echocardiographically detectable RHD, which is epidemiologically implausible for any population and is not supported anywhere else in this 250-paper shelf -- the highest other country-level pooled figure found anywhere is Ethiopia's own 25.83/1000 (2.58%). aliyu-2024 is also internally inconsistent: two paragraphs earlier in the same paper it states 'RHD prevalence in Africa ranges from 2.9 to 30.4 per 1,000 population' (i.e. up to 3.04%), which directly contradicts its own later claim of 33.7-61.5% for four named countries within that same continent.
The innocent explanation, and whether it was checkedLikely explanation, partially checked: aliyu-2024's cited country percentages are almost certainly a units or transcription error in the source material it is citing (e.g. a per-1000 or per-100 figure from a single small/high-risk clinical sample misreported as a general population percentage, or conflation with a different quantity such as clinical-suspicion/screen-positive rate before echo confirmation). The claims file only carries aliyu-2024's own text ('The prevalence of RHD in Ethiopia, Uganda, Tanzania, and Sudan was reported to be 56.74%...') without naming the primary sources behind those four country numbers, so the exact origin of the error could not be traced further within this shelf. Regardless of cause, the number as published in aliyu-2024 is incompatible with abdu-2024's meta-analysis and with aliyu-2024's own stated continental range.
What depends on which is rightCountry-level prevalence figures like these get lifted into policy documents and funding proposals as 'the' national burden. Citing aliyu-2024's Ethiopia/Sudan figures instead of abdu-2024's would overstate national RHD burden by roughly 20-30x, which would badly distort screening-program cost-effectiveness models and country prioritisation within regional RHD control planning.
domain: prevalence and burden
The magnitude and statistical significance of New Zealand's national school-based sore-throat programme's effect on first-presentation ARF, across two primary evaluations covering overlapping years, cited within the same review paper.high
One sideA school-based nurse/community-health-worker sore-throat clinic in New Zealand was associated with a large reduction in first-presentation ARF among school children over ~2 years.
58% reduction (88/100,000 [95% CI 79-111] to 37/100,000 [95% CI 15-83]) · ~25,000 school children, 61 schools · New Zealand, pre/post design, Jul 2012-May 2014
shimanda-2024-preventive-interventions-reduce (citing a pre/post evaluation of 61 schools, Jul 2012-May 2014)
The other sideNew Zealand's national school-based sore-throat services were associated with a decline in national ARF incidence, but the overall intervention effect estimate crossed the null.
23% (95% CI 6%-44%), RR 0.77 (95% CI 0.56-1.06) -- CI includes 1.0, not statistically significant · ~53,376 school children, 244 schools · New Zealand, retrospective cohort comparing baseline 2009-2011 to intervention Jan 2012-Dec 2016
shimanda-2024-preventive-interventions-reduce (citing a retrospective cohort, Jan 2012-Dec 2016)
Why they cannot both be rightBoth figures purport to measure the same intervention (New Zealand's Rheumatic Fever Prevention Programme school-based sore-throat clinics) over almost entirely overlapping years (2012-2014 is a subset of 2012-2016), yet one reports a large, precisely-bounded reduction (58%, CI 15-83 excludes null when read as a rate ratio well below 1) and the other reports a smaller reduction whose confidence interval crosses the null (RR 0.77, CI 0.56-1.06 -- not significant). A programme cannot simultaneously more-than-halve ARF incidence with a tight, clearly significant effect and also show a statistically inconclusive ~23% effect over a longer window covering the same period.
The innocent explanation, and whether it was checkedPartially. The two studies differ in scale (61 vs 244 schools -- the larger study is closer to the true national footprint) and design (simple pre/post two-time-point comparison vs a retrospective cohort with a 2009-2011 baseline and interrupted-time-series-style comparison across 5 years). Pre/post two-point comparisons are known to overstate effects relative to longer time-series designs because they don't control for secular trend or regression to the mean, which would explain why the smaller, cruder study shows a bigger and 'cleaner' effect than the larger, more rigorous one. This is a real methodological explanation, but it is not fully dissolving: shimanda-2024 cites both as evidence in the same review without flagging that they contradict each other on the programme's core question -- whether it produced a real effect -- and a reader citing 'NZ's programme cut ARF by 58%' vs 'NZ's programme effect was not statistically significant' would be equally citing this same review.
What depends on which is rightThis is the number policymakers point to when arguing that national school-based sore-throat screening is worth funding elsewhere. Whether the true effect is a clean >50% reduction or a smaller, statistically inconclusive one changes the strength of the case for replicating NZ's model in other high-burden countries, and the review that would normally reconcile this instead reports both without resolving them.
domain: control programmes, registers and economics
Whether 3-weekly benzathine penicillin G is a proven-superior regimen for preventing recurrent ARF compared with 4-weekly, or an underpowered non-significant trend from a single small trial.high
One side3-weekly intramuscular penicillin did NOT significantly reduce rheumatic fever recurrence compared with 4-weekly injections; the difference did not reach statistical significance.
9 recurrences (3-weekly) vs 16 (4-weekly), RR 0.57, 95% CI 0.26-1.23 · Single trial (Lue 1996), 249 patients aged 3-19 in Taiwan, 12-year follow-up; allocation initially by odd/even hospital numbers (inadequate concealment) · Cochrane-style systematic review reporting the trial's own relative risk and confidence interval, explicitly flagging non-significance
manyemba-2002-penicillin-secondary-prevention
The other sideBenzathine penicillin G every 3 weeks prevents recurrent ARF better than every 4 weeks in highly endemic areas / 3-weekly rather than 4-weekly BPG reduced the number of recurrences of ARF in studies in Taiwan.
Stated as a flat directional finding, no effect size or confidence interval given, citing the same Lue et al Pediatrics 1996 Taiwan data · Same underlying Taiwanese ARF patients (Lue 1996 trial), described narratively rather than with the trial's own statistics · Narrative review citation of the same primary trial, omitting the confidence interval and significance test
seckeler-2011-worldwide-epidemiology-acute AND steer-2009-acute-rheumatic-fever
Why they cannot both be rightAll three papers are citing the exact same primary data (the Lue 1996 Taiwan trial, 9 vs 16 recurrences). The Cochrane systematic review (manyemba-2002) computed the actual RR and CI from that data and the CI crosses 1 (0.26-1.23) -- the textbook definition of a non-significant result. Seckeler 2011 and Steer 2009 cite the identical trial but state the direction of the point estimate as an established, unqualified fact ('reduced', 'better prevention'). A single underpowered trial with a CI spanning unity cannot simultaneously be 'not significant' and 'proven better' -- one characterization is accurate about the statistical evidence and the other overstates it.
The innocent explanation, and whether it was checkedChecked: it does not fully dissolve, but it softens. Manyemba-2002 also reports a second, actually significant endpoint from the same trial (3-weekly reduced streptococcal throat infections by 33%, RR 0.67, 95% CI 0.48-0.92) and notes compliance was comparable between arms. Seckeler and Steer may be summarizing the trial's overall favorable direction across both endpoints rather than misreading the ARF-recurrence result specifically, and it is possible their source citation (refs 86/87 in Steer) includes a second Taiwan paper by the same group not captured in this claims file. But as written, both narrative papers assert '3-weekly reduces ARF recurrence' as if it were the significant finding, which the primary data (as reported by the systematic review) does not support for that specific endpoint.
What depends on which is rightThis is the entire evidence base for whether high-burden programmes should switch patients from 4-weekly to 3-weekly BPG injections -- a real tradeoff, since 3-weekly means 33% more painful injections per year and, per REMEDY registry data, is associated with lower patient-reported adherence (76.0% vs 82.8% for 4-weekly). A clinician or guideline writer reading Seckeler or Steer would conclude the switch is evidence-backed; reading the Cochrane review shows the ARF-recurrence claim rests on a single 249-patient trial with inadequate allocation concealment and a confidence interval that includes no effect. Current WHO/AHA guidance (gerber-2009, leal-2019 in this shelf) reflects the cautious reading -- 3-weekly is reserved for high-incidence populations or breakthrough failures on 4-weekly, not recommended as universally superior.
domain: penicillin prophylaxis
Whether RHD prevalence in children/adolescents is roughly equal between sexes or substantially higher in females, in pooled meta-analyses of population-based echo screening.medium-high
One sidePooled RHD prevalence in males was almost equivalent to (slightly higher than) females
males 10.15/1000 (95% CI 6.84-13.47) vs females 9.72/1000 (95% CI 6.46-12.98) · Sub-analysis of 15 of 22 African population-based echocardiographic screening studies (WHF criteria) that reported sex-disaggregated data; general population/school-age, multiple African countries
abdu-2024-prevalence-pattern-rheumatic
The other sideMales have significantly lower RHD prevalence than females; pooled female excess of about 1.4:1
RR males vs females = 0.70 (95% CI 0.61-0.80) for latent/echo-detected RHD; RR 0.71 (95% CI 0.59-0.86) for definite RHD specifically · 28 of 58 population-based echocardiographic screening studies (WHF/WHO criteria) that stratified by sex, children/adolescents aged 5-20 in RHD-endemic areas worldwide (includes African studies)
mutarelli-2025-global-prevalence-sex
Why they cannot both be rightBoth are pooled meta-analyses of population-based echo screening in children/adolescents using the same WHF/WHO diagnostic criteria, and their overall prevalence estimates agree closely (abdu-2024 definite RHD 8.91/1000 vs mutarelli-2025 definite RHD 9/1000) -- so they are pooling comparable underlying data. Yet on sex, abdu-2024 finds males and females statistically indistinguishable (males even nominally higher), while mutarelli-2025 finds a robust, statistically significant female excess (RR 0.70-0.71, CI excluding 1) for both latent and definite RHD. A ~30% male deficit that is statistically significant cannot coexist with 'almost equivalent' rates in a comparable population using the same criteria -- one characterization of the sex effect is wrong, or the effect genuinely differs between the African subset and the global pool. abdu-2024's finding is also the outlier against nearly every other sex-stratified claim in this corpus (Fiji OR 5.1 female excess, Tanzania +83% female, Brazil 48 vs 35/1000, Australian NT ASRR 1.7-1.9 female excess, Cameroon M/F ratio 0.56, pooled diagnosed-case counts 152 female vs 130 male).
The innocent explanation, and whether it was checkedMost likely innocent explanation: mutarelli-2025's 28 sex-stratified studies are a global mix (Africa, Pacific, Asia, Latin America) and the female excess could be driven by non-African regions, with Africa specifically showing no sex difference -- real regional heterogeneity rather than a true conflict. This is plausible but NOT fully checked here: the claims file does not list which individual studies feed each meta-analysis, so it is unverified whether abdu-2024's 22 African studies are a subset of mutarelli-2025's 58, and whether an African-only sex-stratified subgroup exists inside mutarelli-2025 to compare directly. Until that subgroup comparison is done, the two pooled headline numbers remain in direct tension for anyone citing 'African RHD is sex-neutral' vs 'RHD has a global female excess including in Africa.'
What depends on which is rightThe direction of the sex effect feeds directly into the leading mechanistic explanation for RHD (autoimmune/hormonal susceptibility predicting female excess) and into whether screening programs should be sex-neutral or prioritise girls. Citing abdu-2024 alone would tell a funder or ministry that African RHD screening needs no sex targeting; citing mutarelli-2025 (or the majority of individual country studies in this same shelf) says the opposite.
domain: prevalence and burden
The starting and ending acute rheumatic fever incidence figures reported for the same 10-year Pinar del Rio, Cuba primary/secondary prevention programmemedium
One sideA 10-year community-based primary and secondary prevention programme in Pinar del Rio, Cuba reduced first ARF attack incidence from 12.2 to 2.1 per 100,000
12.2 per 100,000 (1986) -> 2.1 per 100,000 (1996) · Pinar del Rio, Cuba, population aged 5-25 (n=273,933); citing Nordet 2008 directly · not specified (Nordet 2008 primary source)
abrams-2020-integrating-prevention-control
The other sideComprehensive 10-year community programs in Martinique, Guadeloupe, and Cuba achieved dramatic reductions in ARF incidence; in Cuba from 18.6 to 2.5 per 100,000
18.6 per 100,000 -> 2.5 per 100,000 · Cuba, described as the same type of 10-year comprehensive community programme, no separate denominator/age band given · not specified
zuhlke-2013-primary-prevention-rheumatic
Why they cannot both be rightBoth claims describe the identical intervention -- a single 10-year comprehensive primary/secondary ARF prevention programme in Cuba, and both are the ONLY Cuban programme of this description in the corpus (repeatedly cross-referenced elsewhere in the file as the Nordet 2008 Pinar del Rio study, 1986-1996). A single historical time series has one starting value and one ending value. 12.2 and 18.6 per 100,000 cannot both be the 1986 baseline incidence for the same population over the same programme; likewise 2.1 and 2.5 cannot both be the 1996 endpoint. The two numbers are also verbatim-quoted from their respective source PDFs, so this is not a transcription slip in this claims file -- the discrepancy exists in the published literature itself, most likely as a citation-chain error in one of the two secondary reviews.
The innocent explanation, and whether it was checkedChecked: is zuhlke-2013 possibly describing a different, non-Pinar-del-Rio Cuban programme, or national-level data instead of the Pinar del Rio provincial data? The claims file gives no separate citation number or alternate programme name for the zuhlke figure -- it is grouped with Martinique/Guadeloupe under 'these 10-year programs' with no indication it is a different intervention, and every other Cuba entry in this corpus (7 separate claims, spanning cost, compliance, severe-case counts) traces to the same Nordet/Pinar del Rio programme. No second Cuban programme of comparable scale appears anywhere else in the file. This does not rule out that zuhlke-2013 pulled its number from an earlier interim publication of the same ongoing programme (e.g. a mid-1990s progress report before the final 1996 figures were available), which would explain a higher baseline and slightly higher endpoint -- but nothing in the extracted text supports or refutes that reading, so it remains unresolved rather than dissolved.
What depends on which is rightThis is the single most-cited real-world proof point in the corpus that comprehensive ARF/RHD control programmes work at population scale (repeated in at least 7 other claims across different papers, used to argue for register-based programmes in India, Egypt, and elsewhere). If the true reduction was 12.2->2.1 (83% relative decline) versus 18.6->2.5 (87% relative decline), the headline story is similar either way -- but a reader citing the wrong absolute baseline (18.6 vs 12.2) when benchmarking against a different country's current incidence would misjudge how large a starting problem the programme actually solved.
domain: ARF diagnosis, presentation and treatment
How much weight the PARF (peptide associated with rheumatic fever) M-protein/collagen-IV mimicry mechanism deserves as a driver of RHD autoimmunitymedium
One sideThe N-terminal PARF motif of the GAS M protein binds the CB3 region of collagen type IV, forming an autoantigenic complex that induces collagen-cross-reactive antibodies and aggravates inflammation implicated in RHD's pathological progression.
n/a (mechanistic review, summarising Dinkla et al. 2003, 2009; Tandon et al. 2013) · Review-level synthesis, no single named cohort · in vitro / review of prior mechanistic studies
lumngwena-2022-pathophysiology-rhd-outstanding
The other sideOnly 1 of 74 real-world ARF-associated GAS isolates contained the PARF motif, undermining PARF-collagen mimicry as a general or sole mechanism; 'additional and/or complementary mechanisms are likely to be involved'.
1/74 isolates (1.4%) PARF-motif positive · New Zealand GAS strains isolated from confirmed ARF cases · human clinical isolates, molecular typing
zuhlke-2017-group-streptococcus-acute
Why they cannot both be rightIf the PARF-collagen complex is presented as a named pathogenic mechanism contributing to RHD's autoimmune valve damage, that mechanism cannot be doing that work in the 98.6% of real ARF-causing strains that lack the motif at all. Either PARF-collagen mimicry is a minor/strain-specific contributor rather than a general mechanism, or the reviews citing it as 'the' mimicry-to-collagen pathway are overstating its generality relative to the strains that actually cause disease.
The innocent explanation, and whether it was checkedPartially. Zuhlke et al. do not claim PARF is irrelevant in the one strain that carries it, and the original PARF-discovery papers (Dinkla 2003/2009, Tandon 2013) were working with a specific set of GAS reference/nephritis- or rheumatogenic-associated strains, not a random sample of field isolates from a single country's ARF cases. So this could be 'PARF matters in a strain-specific subset' rather than a true logical conflict — but the reviews that cite PARF-collagen binding as a mechanism (without noting the 1/74 prevalence) read as if it generalises, and it doesn't in this NZ isolate collection. Checked: yes, read both source papers' framing in full context.
What depends on which is rightIf PARF-collagen mimicry is a minor, strain-restricted phenomenon rather than a core RHD mechanism, then mechanistic models and any future diagnostic/therapeutic target built on PARF-collagen cross-reactivity would only apply to a small fraction of real-world cases — the field would need to keep looking for the mimicry target(s) that explain the other ~98%.
domain: pathogenesis, immunology and genetics
Direction of circulating IL-2 in ARF/RHD — strongly elevated in acute disease vs. low/deficient in valve diseasemedium
One sideAcute ARF plasma shows strong elevation of IL-2 (alongside IL-7), part of the discriminating proteomic signature versus controls.
Reported as one of the most strongly elevated cytokines (SomaScan relative abundance; no absolute pg/mL given in the claim) · Ugandan children aged 3-17, acute definite ARF vs alternate-diagnosis/healthy/RHD controls · human, acute-phase plasma proteomics
carapetis-2025-plasma-protein-biomarker
The other sideLow levels of IL-2 and deficiency of circulating regulatory T cells are associated with rheumatic mitral-valve disease, with greater Treg deficiency in patients with multiple valve impairment.
Directional only ('low IL-2'), citing refs 44, 52, 53 in the review; no pooled effect size given · Patients with established rheumatic mitral-valve disease (chronic RHVD), mixed literature synthesis · human, chronic/end-stage disease
passos-2021-rheumatic-heart-valve
Why they cannot both be rightBoth claims describe circulating IL-2 as a marker of disease state, in opposite directions: strongly UP in acute ARF plasma, LOW in chronic rheumatic valve disease. Taken as a single stable biomarker (as some review language implies), IL-2 cannot be simultaneously the most elevated discriminating protein in ARF and a marker of deficiency in RHD.
The innocent explanation, and whether it was checkedMost likely yes — ARF (acute rheumatic activity, days to weeks) and RHVD (established, often years-old valve damage) are different disease stages on the same continuum, and an acute IL-2 burst during active T-cell activation followed by IL-2/Treg exhaustion in chronic disease is biologically plausible (consistent with the Treg-deficiency literature in the same review). No single paper in this shelf tracks IL-2 longitudinally from acute ARF through to chronic RHVD in the same patients, so the disease-stage explanation is plausible but not directly confirmed here.
What depends on which is rightIf IL-2 trajectory really does invert between acute and chronic phases, that reframes IL-2 (and Treg biology generally) as a marker of disease PHASE rather than disease presence, which matters for anyone proposing IL-2-based biomarkers or IL-2-directed immunotherapy (e.g. the low-dose IL-2 rat studies elsewhere on this shelf) without specifying which phase they're targeting.
domain: pathogenesis, immunology and genetics
Whether the landmark 1998-2001 South Auckland (NZ) school-based sore-throat clinic RCT (Lennon et al.) found a statistically significant reduction in first-attack rheumatic fever, or no significant reduction.medium
One sideA school-based sore-throat clinic program in South Auckland, New Zealand (1998-2001) showed a statistically significant ~33% reduction in first-attack RF using 1965 Jones criteria.
RR 0.67, 95% CI 0.48-0.93 (29 RF cases in controls per 31,531 person-yr vs 24 cases in intervention schools per 32,254 person-yr) -- CI excludes 1.0 · ~24,000 school children aged 5-18 in schools with >=70% Maori/Pacific students, ~87,000 person-years, 1965 Jones criteria · South Auckland, New Zealand, single RCT, 1998-2001
lennon-2009-meta-trials-streptococcal
The other sideA New Zealand randomised trial of school-based sore-throat clinics showed no significant reduction in ARF incidence.
no effect size given, described only as 'no significant reduction' · children enrolled through the school-based sore-throat clinic programme, same NZ trial referenced as 'a previous randomised trial' · New Zealand, cited secondhand in a Strep A vaccine development review
sheel-2015-development-group-streptococcal
Why they cannot both be rightBoth claims describe the same single RCT (Lennon et al. 2009, South Auckland) and its effect on first-attack ARF incidence. Lennon's own reported result under 1965 Jones criteria is RR 0.67 with a 95% CI of 0.48-0.93 -- a confidence interval that does not cross 1.0, which is the textbook definition of statistically significant. Sheel-2015 characterises the exact same trial as showing 'no significant reduction.' A single trial's primary result cannot be both significant and not significant when both descriptions are talking about the same comparison.
The innocent explanation, and whether it was checkedPartially, and this is the most likely explanation. The extracted claim for lennon-2009 explicitly qualifies its significant result as 'using 1965 Jones criteria' -- language that only makes sense if the same trial was also analysed under a second, stricter/revised Jones criteria (1992-era), which is a well-known feature of RF trials spanning that criteria transition. It is plausible the trial's result was significant under the older, looser 1965 criteria but not under the revised criteria, and that Sheel-2015 is citing the non-significant analysis while lennon-2009 (self-reporting) highlights the significant one. I could not confirm this from the CLAIMS-FLAT extracts alone -- no claim in this shelf reports the trial's revised-Jones-criteria result explicitly, so the dual-criteria explanation is inferred, not verified. Even if that is the mechanism, it means a reader citing either paper's characterisation without stating which Jones criteria was used walks away with the opposite conclusion about whether NZ's foundational RF-prevention RCT worked.
What depends on which is rightThis is the single RCT most often cited as the evidence base for school-based sore-throat screening programmes globally. Whether it 'worked' (significant) or 'didn't' (not significant) is cited in both directions elsewhere in this literature to argue for or against replicating the model, and the true answer depends on which diagnostic criteria the citing paper is silently using.
domain: control programmes, registers and economics
What fraction of a first episode of acute rheumatic fever (ARF) actually progresses to rheumatic heart disease (RHD), measured empirically over roughly a decade of follow-up in a comparable high-burden, Indigenous-majority, high-income-country population with an active national ARF/RHD register.medium
One sideProgression from first ARF episode to RHD is fast and common: highest incidence in year 1, with cumulative progression already over a quarter of the cohort by 1 year and over half by 10 years.
Incidence 35.92 per 100 person-years in year 1; cumulative progression 27.1% at 1 year, 44.0% at 5 years, 51.9% at 10 years · 572 first-episode ARF patients (97% Indigenous) on the Northern Territory (Australia) RHD Register, 1997-2013; RHD ascertained via the register, which combines clinical diagnosis with the NT's echocardiographic active-surveillance control program · Up to ~17 years (1997-2013 enrolment window); cumulative incidence reported at 1, 5 and 10 years
he-2016-long-term-outcomes
The other sideProgression from first ARF to RHD is much rarer and slower: at the cohort's actual median follow-up of about a decade, fewer than 1 in 12 initial-ARF patients had been hospitalized for RHD, and even a theoretical lifetime (26.8-year) extrapolation puts RHD probability under a quarter.
7.9% hospitalized for RHD (173/2,182) and 13.6% experienced any disease progression (recurrent ARF, RHD hospitalization or circulatory death) by end of observation (median follow-up 10.4 years); Kaplan-Meier extrapolation to 26.8 years estimates 23.5% probability of RHD hospitalization specifically · 2,182 patients hospitalized with initial ARF in New Zealand, 1989-2012, predominantly Māori and Pacific Islander, followed via linked national hospital/mortality records to end-2015 · Median 10.4 years actual observed follow-up (max 26.8 years); outcome ascertained via hospital discharge/mortality record linkage, not an active screening register
oliver-2021-ethnically-disparate-disease
Why they cannot both be rightBoth papers answer the same natural-history question -- what share of a first ARF episode converts to RHD -- in structurally similar settings: high-income countries with an Indigenous population carrying a disproportionate ARF/RHD burden and an established national control program, over comparable ~10-year windows. Yet He-2016 finds just over half of first-episode ARF patients have progressed to RHD by 10 years (and over a quarter already by year 1), while Oliver-2021 finds only about 8% had been hospitalized for RHD by a similar (in fact slightly longer, median 10.4-year) follow-up, rising to at most ~23.5% even under a full-lifetime 26.8-year extrapolation. A ~6-7x gap at the same time horizon, in two population-based cohorts from comparable high-burden settings, cannot both describe the same underlying biological progression rate.
The innocent explanation, and whether it was checkedThe leading candidate explanation is ascertainment method, and it was checked as far as this corpus allows: He-2016's outcome is drawn from the NT RHD Register, which is fed by the Northern Territory's long-running active echocardiographic screening/control program and therefore captures RHD diagnosed on screening (including mild, asymptomatic disease) as well as clinically apparent disease. Oliver-2021's outcome is restricted to RHD severe enough to generate a hospital admission, via administrative ICD-coded discharge data, with no equivalent nationwide active-surveillance register described in the draft. This alone plausibly explains a large part of the gap (register-based ascertainment will always find more disease than hospitalization-only ascertainment), and the widely-repeated 'up to 60% of ARF progresses to RHD' estimate that recurs elsewhere in this corpus (antunes-2020, muhamed-2019, both citing Carapetis et al. 2005) sits closer to He-2016's number than to Oliver-2021's, which is circumstantial support for He-2016 rather than independent confirmation (all three ultimately trace to variants of active-surveillance-based case-finding). What was NOT independently confirmed: whether Oliver-2021's non-hospitalized RHD is truly absent from national ascertainment (New Zealand does not appear, from this corpus, to run an NT-style universal echo-screening register), and whether case-mix (age at first ARF, carditis severity at presentation, access to primary care) differs enough between the NT and NZ cohorts to independently explain part of the gap. Given the size of the gap (6-7x) is large even for an ascertainment difference, and it recurs at the 1-year mark too (27.1% NT vs presumably far under that in NZ, not separately reported), this is flagged as a real, only partially-dissolved conflict rather than fully innocent.
What depends on which is rightThis number is the anchor statistic for RHD burden-of-disease estimates, control-program cost-effectiveness models, and any claim about how urgently a first ARF episode needs secondary prophylaxis to prevent RHD. Using He-2016's ~52% would argue for aggressive, register-based active surveillance as the only way to see the true burden; using Oliver-2021's ~8-24% (from hospitalization records alone) would substantially understate how many first-ARF patients are heading toward RHD, and would make a hospitalization-only surveillance system look reassuring when it may simply be blind to most of the disease.
domain: natural history and progression
Whether standard 4-weekly benzathine penicillin G actually maintains protective serum levels for most of the dosing interval (pharmacokinetic evidence) versus whether the standard regimen is highly effective at preventing recurrence in practice (population-level clinical outcome evidence).medium
One sideNo participant on standard monthly BPG maintained benzylpenicillin concentrations above the 0.02 mg/L (~20 ng/mL) target for the entire interval between injections; higher-BMI patients spent essentially none of the cycle above target.
0 of 18 participants covered the full interval; median duration above target = 9.8 days (of 28) in lower-BMI participants, 0 days in higher-BMI participants. Ethiopian cohort (ketema-2021): patients stayed above a 20 ng/mL target for only about 2 of the 28 days on average, and above a lower 10 ng/mL target for only ~20 days (73%) of the cycle. · Hand 2019: 18 children/adolescents with RHD on monthly BPG, Perth, Australia, 256 plasma concentrations over 6 cycles. Ketema 2021: 74 Ethiopian ARF/RHD patients on 4-weekly IM BPG, modelled serum concentrations over one 28-day cycle.
hand-2019-population-pharmacokinetic-benzathine (supported by ketema-2021-high-risk-early)
The other sideAmong patients on regular BPG in a large multi-country secondary-prevention programme, ARF recurrence was rare, and the great majority of recurrences occurred in patients who were NOT on regular BPG.
ARF reoccurrence in 0.4% of patient-years; only 2 of 53 recurrent cases were on regular BPG · WHO Phase I multi-country secondary-prevention programme, 16 developing countries, 4 years (1986-90), 33,651 registered patients on the standard 4-weekly-equivalent regimen
abrams-2020-integrating-prevention-control
Why they cannot both be rightIf, as the PK data show, the overwhelming majority of the 28-day cycle is spent below the accepted protective threshold on the standard regimen, clinical breakthrough recurrence should be substantially more common among 'regular' patients than the WHO programme observed. Instead the WHO data show 'regular' BPG is associated with a roughly 96% share of protection against recurrence (51 of 53 recurrences were in irregular patients). Sub-therapeutic levels for most of the cycle and near-complete clinical protection from the same regimen is a genuine tension, not just noise.
The innocent explanation, and whether it was checkedPartially, and the authors of hand-2019 say so explicitly: 'The discordance of this observation with reported efficacy of BPG to prevent rheumatic fever implies a major knowledge gap.' Two innocent explanations were checked: (1) the 0.02 mg/L / 20 ng/mL PK target is derived from in-vitro MIC breakpoints and, per hand-2019 and ketema-2021 themselves, 'may not reflect the concentration needed to prevent pharyngeal GAS acquisition or colonisation' -- the threshold itself could simply be set too high, meaning sub-threshold does not equal unprotected; (2) 'regular BPG' in the WHO programme is an administrative/registry definition (doses received on schedule) that was never validated against actual serum levels in that cohort, so it is possible some 'regular' patients were also running sub-therapeutic without it being detected as a recurrence (asymptomatic GAS reinfection without progressing to ARF). Neither explanation is confirmed by direct evidence in this shelf, so the tension stands as a real and currently unresolved discordance rather than a dissolved artifact.
What depends on which is rightThis is the central rationale question behind interval and dosing debates (whether to move to 3-weekly, higher doses, or new long-acting formulations). If the PK finding is taken at face value, the standard regimen looks pharmacologically inadequate and would justify a costly, painful move to more frequent or higher-dose injections. If the clinical outcome data is taken at face value, the standard regimen already works about as well as anything realistically achievable, and the PK target itself is likely the wrong benchmark. Guideline bodies cannot act on both readings simultaneously.
domain: penicillin prophylaxis
What counts as 'adherent'/'optimal' varies enough between papers, even within similar hospital-registry settings, to make the headline adherence percentages non-comparable -- a definition conflict, not a data conflict.medium
One sideOnly 32% of surveyed RHD patients in Khartoum, Sudan were 'optimally adherent' to monthly BPG.
32% (127/397) · 397 RHD registry patients aged 12-90, Khartoum, Sudan · Self-report of >=6 BPG injections in the preceding 6 months -- i.e. essentially 100% dose completion for a monthly regimen, not the >=80% threshold used almost everywhere else in this shelf
edwards-2021-health-system-patient / edwards-2021-health-system-factors
The other sideGood adherence to BPG was 63-64% in comparable Ethiopian hospital-registry cohorts.
63.6% (Bahir Dar, n=346) and 63% (Jimma zone, n=253) · Ethiopian RHD patients on BPG for >=1 year, attending cardiac clinics · Good adherence = >=80% of expected annual injections (>=10 of 13 doses/year)
kelemu-2025-adherence-benzathine-penicillin / adem-2020-rheumatic-heart-disease
Why they cannot both be rightNot a true contradiction in the underlying reality -- these are different yardsticks measuring different things, which is itself the finding. The Khartoum 'optimal adherence' bar (100% dose completion, no missed doses in 6 months) is structurally stricter than the near-universal >=80%-of-doses standard used in the Ethiopian cohorts and the pooled 33-study meta-analysis (bimerew-2024, pooled adherence 58.5% at the >=80% cutoff). A reader who does not check the fine print would compare Khartoum's 32% against Ethiopia's 63-64% and conclude Sudanese care is roughly half as good, when the gap could be substantially a measurement artifact rather than a real difference in patient behaviour.
The innocent explanation, and whether it was checkedThis IS the dissolution -- the apparent gap between 32% and 63-64% is explained, not just excused, by the stricter Khartoum definition. What it does not dissolve is the underlying problem: because every 'adherence' number in this literature carries its own implicit threshold and measurement method (self-report vs register data, patient-level vs group-level, 6-month vs annual window), none of the headline percentages in this shelf can be safely ranked against each other without re-deriving them from the same definition first.
What depends on which is rightThis is exactly the kind of definitional drift that makes cross-country adherence league tables misleading. Any argument that 'adherence is worse in country X than country Y' needs the definitions checked line by line before it can be used to direct resources or set targets; bimerew-2024's own pooled analysis shows adherence ranged from 10.7% to 92.9% across 33 studies using nominally the same >=80% cutoff, which is itself only explicable by further definitional and measurement-method heterogeneity (self-report vs pharmacy/register data, patient-level vs group-level denominators).
domain: penicillin prophylaxis
Whether echo-detected RHD prevalence differs by sex in African population-based screening studies, pooledmedium
One sidePooled RHD prevalence (combined definite+borderline, 'RVHD') in males was almost equivalent to that in females
Male 10.15/1000 (95% CI 6.84-13.47) vs Female 9.72/1000 (95% CI 6.46-12.98) — CIs almost fully overlapping · 15 of 22 African population-based echocardiographic screening studies (2015-2023) that reported sex-disaggregated data · Random-effects meta-analysis of study-level prevalence estimates, WHF echocardiographic criteria, combined definite+borderline disease
abdu-2024-prevalence-pattern-rheumatic
The other sideLatent/echo-detected RHD prevalence is significantly lower in males than females, and the same holds when restricted to definite RHD only
RR males-vs-females 0.70 (95% CI 0.61-0.80, p<0.01) for combined latent RHD, pooled female:male ratio ~1.4:1; RR 0.71 (95% CI 0.59-0.86, p<0.01) for definite RHD alone · 28 of 58 population-based echocardiographic screening studies (2001-2022, children/adolescents aged 5-20) in RHD-endemic areas worldwide that stratified prevalence by sex; Africa is the second-highest-prevalence continent in this same 58-study set (25/1000), so African studies are a substantial share of the 28 · Direct risk-ratio comparison pooled across sex-stratified studies (not adjusted meta-regression), same WHF/WHO echo criteria family as abdu-2024
mutarelli-2025-global-prevalence-sex
Why they cannot both be rightBoth are 2020s meta-analyses of the same underlying literature (population-based echo screening of African children/adolescents using WHF criteria, overlapping 2015-2022 windows) asking the identical question — is combined (definite+borderline) RHD prevalence higher in females. abdu-2024's African-only sex-stratified subset (k=15) finds the sexes statistically indistinguishable, with the male estimate numerically higher. mutarelli-2025, pooling a larger and more geographically broad set of sex-stratified studies (k=28, of which Africa is the largest non-Oceania continent by prevalence), finds a robust ~40% female excess whose CI excludes the null, and the same direction holds when restricted to definite RHD only. A reader cannot take both pooled estimates as representative of the same population and get materially different answers to 'does sex matter for RHD risk.'
The innocent explanation, and whether it was checkedPartly plausible but not confirmed. abdu-2024's k=15 is almost certainly a subset of mutarelli-2025's k=28 (both drawing on the same African echo-screening literature 2015-2023), so if the studies overlap heavily this could be a case of the smaller, Africa-restricted pool washing out a real but continent-varying effect, or of publication-bias-adjusted vs unadjusted pooling (abdu's own trim-and-fill roughly doubled its overall prevalence estimate, showing the pooled figures are sensitive to which studies are included). abdu-2024's authors themselves offer a non-quantitative post-hoc explanation ('more sensitive screening methods') rather than a tested one. I did not obtain the actual study list for either k=15 or k=28 to confirm the overlap, so this is checked (context read, overlap argued) but not fully dissolved — it remains a genuine, unresolved discrepancy between the two most recent African/global sex-stratified pooled estimates.
What depends on which is rightThis is the exact sex-ratio question the domain brief flags as a known live case. It matters for screening design (should girls-only or mixed-sex school screening be prioritized?) and for how confidently a communications piece can claim 'RHD hits girls twice as hard' versus the more cautious 'RHD may be somewhat more common in girls, but the African evidence itself is split.'
domain: risk factors and disparities
Whether rural residence is a significant risk factor for RHD prevalencemedium
One sideNo significant difference in RHD prevalence between rural and urban settings
Non-significant in multivariable meta-regression (reported alongside a non-significant school-vs-community coefficient of -0.0068, p=0.899) · 82 population-based echocardiographic screening studies worldwide, 1,090,792 participants, ages 3-74, 1996-2017 · Multivariable meta-regression across all 82 studies simultaneously, adjusting rural/urban alongside diagnostic criteria, country income, WHO region and study setting
noubiap-2019-prevalence-progression-rheumatic
The other sideChildren in urban areas have about half the latent RHD prevalence of those in rural areas
RR (urban vs rural) 0.49, 95% CI 0.26-0.93, p=0.03, I2=88% · 10 of 58 population-based echocardiographic screening studies (2001-2022) that directly compared rural vs urban populations, children/adolescents aged 5-20 in RHD-endemic areas · Direct unadjusted risk-ratio comparison across the 10 studies that reported rural-vs-urban prevalence, not adjusted for other covariates
mutarelli-2025-global-prevalence-sex
Why they cannot both be rightnoubiap-2019 concludes rural/urban residence carries no significant prevalence signal once folded into a global multivariable model; mutarelli-2025 concludes urban children have roughly half the rural prevalence, with a CI that (barely) excludes 1. Same disease, same diagnostic-criteria family (WHF/WHO echo screening), same broad outcome (population-based latent RHD prevalence) — one says the variable does not matter, the other says it roughly halves risk.
The innocent explanation, and whether it was checkedLargely, yes — this is the weaker of the pair. noubiap-2019 is an unrestricted-age (3-74y) global meta-regression that adjusts rural/urban simultaneously with diagnostic criteria, country income and WHO region, so a real rural effect could be absorbed by the correlated income/region terms. mutarelli-2025's comparison is a simple pairwise RR from only 10 studies restricted to children/adolescents 5-20, with very high heterogeneity (I2=88%) and a CI (0.26-0.93) that only just clears the null — this is a fragile significant result, not a strong one. The most likely honest read is 'weak, heterogeneity-laden evidence for a rural excess that a more conservative, covariate-adjusted, all-ages model does not detect,' rather than two solid findings that flatly disagree.
What depends on which is rightProgram design money follows this: if rural residence independently predicts RHD, screening and prevention budgets should be weighted toward rural school programs; if it washes out once income and region are controlled, the same money is better spent targeting poverty/access directly regardless of rural/urban label.
domain: risk factors and disparities
Whether acute rheumatic fever (ARF) incidence itself — as opposed to RHD prevalence — differs by sexmedium
One sideMales and females are equally likely to have acute rheumatic fever (RHD, not ARF, is the one that is more common in females)
No effect size given — stated as a categorical 'equally likely' · General/global ARF-RHD population (review-level generalization); the same claim, attributed to the REMEDY registry, appears independently in lupieri-2025-rheumatic-heart-valve ('ARF appears to be equally found in males and females') · Narrative synthesis in a clinical review, not a primary incidence study
karthikeyan-2018-acute-rheumatic-fever
The other sideAmong Indigenous people in the Northern Territory, females had substantially higher first-episode ARF incidence than males
49% higher ARF incidence in females than males, adjusted for age and year · NT Indigenous residents, first ARF episode, 1997-2010 (population-based incidence registry) · Adjusted incidence rate comparison within a whole-population disease registry
lawrence-2013-arf-rhd-northern-territory
Why they cannot both be rightThe 'ARF is sex-neutral' line is stated as a general clinical fact by two independent review sources (karthikeyan-2018 and lupieri-2025, the latter citing REMEDY). But the two largest, most rigorous population-based ARF incidence studies in this evidence base — both drawn from the same high-burden Australian Indigenous population — directly measure a female excess in ARF itself: 49% higher incidence in the NT (lawrence-2013, 1997-2010) and 1.3x the risk in the Top End (carapetis-2001, 1987-1996, 95% CI 1.0-1.7). A textbook 'equally likely' cannot be squared with two of its own evidence base's largest whole-population ARF registries both finding a significant female excess.
The innocent explanation, and whether it was checkedPartially — population-specificity is the obvious candidate explanation, and there is direct support for it in the same shelf: a single rural Indian ARF incidence study (Grover 1993, cited within a different systematic review) found no sex difference, consistent with the 'equal' generalization holding in at least one non-Australian setting. So the honest resolution is probably 'globally/on average ARF may be closer to sex-neutral, but the Australian Indigenous population — the single best-instrumented ARF registry in this literature — is a documented, sizeable exception,' not that either paper made an error. I checked this (found the Indian counter-example) but could not confirm whether karthikeyan-2018's 'equally likely' claim was itself derived from a global synthesis that already includes or excludes the Australian Indigenous data, so it is not fully dissolved.
What depends on which is rightTextbook teaching material for clinicians ('ARF hits boys and girls equally, RHD favors girls') is the kind of one-liner that gets repeated in training decks. In the one population where it has been tested most rigorously, it is off by roughly 30-50% in the female direction — worth knowing before that line goes into any Lead+D-adjacent teaching content on the topic, and worth knowing for Indigenous-health-specific screening/education framing.
domain: risk factors and disparities
Whether task-shifted handheld/portable echocardiography screening (non-experts, read against standard echo) has high specificity (~85%) or very low specificity (~20%) for detecting latent RHD in a school-age population.medium
One sidePooled across 7 screening studies of school-aged children/adolescents in high-prevalence areas, handheld echocardiography versus standard echocardiography for screening any RHD (borderline or definite) had high sensitivity and specificity, rated high certainty by GRADE.
sensitivity 0.79 (95% CI 0.73-0.84), specificity 0.85 (95% CI 0.80-0.89), AUC 0.90 (0.85-0.94) · 7 pooled screening studies, school-aged children/adolescents (5-20y) in high-prevalence RHD regions; reference standard = standard echocardiography
providencia-2024-handheld-echocardiography-screening
The other sideScreening handheld echocardiography by nonphysicians has high sensitivity but very low specificity for any latent RHD when checked against confirmatory standard echocardiography.
sensitivity 95.5% (95% CI 84.3-99.4), specificity 19.7% (95% CI 10.9-31.3), PPV 5.3%, NPV 98.9%, in a prevalence scenario around 4.5% · 1,390 Brazilian schoolchildren (ages 11-19, mean 14), 6 low-income public schools, metropolitan Belo Horizonte, PROVAR+ substudy 2022-2023, 2012 WHF criteria; 110 screen-positive children got confirmatory standard echo, 1,374 screen-negative children were 'presumably' negative and not independently re-examined
diniz-2024-agreement-between-handheld
Why they cannot both be rightBoth purport to measure the same quantity: the specificity of a non-expert-operated handheld/portable echo screen for latent (borderline-or-definite) RHD in a school-age population, checked against standard echocardiography. Providencia-2024's pooled estimate (0.85, CI 0.80-0.89) and diniz-2024's estimate (0.197, CI 0.109-0.313) do not overlap at all -- roughly a fourfold gap. That gap is not a rounding difference: at 85% specificity a screening program produces about 15 false positives per 100 disease-free children; at 20% specificity it produces about 80, meaning four out of five children referred for confirmatory echo do not have RHD. Whether task-shifted handheld screening is a workable frontline tool, which is the headline conclusion of the WHF's own review (rwebembera-2023, 'a reasonable alternative option'), depends entirely on which of these two numbers is true. Notably diniz-2024 is not an isolated low outlier against just one comparator -- ali-2021 (pooling Lu 2015 and Ploutz 2016) reports specificity 82.4-87.3%, jaimesreyes-2022 reports 85% for the same PROVAR non-expert program, and lamichhane-2023 pools three field studies of non-expert HHE at 79-92% specificity. Every other specificity figure for this exact modality in this shelf clusters at 79-92%; diniz-2024 alone reports 19.7%.
The innocent explanation, and whether it was checkedPartially checked. The most likely mechanism is a partial-verification design: diniz-2024 only sent the 110 screen-positive children to confirmatory standard echo; the 1,374 screen-negative children were assumed disease-free and never independently examined. The claim text itself flags this ('false negatives were not verified by standard echo, limiting true accuracy estimates') and the reported specificity is stated 'in a prevalence scenario around 4.5%', which reads as a figure back-derived from an assumed background prevalence via Bayes' theorem rather than counted directly from a fully cross-verified 2x2 table -- which is how the other studies (all with both arms examined by standard echo) computed their 79-92% range. I could not confirm the exact statistical method from the extracted claims alone; that would need the full diniz-2024 methods section, which is outside this shelf. Even granting that explanation, it does not fully dissolve the contradiction as it exists in the literature: both figures are published as 'specificity of handheld echo screening for latent RHD,' and a reader who pulls one number without checking the verification design will draw the opposite conclusion about whether task-shifted screening is fit for purpose.
What depends on which is rightThis is the number that decides whether an RHD screening program can task-shift handheld echo to non-expert operators and trust the positive/negative calls at the point of care, or whether every screen-positive needs expert confirmatory echo before any action -- with large workforce and cost consequences, since a program run at 20% specificity floods confirmatory services with roughly four false alarms for every true case.
domain: screening and diagnostic criteria
Whether pooled RHD prevalence differs significantly between the 2012 WHF and the WHO echocardiographic diagnostic criteria, across the same general body of population-based screening literature.medium
One sidePooled RHD prevalence differed substantially by diagnostic criteria, more than double for WHF versus WHO.
WHF 26.1/1000 (95% CI 19.2-33.1) vs WHO 11.3/1000 (95% CI 7.2-16.2); reported as a stated comparison (p<0.0001 pattern applied elsewhere in the same meta-regression) · 82 population-based echocardiographic screening studies worldwide, 1,090,792 participants ages 3-74, published/conducted 1996-2017
noubiap-2019-prevalence-progression-rheumatic
The other sidePooled RHD prevalence by WHF and WHO criteria was not significantly different.
WHF 25/1000 (95% CI 19-32, 46 studies) vs WHO 24/1000 (95% CI 13-40, 8 studies), 'no significant difference' · 58 population-based echocardiographic screening studies, 215,552 subjects, children/adolescents ages 5-20 in RHD-endemic areas, recruitment 2001-2022
mutarelli-2025-global-prevalence-sex
Why they cannot both be rightBoth are meta-analyses of the same broad literature (population-based echocardiographic screening for RHD) explicitly stratifying pooled prevalence by WHF-2012 versus WHO diagnostic criteria, and their WHF-only pooled estimates are nearly identical (26.1 vs 25/1000) despite different study sets, date ranges and age restrictions. Yet their WHO-only pooled estimates diverge sharply (11.3 vs 24/1000, more than double), and the two reviews reach opposite conclusions about whether the criteria sets are interchangeable: noubiap-2019's numbers show WHF finding roughly 2.3x more RHD than WHO, while mutarelli-2025 states explicitly there is no significant difference. If sampling or population differences alone explained the gap, the WHF estimate should have moved too, and it did not. A reader asking whether historical WHO-criteria screening data can be trend-lined against newer WHF-criteria data gets opposite answers depending which review they cite.
The innocent explanation, and whether it was checkedPartially checked, not confirmed. 'WHO criteria' is not a single fixed rubric across the field's history -- there is an original WHO/Bland-Jones-type auscultation-anchored definition and a later 'modified NIH/WHO' echo criteria used by some groups (e.g. colquhoun-2014's Fiji study cites 'modified NIH/WHO criteria' separately from its own post-hoc WHF reclassification), so the two reviews may be pooling different sets of studies under the same label depending on their inclusion rules and date windows (noubiap 1996-2017, all ages 3-74; mutarelli 2001-2022, ages 5-20 only). Both reviews also report very high heterogeneity for the WHO-criteria subgroup specifically (I2 = 97% in mutarelli, only 8 studies), so a small, heterogeneous subgroup is inherently unstable and sensitive to exactly which primary studies are included. I did not have the underlying study-inclusion lists for either review (only the extracted claims), so I could not confirm whether the two 'WHO criteria' buckets share the same primary studies -- that would need the source papers.
What depends on which is rightWhether WHF and WHO criteria are interchangeable determines whether older WHO-criteria surveillance data can be pooled with newer WHF-criteria data to track a program's impact over time, or whether an apparent trend in RHD burden is actually just an artifact of which criteria set was used in which era.
domain: screening and diagnostic criteria
What strain coverage the 30-valent StreptAnova M-protein vaccine actually provides in North America / the developed-country markets it was specifically designed for.medium
One sideThe 30-valent vaccine was formulated to contain M types prevalent in the United States, Canada and Europe, with the potential to provide immunity against 85% of cases of pharyngitis and invasive infection in these geographic locations.
~85% coverage · GAS pharyngitis and invasive Strep A cases · US, Canada and Europe (pooled)
dale-2020-update-group-streptococcal
The other sideTheoretical coverage of StreptAnova across S. pyogenes strains, based on a large-scale genomic analysis, is 48% globally, ranging from 28% in East Africa to 75% in North America.
75% (North America), 48% (global) · S. pyogenes strains from genomic surveillance · North America specifically, plus global
walkinshaw-2023-streptococcus-pyogenes-vaccine
Why they cannot both be rightBoth claims describe the same vaccine candidate (StreptAnova, 30-valent M-protein) in essentially the same target market. dale-2020 states it was designed to protect against 85% of disease in the exact region walkinshaw-2023 measures at 75% for North America alone (and only 48% worldwide). A vaccine cannot simultaneously be on track for 85% coverage in its core design market and only 75% coverage in the largest single market inside that same pooled region, unless Europe alone is running far above 85% to pull the three-region average up past the North-America-only figure.
The innocent explanation, and whether it was checkedPartially plausible but unconfirmed: if European emm-type coverage is substantially higher than North America's, a pooled US+Canada+Europe average of 85% is arithmetically compatible with a 75% North-America-only figure. I searched the corpus for a Europe-specific StreptAnova coverage number to test this and found none, so I could not confirm or rule it out. The more likely explanation is a real, and reportable, finding: dale-2020's 85% is the vaccine's original design-era target (based on the emm surveillance data available when the 30-valent formulation was fixed), while walkinshaw-2023's 75%/48% comes from a subsequent large-scale genomic reanalysis against more recent circulating strains -- i.e. emm-type replacement since the vaccine was designed, not a measurement error in either paper.
What depends on which is rightStreptAnova is one of only two GAS vaccine candidates that have reached human trials, and downstream cost-effectiveness and DALY-averted models elsewhere in this corpus assume high (80%+) efficacy/coverage to justify introduction. If real-world genomic coverage in the vaccine's own core market is closer to 75% (and 48% globally) rather than the 85% design target, those health-economic cases -- and any regulatory or funding decision that cites the manufacturer's original 85% figure -- are working off an optimistic number.
domain: group A Streptococcus and vaccine development
Whether the only community/school-based RCT of primary prevention (treating sore throats to prevent ARF) -- the Lennon 2009 South Auckland, New Zealand trial -- showed a statistically significant reduction in rheumatic fever.medium
One sideA school-based sore-throat clinic program in South Auckland, New Zealand (1998-2001) showed a statistically significant ~33% reduction in first-attack RF using 1965 Jones criteria.
RR 0.67, 95% CI 0.48-0.93 (CI excludes 1, significant) · ~24,000 school children aged 5-18, ~87,000 person-years, schools with >=70% Maori/Pacific students · South Auckland, New Zealand
lennon-2009-meta-trials-streptococcal
The other sideThe only randomized controlled trial of community (school)-based primary prevention of ARF did not show a statistically significant effect on ARF incidence.
RR 0.81, 95% CI 0.47-1.39 (CI crosses 1, not significant) · School-age children, >85,000 person-years · New Zealand
colquhoun-2009-rheumatic-heart-disease
Why they cannot both be rightThe person-years (~87,000 vs >85,000) and setting make clear these are the same single trial, described as the field's only RCT-grade evidence on primary prevention. One description reports a significant ~33% risk reduction with a confidence interval that excludes the null; the other reports a non-significant ~19% reduction with a CI that comfortably crosses 1. A single trial's headline result cannot be both statistically significant and not significant. The direction of colquhoun-2009's read (no proven effect) is independently echoed by marijon-2012, zuhlke-2013, and sheel-2015 in this corpus, all citing the same trial as showing no significant benefit -- so the corpus is split 1 paper (RR 0.67, significant) against at least 4 papers (null result, RR ~0.81 where a number is given) describing the same trial.
The innocent explanation, and whether it was checkedMost likely yes, and this is the probable explanation: the original Lennon trial paper is known (elsewhere in this corpus, e.g. the Coulehan/Chun/Nordet entries in the same lennon-2009-meta-trials-streptococcal digest) to report results under more than one rheumatic-fever case definition. The RR 0.67 figure this corpus extracted is explicitly labelled 'using 1965 Jones criteria' -- a narrower, confirmed-case analysis. The RR 0.81 figure that colquhoun-2009 (and by direction, marijon-2012/zuhlke-2013/sheel-2015) cite as 'the' trial result is plausibly the trial's primary, broader-case-definition analysis, which several downstream papers chose to headline instead. I could not verify this against the original PDF beyond what is captured in the claims file, since our digest of lennon-2009 only captured the 1965-criteria subgroup number, not a primary/overall RR.
What depends on which is rightThis is the single trial that answers whether population-scale sore-throat treatment programs actually lower ARF rates, and it is cited on both sides of the ongoing policy debate over funding primary-prevention screening versus concentrating resources on secondary prophylaxis and echo screening (a debate visible elsewhere in this corpus, e.g. wyber-2020's finding that primary prevention scale-up alone has a benefit-cost ratio of just 0.2 through 2030). Whether the field's best RCT evidence says 'significant benefit' or 'no proven benefit' materially changes how strongly that debate should lean toward secondary prevention.
domain: group A Streptococcus and vaccine development
Whether mechanical valves give better long-term survival than bioprosthetic valves in young RHD patients, or the reverse -- pooled/general claim vs specific young-cohort studies.medium
One sideIn a retrospective study of 3,638 RHD patients undergoing mitral valve replacement, mechanical valves had superior long-term outcomes (lower all-cause mortality and reoperation) than bioprosthetic valves, particularly in patients under 65; bioprosthetics are only flagged as possibly preferable in low-income regions because of poor INR management
mechanical valves = lower all-cause mortality and lower reoperation, framed as generalising to patients under 65 · 3,638 rheumatic patients undergoing mitral valve replacement (country/setting not specified in the claim -- cited secondhand inside a review) · not stated; carved out as generally applicable except in explicitly 'low-income regions such as Africa'
zhang-2025-recent-advances-prevention
The other sideIn young adults in Saudi Arabia undergoing double valve replacement for RHD, 15-year survival was 92% with bioprosthetic valves versus 76% with mechanical valves; separately, in Maori and Pacific Island women in New Zealand, mechanical valves carried a 2.2x higher relative risk of death than tissue valves despite a 3-fold higher reoperation rate with tissue
15-year survival: bioprosthetic 92% vs mechanical 76% (Saudi Arabia); relative risk of death 2.2x higher with mechanical (New Zealand) · Young adults, RHD, double valve replacement (Saudi Arabia); Maori/Pacific Island women receiving valve replacement (New Zealand) · Saudi Arabia and New Zealand -- both upper-middle/high-income health systems, not the 'low-income regions' zhang-2025 names as the exception
zilla-2024-mechanical-valve-replacement
Why they cannot both be rightzhang-2025's synthesis states mechanical valves are superior 'particularly in patients under 65' with the only stated exception being low-income settings with poor INR access. The Saudi cohort zilla-2024 cites is exactly the population zhang-2025's rule should favour mechanical valves in (young adults, double valve replacement) and in a country that is not low-income, yet it shows the opposite direction at a large magnitude (92% vs 76% survival at 15 years -- a 16-point gap favouring bioprosthetic). The New Zealand finding adds a second population, in an unambiguously high-income country, where mechanical valves carried significantly higher mortality risk despite lower reoperation. A rule with an income-based exception cannot be reconciled with two contrary results from health systems outside that exception unless the true moderator is something other than national income classification.
The innocent explanation, and whether it was checkedMost likely explanation, partially checked: the real moderator may be reliability of anticoagulation monitoring/access rather than national income per se -- zilla-2024's own argument elsewhere in the same paper is that Indigenous/remote populations (Maori/Pacific Islanders in New Zealand, Indigenous Australians) have disproportionately poor access to anticoagulation monitoring despite living in high-income countries, which would explain the New Zealand result without contradicting the underlying mechanism zhang-2025 invokes (poor INR management hurts mechanical-valve outcomes). This is plausible but not fully checked for the Saudi cohort: the claims file gives no INR/monitoring-quality data for that study, so whether poor anticoagulation access (rather than national income) explains the Saudi reversal specifically is unverified. It is also not simply reconcilable by age, since the Saudi cohort is explicitly the young population zhang-2025's rule targets.
What depends on which is rightThis is the central decision in RHD valve surgery for young patients in resource-constrained settings: pick mechanical (durable, but demands lifelong reliable INR monitoring) or bioprosthetic (avoids anticoagulation risk, but degenerates faster in young patients). If the moderator is genuinely 'low income' as zhang-2025 frames it, a wealthy-country young patient should default to mechanical; if the moderator is actually 'reliable anticoagulation access' as the Saudi and New Zealand data suggest, that same recommendation could be actively harmful for patients in well-resourced countries who nonetheless face real barriers to INR monitoring (rural, Indigenous, or otherwise underserved populations).
domain: surgery, valve intervention and anticoagulation
Whether mitral valve repair (MVP) should be recommended over replacement (MVR) 'whenever possible' as an unqualified rule, given what the pooled RHD repair-vs-replacement evidence actually shows once confounding and reoperation are accounted for.medium
One sideMitral valve repair yields better outcomes than replacement in rheumatic mitral regurgitation and should be undertaken whenever possible; experienced centres can repair 75-80% of patients with superior long-term survival versus replacement
75-80% of patients repairable, with survival stated as superior to replacement, no reoperation trade-off mentioned · Patients with severe rheumatic MR requiring surgery (general recommendation, not confounder-adjusted) · General clinical guidance, not restricted by country income
marijon-2012-rheumatic-heart-disease; kumar-2020-contemporary-diagnosis-management
The other sideIn a meta-analysis of the same repair-vs-replacement RHD literature (11 studies, 5,654 patients), MVP's long-term survival advantage over MVR loses statistical significance once analysis is restricted to propensity-score-matched studies, and MVP carries a significantly HIGHER reoperation rate than MVR
Overall survival HR 0.72 (95% CI 0.55-0.95, p=0.020) in favour of MVP, but HR 0.62 (95% CI 0.35-1.11, p=0.110) in PSM-only studies -- not significant; reoperation HR 2.60 (95% CI 1.89-3.57, p<0.001) against MVP; 12-year freedom from reoperation 74.45% (MVP) vs 86.66% (MVR) · Meta-analysis of 11 RHD mitral valve surgery studies (5,654 patients: 1,951 MVP, 3,703 MVR), predominantly retrospective cohorts, mixed PSM/unmatched designs, published up to Dec 2020 · Pooled, mixed-country retrospective cohorts
fu-2021-mitral-valve-surgery
Why they cannot both be rightThe blanket claim 'repair yields better outcomes... should be undertaken whenever possible' asserts an unqualified net benefit with no acknowledged downside. The pooled meta-analysis of essentially the same underlying evidence base shows two things that undercut that framing: (1) the survival benefit for repair shrinks and loses significance once studies are restricted to propensity-score-matched (confounding-adjusted) designs, meaning the raw survival advantage may substantially reflect that healthier patients with more favourable anatomy get selected for repair, not a true causal effect of repair itself; and (2) repair carries a real, statistically robust, quantified reoperation penalty (12 percentage points worse freedom-from-reoperation at 12 years). A recommendation that repair should happen 'whenever possible' with no caveat is not compatible with a trade-off where the survival case is confounding-fragile and the reoperation cost is confirmed.
The innocent explanation, and whether it was checkedPartially dissolves: repair candidacy in kumar-2020/marijon-2012 is inherently about anatomically favourable valves in experienced centres, so the same selection that inflates the raw survival numbers in fu-2021's unmatched studies is also what real-world surgeons are doing on purpose -- picking repairable valves for repair. That does not fully dissolve the conflict, though, because neither marijon-2012 nor kumar-2020 (as captured in this corpus) mentions the reoperation trade-off at all, and 'should be undertaken whenever possible' is stated as if there were no downside to weigh against the survival case. Checked: searched both papers' full sections in the claims file for any mention of reoperation risk alongside the repair recommendation -- none found.
What depends on which is rightThis determines how strongly a surgeon in a resource-constrained programme should push for repair over replacement, especially since reoperation in these settings is disproportionately dangerous (South African data elsewhere in this same shelf show redo mitral surgery presents as an emergency in 73% of cases). If the reoperation downside is real and the survival upside is confounded, the 'repair whenever possible' rule needs a caveat about where redo capacity is limited; if the guidance is right as stated, patients are being needlessly steered toward replacement in some programmes.
domain: surgery, valve intervention and anticoagulation
Whether task-shared (non-physician-led) programmes for RHD tertiary prevention/care in LMICs have any published evidence at all.low
One sideA systematic review of all study designs on task sharing for GAS/ARF/RHD in limited-resource settings found zero eligible studies, and specifically no studies addressing tertiary prevention.
0 of 18 candidate full-text articles met inclusion criteria; 0 tertiary-prevention studies identified · LMIC/limited-resource settings, all study designs, searched PubMed, Cochrane CENTRAL, Embase, Scopus, Web of Science, WHOLIS, Africa Wide, CINAHL, English-language, no date restriction · systematic review, published 2019
abdullahi-2019-task-sharing-diagnosis
The other sideDecentralised, nurse-led follow-up of patients with advanced RHD after cardiac surgery at rural Rwandan district hospitals was feasible and produced generally good outcomes, including 92.6% survival, 96% INR-monitoring adherence at visits, and 100% of alive patients still on penicillin secondary prophylaxis at last visit.
54 patients, 92.6% alive at end of follow-up, 96% of visits had INR checked, 83%->4% NYHA III/IV before vs after surgery · 54 RHD patients post cardiac valve surgery, 2007-2015, followed median 3 years (range 0.2-7.9) · 3 rural PIH/IMB-supported district hospital NCD clinics in Rwanda (Butaro, Kirehe, Rwinkwavu), retrospective chart review, published 2018
rusingiza-2018-outcomes-patients-rheumatic
Why they cannot both be rightAbdullahi-2019's review explicitly states it found no studies addressing task sharing for tertiary prevention of RHD anywhere in the searched literature. Rusingiza-2018 is exactly this: a published, outcome-reporting study of nurse-led (task-shared, non-physician) follow-up of post-surgical RHD patients -- the textbook definition of tertiary-prevention task sharing -- with a full year before the review's publication date.
The innocent explanation, and whether it was checkedYes, most likely, though not fully confirmed. Two plausible innocent mechanisms: (1) search-cutoff timing -- systematic reviews typically finalise their search several months to a year before publication, so a review published in 2019 may have run its search in 2017 or early 2018, before Rusingiza-2018 was indexed; (2) vocabulary mismatch -- Abdullahi's review searched on 'task sharing/task shifting' terminology, while Rusingiza-2018's title and likely its indexing terms use 'decentralised, nurse-led follow-up' language instead, which a keyword-based systematic-review search could plausibly miss even if it describes the identical intervention model. I could not confirm either mechanism directly (the CLAIMS-FLAT extracts don't give Abdullahi's exact search date or Rusingiza's indexing terms), so this is the likely but unverified explanation, and I am downgrading confidence accordingly.
What depends on which is rightIf 'no evidence exists for task-shared tertiary RHD care' becomes the standing citation (as Abdullahi-2019's conclusion is positioned to be), funders and guideline writers may treat nurse-led post-surgical follow-up models as unproven and deprioritise scaling them, when at least one published, favourable outcomes study existed at the time. It also flags a structural risk in this literature: keyword-based systematic reviews on 'task sharing' may systematically undercount task-shared care models that are described using different operational language (decentralisation, nurse-led clinics, NCD-integrated follow-up).
domain: control programmes, registers and economics

What happened when we tried to kill them

Each of the top 12 was handed to a separate reader whose only job was to find the paper that already did the work. They searched live literature and this shelf, and were told to default to ‘refuted’ when unsure.

9refuted outright
3weakened
0survived intact
212searches run
The honest headline: nothing survived intact. Mining papers for their stated limitations produces candidates cheaply and most of them are false, because a limitations section is a paper being modest about itself — not a report on the field. Several times the answer was inside the very papers claiming it was missing. What is left below is the narrow residue each reader could not kill, which is where anything real would be.
At what age a GAS vaccine should be givenWEAKENED
VerdictThe stated obstacle ("nobody ever built the model") is false — the same author group built it at least three times, comparing infant vs age-5 GAS vaccination on cost, cost-effectiveness and DALYs/deaths averted; what actually survives is a narrower causal question those models explicitly admit they didn't solve, and it can't be closed by more modelling alone.
Already done byGiannini F, Cannon JW, Cadarette D, Bloom DE, Moore HC, Carapetis J, Abbas K. Modeling the potential health impact of prospective Strep A vaccines. npj Vaccines. 2023.
Static cohort model, 183-205 countries, explicitly runs vaccination-at-birth vs vaccination-at-5-years as two arms and reports cases/deaths/DALYs averted separately by age of administration, across 6 coverage/waning scenarios. On our own shelf (wiki/evidence/_drafts/giannini-2023-modeling-potential-health.md).
Already done byLee/Cannon et al. The potential global cost-effectiveness of prospective Strep A vaccines and associated implementation efforts. npj Vaccines. 2023.
Full global cost-effectiveness analysis with maximum-cost-per-fully-vaccinated-person thresholds computed separately for routine vaccination at birth vs at 5 years, across HIC/UMIC/LMIC/LIC and per disease endpoint (pharyngitis, RHD, invasive, impetigo, cellulitis) — the exact 'cost and cost-effectiveness by age' quantification the candidate claims doesn't exist. On our own shelf as wiki/evidence/_drafts/lee-2023-potential-global-cost.md, 19 extracted claims.
Already done byCadarette D et al. The full health, economic, and social benefits of prospective Strep A vaccination.
Models aggregate lifetime economic benefit and break-even cost per vaccinated person separately for infant vs age-5 administration (e.g. $2.3T vs $3.8T baseline benefit over 30 cohorts; break-even <$1,000 LMIC to ~$4,300 UMIC for age-5 dosing), across a battery of discount-rate/VSLY sensitivity scenarios.
Already done byHealth-Economic Value of Vaccination Against Group A Streptococcus in the United States. Clin Infect Dis. 2022.
US-specific model comparing costs averted for a 12/18-month schedule, a 12/18-month-plus-age-5-booster schedule, and adult (age 65) vaccination separately, i.e. an explicit age-schedule comparison with dollar figures attached.
What it could not killWhat proportion of RHD incidence detected at age 5+ is causally attributable to GAS infections occurring in the first years of life (and therefore preventable by blocking that antecedent infection with an infant-timed vaccine) is not modeled by any published Strep A impact/cost-effectiveness paper. Every model found (Giannini 2022/2023, Lee/Cannon 2023, Cadarette 2023, the 2022 US CID paper) uses a static cohort design that assumes vaccination directly prevents incident RHD in the vaccinated age cohort — a simplifying surrogate the Giannini papers explicitly flag by name as likely overestimating the impact of 5-year vaccination and underestimating infant vaccination, precisely because the infection-to-RHD etiological lag isn't represented. Closing this needs an age-structured transmission/natural-history model linking early-life GAS exposure (the kind of data Keeley et al. 2025's Gambian SpyCATS cohort is starting to generate) to downstream RHD risk — no such model has been published. On top of that, the true optimal age also needs a real vaccine's measured efficacy-by-age and duration of protection, which cannot exist before a correlate of protection and Phase 2/3 trial data exist — an upstream dependency, not a modelling gap.
The only version still worth doingAn age-structured transmission/natural-history model that couples early-life GAS acquisition and reinfection risk (calibrated on the Gambian SpyCATS carriage and serology cohort underlying Keeley et al. 2025) to downstream ARF/RHD risk at age 5+, so that the RHD burden at age 5 attributable to infections occurring in years 0-5 can be estimated counterfactually and compared against a birth-dose vaccination scenario. This is the one component none of the four existing economic/impact models attempt — they are all static, non-dynamic, direct-effect cohort models by explicit design. Even this would still be a projection pending real trial data on vaccine efficacy and duration of protection by age, so its output would be a bound on the answer, not a settled optimal age — the field cannot fully close this gap until a licensed candidate and a correlate of protection exist.
Objection beyond coverageThe obstacle as written ('nobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantified') is not thinly covered, it is factually wrong, and wrong in a specific way: it appears to have been extracted from the limitations paragraph of one paper (Giannini 2022/2023) without checking whether that same limitation had already been addressed elsewhere in the same small author cluster's own body of work — Lee/Cannon 2023 and Cadarette 2023, both companion papers from the SAVAC-affiliated group, ARE the cost/cost-effectiveness-by-age model the obstacle claims is missing, and one of them (giannini-2023) is even listed as one of the candidate's own three source papers, i.e. the gap-candidate cites, as evidence of a missing model, a paper that itself IS the model. Separately, the three example gap sentences quietly conflate two different kinds of unknown: 'requires evaluation in modelling' (done, repeatedly, with numbers) and 'requires evaluation in vaccine trials' (genuinely not done, but for a reason that has nothing to do with anyone failing to build a model — no licensed candidate and no correlate of protection exist yet). Bundling an answered question with an unanswerable-until-upstream-work-is-done question into one 'gap' makes the whole thing look more open than it is.
21 searches run against this one.
An affordable point-of-care test for GAS pharyngitisWEAKENED
VerdictThe exact combination (molecular/RADT test x LMIC unit costs x Markov CEA) hasn't been published, but every piece it would require already has been — an LMIC cost-effectiveness model for GAS diagnosis-and-treat strategies exists (Irlam 2013, South Africa), a full systematic review + economic model of POC GAS tests exists (Fraser 2020, UK HTA), and the candidate's own cited paper (Armitage 2025) already supplied LMIC accuracy data showing the tests currently perform too poorly to be worth costing out — and the real reason WHO tells endemic countries to skip testing is stated explicitly as poor test ACCESS, not missing analysis.
Already done byIrlam JH, Mayosi BM, Engel ME, Gaziano TA. A cost-effective strategy for primary prevention of acute rheumatic fever and rheumatic heart disease in children with pharyngitis. S Afr Med J. 2013;104(2):125-130 (also indexed as PLoS ONE 2013 Markov model paper)
Markov cost-effectiveness model in urban South African primary care (children 3-15) comparing treat-all, culture, and clinical-decision-rule (CDR 2+/3+) strategies for pharyngitis, with ICER/QALY outputs against a WHO cost-effectiveness threshold, plus sensitivity analysis across GAS prevalence 1.6-30%. Already sits on our own shelf (wiki/evidence/CLAIMS-FLAT-rhd.md, section 'irlam-2013-primary-prevention-acute').
Already done byFraser H, Gallacher D, Achana F, Court R, Taylor-Phillips S, Nduka C, Stinton C, Willans R, Gill P, Mistry H. Rapid antigen detection and molecular tests for group A streptococcal infections for acute sore throat: systematic reviews and economic evaluation. Health Technol Assess. 2020;24(31)
Full NIHR HTA: systematic review of accuracy for 21 POC GAS tests plus a de novo decision-tree economic model comparing POC testing + CDRs against CDRs alone, concluding none of the 14 modelled tests were cost-effective at NICE thresholds. Proves the modelling apparatus the candidate says is missing (systematic-review-plus-decision-tree for POC GAS tests) already exists and is a straight parameter swap, not new science, to re-run with LMIC unit costs and prevalence.
Already done byArmitage EP et al. Evaluating clinical decision rules and rapid diagnostic tests for the diagnosis of Streptococcus pyogenes pharyngitis in Gambian children: A diagnostic accuracy study, 2025 (one of the two papers the gap candidate itself cites)
Prospective diagnostic-accuracy study of a molecular POC test (ID NOW), a lateral-flow antigen test, and five clinical decision rules in 376 Gambian children with pharyngitis — i.e. exactly the LMIC/RHD-relevant population and exactly the test types the gap names. Findings: ID NOW PPV only 28.7% against culture; the antigen LFT had sensitivity 55.7%/specificity 80% against PCR; authors conclude none of the evaluated tests or CDRs 'appears suitable for adoption in this setting.' This supplies the accuracy inputs a cost-effectiveness model needs and shows the likely answer is unfavourable on clinical grounds before cost is even added.
Already done byWHO 2024 RHD guidelines, as summarised in ali-2026-implementing-world-health (wiki draft on our shelf)
States plainly that WHO recommends treating suspected GAS pharyngitis on clinical grounds alone in highly endemic countries 'because access to rapid antigen tests and culture is poor' — i.e. the global guideline body already names the barrier as procurement/access, not as absence of a cost-effectiveness analysis.
What it could not killA single published economic model that (a) uses molecular or rapid-antigen POC tests specifically (not just culture/CDR, which Irlam 2013 already modelled) as the tested arm, (b) is parameterised with genuine LMIC/RHD-endemic unit costs, prevalence and care-seeking behaviour, and (c) reports NNT/NNH and AMR-scale-up implications, has not surfaced in this search. That precise triple intersection is real and narrow. What does not survive is the framing that this is scientifically unknown territory: the Markov/decision-tree structure (Irlam), the accuracy inputs from an LMIC population (Armitage, the candidate's own cite), and the POC-test economic-modelling method (Fraser) all already exist and would need to be assembled, not invented. The skin-sore/impetigo half of the gap is weaker still: there is no widely validated POC RDT for GAS skin infection to cost out in the first place, so 'feasibility of POC RDTs to enable directed treatment' for skin sores is arguably not yet a research question — it is a missing product.
The only version still worth doingReframe it as a threshold/price-target analysis rather than a generic cost-effectiveness study: take Armitage 2025's Gambian (or another current LMIC) accuracy data for a molecular POC test, plug it into an Irlam-style Markov/decision-tree structure with local BPG and complication costs and local ARF/RHD incidence, and solve for the maximum per-test price at which POC testing beats treat-all or CDR-based care. That number is directly usable by donors and ministries negotiating with manufacturers (analogous to what lee-2023-potential-global-cost already did for a Strep A vaccine), which is the one output from this exercise that could plausibly change a procurement decision rather than just adding a citation.
Objection beyond coverageThe candidate frames this as an analytic gap ('the model was never built') but the source it leans on for policy relevance is a 2024 WHO guideline that already made the decision — treat on clinical suspicion in endemic settings — and named the reason as poor test ACCESS, not missing cost-effectiveness data. A cost-effectiveness paper does not put a $2 molecular cartridge into a rural clinic in Chad or the Gambia; supply chains, manufacturer pricing and procurement volume do. Even if the missing CEA were built tomorrow with a favourable ICER, WHO's practical recommendation would not change until the test is actually purchasable at scale in that setting, because the guideline's stated constraint is availability, not knowledge. That makes this substantially a market/procurement problem wearing a research costume: the paper that 'closes' it would add a number to cite, not change what a health worker in an endemic district is handed today. The one place a CEA genuinely would move a decision is upstream, as a price target for donors/manufacturers ('a POC test needs to cost under $X to be worth subsidising') — but the candidate's gap language ('the cost-effectiveness... has not been established') doesn't ask that sharper, decision-relevant question; it asks for the generic analysis, which is the kind of gap that survives searches mainly because nobody has bothered to assemble existing pieces, not because the pieces don't exist.
21 searches run against this one.
School-based sampling versus true community prevalenceWEAKENED
VerdictSomeone already ran exactly this test at meta-analytic scale in 2014 and found no significant difference between school-based and community-based RHD prevalence surveys, so the premise that this bias is unexamined and unquantified is wrong even though the specific test was underpowered.
Already done byRothenbühler M, O'Sullivan CJ, Stortecky S, Stefanini GG, Spitzer E, Estill J, Shrestha NR, Keiser O, Jüni P, Pilgrim T. Active surveillance for rheumatic heart disease in endemic regions: a systematic review and meta-analysis of prevalence among children and adolescents. Lancet Glob Health. 2014.
Pooled 37 study populations (25 for the main prevalence estimate) of active RHD surveillance in children/adolescents across endemic regions worldwide, and ran a pre-specified sensitivity analysis testing surveillance type (school-based, 34 populations, vs community-based, 3 populations) as a moderator of pooled prevalence. Found no significant interaction (p=0.200). This is already sitting in our own shelf at wiki/evidence/_drafts/rothenbuhler-2014-active-surveillance-rheumatic.md, claim #8.
Already done byPandey (2025) — the very paper the candidate cites as evidence of the Nepal gap.
The candidate's own source gap file states verbatim: 'Existing literature on RHD in Nepal has not adequately synthesized the evidence to provide a robust national prevalence estimate, nor has it systematically quantified urban-semi urban disparities...' — i.e. this is the PRIOR gap the paper names, and the paper itself is the review that did the quantifying. The Nepal instantiation of the candidate's gap was closed in the same paper that is cited as evidence it exists.
What it could not killA narrow, individual-level version: within one specific high-burden country, has anyone drawn a probability sample stratified by school-enrollment status and directly compared echocardiographic RHD prevalence in enrolled vs non-enrolled children of the same age band and district, adequately powered to detect a policy-relevant difference? I did not find one, and I looked specifically for that phrasing (school-going vs non-school-going, individual-level stratification, enrollment status). That precise design is genuinely absent from what I could retrieve.
The only version still worth doingIf someone still wanted to fund this, the only version worth running is a single-country, adequately-powered cross-sectional survey using a two-stage design — draw a random sample of villages/wards, then within each cluster enumerate ALL children by house-to-house listing (capturing both enrolled and non-enrolled), and echo-screen a random subsample from each enrollment stratum, powered a priori to detect a plausible effect size (e.g. a 1.5-2x prevalence ratio) between enrolled and non-enrolled children — not another systematic review or re-analysis of existing pooled study-level data, which is what both cited 2025 papers already are and what produced the null Rothenbühler result. Anything short of new primary household-level data collection just re-runs the same underpowered study-level comparison that already came back null.
Objection beyond coverageThe candidate treats 'never quantified' as equivalent to 'unknown and answering it would change a decision,' but it conflates two different things. First, direction is not obviously the one the candidate assumes: Rothenbühler et al.'s own meta-analysis (which we already hold) shows RHD prevalence tracks the Gini coefficient (a 0.1 rise in Gini predicts a 1.4x rise in prevalence) — so if out-of-school children are poorer, the bias direction plausibly makes school-based estimates an UNDERESTIMATE of true community burden, which is directionally intuitive and already assumable without a new study. Second, and more importantly, this exact comparison — school-based vs community-based surveillance type as a predictor of pooled RHD prevalence — was already run as a sensitivity analysis on 37 pooled study populations and came back null (p=0.200). A health ministry deciding where to put prevention money does not need a bias-magnitude estimate to act: it already has cluster/population-based prevalence figures (used routinely, e.g. the Africa-wide population-based pooled estimate of 18.4/1000 from 22 studies) it can and does use instead of school-only figures when it wants a whole-population number. The candidate's 'obstacle' framing — 'the datasets already exist but were never joined' — is also factually wrong for the Nepal half of its own evidence base: the cited Pandey (2025) paper IS the joining exercise, done by the same authors the candidate cites as proof the joining hadn't happened.
32 searches run against this one.
Why RHD affects women roughly twice as often as menREFUTED
VerdictA biological mechanism for the female excess in RHD has already been published in Circulation (Passos et al. 2022): an estrogen-receptor-alpha-linked, prothymosin-alpha-driven CD8+ T-cell autoimmune pathway against valve collagen, and it is even cited inside two of the candidate's own seven source papers (mutarelli-2025, lupieri-2025), so the gap sentence claiming this 'remains unexplored' is contradicted by the candidate's own reading list.
Already done byPassos LSA, Jha PK, Becker-Greene D, Blaser MC, Romero D, Lupieri A, Sukhova GK, Libby P, Singh SA, Dutra WO, Aikawa M, Levine RA, Nunes MCP, Aikawa E. Prothymosin Alpha: A Novel Contributor to Estradiol Receptor Alpha-Mediated CD8+ T-Cell Pathogenic Responses and Recognition of Type 1 Collagen in Rheumatic Heart Valve Disease. Circulation. 2022.
Proteomic screen of 30 human heart valves identified prothymosin-alpha (ProTa) as enriched in rheumatic valves; ProTa expression correlates with estrogen receptor alpha (ERa) in circulating CD8+ T cells, ProTa induces perforin/granzyme B cytotoxicity and raises ERa expression, and an ERa antagonist blocks the cytotoxic effect. In silico molecular mimicry between a human type-1-collagen epitope and a Streptococcus pyogenes collagen-like protein was also identified. This is precisely the 'estrogen-regulated pathway' gap sentence 5 says is unexplored -- it is not only explored, it is published with mechanistic in vitro validation in a flagship cardiology journal, and the authors state its purpose explicitly: 'implicating ProTa as a potential regulator of sex predisposition in RHVD.'
Already done byMutarelli A, et al. Global prevalence and sex differences in rheumatic heart disease: a systematic review and updated meta-analysis. 2025. (one of the candidate's own 7 source papers; draft on our shelf: wiki/evidence/_drafts/mutarelli-2025-global-prevalence-sex.md)
Pools 58 echo-screening studies (215,552 subjects) and confirms the female excess quantitatively (latent RHD RR for males vs females 0.70, 95% CI 0.61-0.80; definite RHD RR 0.71, 95% CI 0.59-0.86), then in its own discussion names two candidate mechanisms for the sex gap: women's greater caregiving-driven exposure to group A strep, and the ProTa/estrogen-receptor-alpha autoimmune pathway from Passos 2022. The paper the candidate cites as evidence of the open gap is itself citing a proposed mechanism for it.
Already done byLupieri A, et al. Rheumatic Heart Valve Disease pathophysiology and mechanisms (review). 2025. (one of the candidate's own 7 source papers; draft on our shelf: wiki/evidence/_drafts/lupieri-2025-rheumatic-heart-valve.md, claim #10)
Review by an overlapping author group (shares Levine RA, Nunes MCP, Aikawa E with Passos 2022) states as claim #10 that ProTa is elevated in rheumatic mitral valves, promotes CD8+ cytotoxicity and VLA-2-mediated collagen recognition, 'with estrogen receptor-alpha activity amplifying this effect,' annotated by our own extraction as 'potentially explaining female predisposition.' A paper the candidate lists as one of the seven building blocks of the gap already synthesizes the mechanistic answer.
Already done byFairweather D, Petri MA, Coronado MJ, Cooper LT. Autoimmune heart disease: role of sex hormones and autoantibodies in disease pathogenesis. Expert Rev Clin Immunol. 2012.
Establishes the general framework the candidate's gap sentence 5 gestures at (sex-hormone-driven autoimmunity in cardiac disease) a decade before the candidate's example sentences were drafted -- shows the mechanistic research programme into sex differences and autoimmune heart disease was already active well before 2022, not a fresh unexplored territory.
What it could not killTwo narrower pieces are genuinely thin. First, nobody has run a study that directly pits the socio-cultural (health-seeking-behavior, differential care access) hypothesis against the biological (estrogen/autoimmune) hypothesis in the same population with a design that could actually apportion variance between them -- every search for that specific comparative design (health-seeking-behavior gender diagnosis delay, streptococcal exposure by household caregiving role) returned zero hits, so the 'socio-cultural versus biological, unresolved' framing in gap sentence 2 is the one piece of the six gap sentences that survives as stated, though it survives as an underpowered/never-attempted comparative design question, not as a total absence of proposed mechanisms. Second, the precise age at which the sex divergence first becomes measurable (gap sentence 6) was not directly located in any search and is plausible as a genuinely unaddressed descriptive-epidemiology question, though it is a minor, low-stakes datum next to the mechanism question, and it would not on its own justify screening pregnant women for RHD.
The only version still worth doingIf anything is worth doing here, it is not a bench mechanism study (that lane is already occupied by an active, well-funded, high-impact-journal research programme -- Aikawa/Levine/Nunes/Dutra and collaborators, three papers 2021-2025 and counting). The one design that would add real information is a prospective, mixed biological-and-behavioral cohort in a single high-burden population (e.g. a site already running RHD registries in Uganda or Fiji) that collects, in the same children, (a) confirmed GAS exposure/household caregiving-role data, (b) validated health-seeking/access-to-care measures, and (c) a biological panel (autoantibody titers, ProTa/ERa where feasible) at the point of first echo screening -- allowing an actual apportionment between exposure, access and biology, rather than another prevalence meta-analysis that just re-confirms the ratio. Absent that specific comparative design, another descriptive prevalence-by-sex paper (which is what six of the candidate's seven source papers already are) adds nothing.
Objection beyond coverageThe obstacle statement bundles a specific, already-published biological mechanism (estrogen/ProTa/CD8+ autoimmunity) together with a genuinely unanswered comparative-design question (how much of the effect is behavioral vs biological) and treats the whole bundle as one untouched gap. That is doing real work for the candidate: as soon as you split it, most of the load-bearing claim collapses. It is also not clear the remaining sliver would change the decision it is attached to. The stated decision is 'whether to screen pregnant women for RHD and how to manage the ones found' -- but that decision does not actually depend on knowing *why* women get RHD more often. It depends on knowing that they do (settled, RR ~0.7-1.4x depending on definition, replicated across dozens of studies) and on maternal/fetal outcome and screening-yield data, which is itself a mature, separate literature (Beaton 2019 Heart, Khanna 2021, Yang 2024 JACC Adv scoping review, multiple national pregnancy-cardiac registries). A ministry of health does not need a resolved etiological mechanism to decide to screen pregnant women in an endemic region; WHO and multiple national programmes already screen and manage RHD in pregnancy today using risk-classification tools (modified WHO classification) built without knowing the mechanism. So even in its strongest surviving form, this gap is not a blocker on the cited decision -- it is a mechanism-curiosity question wearing a policy-relevance costume.
15 searches run against this one.
Cost-effectiveness and sustainability of echo screening programmesREFUTED
VerdictThe claim that nobody has modelled the cost-effectiveness of RHD echo screening is false: at least seven published cost-effectiveness/cost-utility models exist since 2013 across Australia, Brazil, India and Rwanda, one of them (Ubels/Nascimento 2020, Brazil PROVAR+) explicitly costing handheld devices with task-shifted, telemedicine-read screening, and Nepal has already run a cluster-randomized school-based echo screening trial.
Already done byRoberts K, Cannon J, Atkinson D, et al. Echocardiographic Screening for Rheumatic Heart Disease in Indigenous Australian Children: A Cost-Utility Analysis. J Am Heart Assoc. 2017.
Built a multistate cost-utility model of RHD progression from Australian register data and ran it against two screening strategies, producing an ICER (AU$/DALY averted) and identifying which age/frequency strategy is cost-effective. This is exactly 'the model' the candidate claims does not exist, published nine years ago.
Already done byUbels J, Sable C, Beaton AZ, et al. Cost-Effectiveness of Rheumatic Heart Disease Echocardiographic Screening in Brazil: Data from the PROVAR+ Study. Glob Heart. 2020.
A 13-state Markov cost-utility model built on primary large-scale screening data, specifically costing handheld-device screening performed by non-physicians with telemedicine remote reading (i.e. the exact WHO-recommended handheld/task-shifted model the candidate says is unevaluated). ICER $10,148/DALY, cost-effective in 70% of probabilistic iterations.
Already done byZachariah JP, Samnaliev M. Echo-based screening of rheumatic heart disease in children: a cost-effectiveness Markov model. J Med Econ. 2015.
Markov model combining two-stage echo screening with secondary penicillin prophylaxis, exactly the pairing named in gap sentence 3 ('cost-effectiveness of echocardiographic screening and penicillin prophylaxis has not been established'); found screening dominates on cost in >80% of simulations.
Already done byManji RA, Witt J, Tappia PS, et al. Cost-effectiveness analysis of rheumatic heart disease prevention strategies. Expert Rev Pharmacoecon Outcomes Res. 2013.
Earliest of the batch: Markov model comparing primary prophylaxis, universal antibiotic prophylaxis and echo-screening-plus-secondary-prophylaxis; concludes screening-plus-secondary-prophylaxis is the best strategy. Directly answers gap sentence 3, twelve years before the candidate's example sentences were drafted.
What it could not killOne narrow, genuinely open piece: empirical (not modelled) multi-region comparative cost and long-term sustainability data for handheld-device, task-shifted active case finding, collected prospectively across several different high-prevalence settings rather than as country-specific Markov simulations built on borrowed parameters. Providencia 2024's own systematic review calls this out explicitly, and it is accurate as far as it goes -- the six CEA papers above are each single-country models, not a pooled empirical multi-country costing exercise, and the one study designed to produce real (not simulated) sustainability and cost data (NEARER SCAN/LENO BESIK) had not published results as of my search and covers only Australia and Timor-Leste. That is a real, specific residual gap -- but it is a refinement of a well-populated literature, not the empty field the candidate's obstacle statement describes.
The only version still worth doingA prospective, pooled multi-country costing add-on nested inside the running task-sharing implementation trials (NEARER SCAN/LENO BESIK plus a parallel arm in a low-resource, non-Pacific setting such as Nepal or a sub-Saharan African site) that collects real per-child screening cost, task-shifted labour cost, device amortization and 3-5 year programme continuation/fidelity data, then re-fits the existing Markov models' assumed parameters against measured ones. That would meaningfully upgrade six assumption-driven models into one empirically anchored costing tool -- but it is a validation and pooling study on an already-substantial literature, not a first-of-its-kind gap-filling study, and should be pitched that way rather than as 'the first model.'
Objection beyond coverageThe obstacle as written -- 'nobody ever built the model' -- is a category error dressed as a gap. It quietly redefines 'unquantified' to mean 'not quantified for every device generation, every country and every age band simultaneously,' which is a standard applies to no health-economics question ever and can never be fully satisfied; by that bar cost-effectiveness of literally any global-health intervention is 'unquantified.' The actual decision the candidate says is blocked (should a ministry screen, whom, what age, what device, who reads it) has already been answered operationally in several jurisdictions using exactly this evidence base: Australia's RHD Control Programs are running and measured (Stacey 2023), Brazil scaled PROVAR+'s task-shifted handheld model on the back of its own cost-utility analysis, and WHO's 2023 recommendation to use handheld echo was itself informed by this literature (the candidate's own source paper, Rwebembera 2023, is the WHF document that synthesizes it). A gap statement that is contradicted by the policy action its own source paper already recommends is not describing an open decision -- it is restating a call for more precision after the coarse decision was already made and acted on.
12 searches run against this one.
The economic value case for a GAS vaccineREFUTED
VerdictThe candidate's own four citations are already a full health-economic model of a Strep A vaccine (trillion-dollar benefit ranges, income-stratified cost-effectiveness thresholds, DALYs and deaths averted by age of vaccination), so "nobody ever built the model" is false on its face.
Already done byCadarette D, Boyer M, Duffy J, et al. 2023. The full health, economic, and social benefits of prospective Strep A vaccination. NPJ Vaccines / health-economic modeling report.
Builds a global CEA-style benefit model across 30 birth cohorts, 6 scenarios, income-stratified break-even costs (e.g. <$1,000 per vaccinated person in LMICs at age 5), and explicitly quantifies AMR-linked benefits: ~800 deaths and 69,000 DALYs averted from AMR (citing Kim et al. 2023) and a 32%/7% cut in antibiotic prescriptions for Strep A pharyngitis (citing Miller et al. 2023). This directly falsifies the gap sentence claiming AMR impact 'has not been quantified.'
Already done byGiannini F, et al. 2022. The potential health impact of prospective Strep A vaccines: a modeling study.
Static cohort transmission-adjacent burden model across 183-205 countries quantifying cases, deaths and DALYs averted by age of vaccination (birth vs. 5 years) and income group; the paper itself flags, in its own limitations, that it excludes herd effects and is therefore conservative -- this is the one place a real residual gap sits.
Already done byGiannini F, et al. 2023. Modeling the potential health impact of prospective Strep A vaccines.
Companion/expanded version of the 2022 model; same structure, same explicit self-flagged limitation on herd effects, ARF/APSGN exclusion, and etiological-pathway timing.
Already done byLee KY, et al. 2023. The potential global cost-effectiveness of prospective Strep A vaccines and associated implementation efforts.
Full four-income-group cost-effectiveness analysis with disease-specific threshold prices (pharyngitis, RHD, invasive, impetigo, cellulitis), univariate and multivariate sensitivity analysis on discounting and economic-burden assumptions, and an explicit cite to a prior systematic review (Lee et al. 2021) mapping the existing CEA literature landscape and its gaps -- i.e. this field already has its own gap-mapping paper.
What it could not killA formal cost-effectiveness model that endogenizes GAS transmission dynamics (herd/indirect effects) inside the economic evaluation, rather than assuming direct-effect-only and flagging the omission as a conservative-bias caveat, has not been published. The epidemiological piece (a GAS-specific transmission-dynamic model with a vaccine-impact analysis, Chisholm et al. 2020) and the economic piece (Cadarette/Giannini/Lee) both exist separately but have never been merged the way pneumococcal and meningococcal B vaccine CEAs now routinely do (e.g. the Dutch MenB DyCE model). That merge is a real, specific, still-open methods gap -- but it is a refinement of an existing decision-ready model stack, not the absence of one.
The only version still worth doingThe one study worth doing is narrow and methodological, not foundational: extend a GAS-specific transmission-dynamic model (building on Chisholm et al. 2020's agent-based structure, or a compartmental analogue) to output age- and setting-specific attack-rate reductions under vaccination, then feed those into the existing Lee 2023 / Cadarette 2023 cost-effectiveness framework in place of the direct-effect-only assumption, for one or two high-burden LMIC settings where herd effects are plausibly largest. This would quantify how much the giannini/cadarette estimates undersell the vaccine (both papers already say 'likely conservative' but give no number), which matters for age-of-vaccination and dose-schedule decisions specifically, not for the license/fund go/no-go decision, which the existing threshold tables already answer under a wide assumption range.
Objection beyond coverageThe 'obstacle' sentence claims the model doesn't exist, but the candidate's own paper list of 4 is that model, published and cited together as a coherent 2022-2023 body of work (the Cadarette/Giannini/Lee cluster came out of the same WHO Strep A vaccine roadmap effort and cross-cite each other, including a 2021 systematic review of the existing CEA literature). A gap statement that is refuted by its own evidence list is not a gap in the literature, it is a gap in how the sentence was written -- most of the individual 'gaps' listed (no empirical LMIC VSLY, no Strep-A-specific WTP study, extending life vs improving quality-of-life-value equivalence untested) are not Strep-A-specific limitations at all: they are the standard, field-wide caveats stapled onto essentially every LMIC vaccine CEA (typhoid, rotavirus, malaria) because empirical VSLY and disease-specific WTP studies barely exist anywhere in global health economics. Naming them as a Strep A gap implies a Strep-A-specific fix is needed and possible; it isn't -- fixing it would require a field-wide advance in stated-preference methods, not a Strep A study. That makes 'the economic value case' a poor unit for a research gap: it bundles one real narrow methods gap (transmission-dynamics-integrated CEA) with several field-wide, non-fixable-by-this-study caveats, and the decision at stake (license/fund/schedule, which antigen/age) is already answerable from the existing cost-effectiveness threshold tables across a wide range of discount rates and VSLY assumptions -- Lee 2023's own sensitivity analysis shows the qualitative conclusion (favorable in LMIC/LIC at very low per-dose cost, viable in HIC/UMIC at higher cost) is robust to the very assumptions the gap sentences complain about.
13 searches run against this one.
Economic evaluation of prevention strategies other than vaccinesREFUTED
VerdictThe model has been built at least five times since 2002, including a 55-country African investment case (2021) and an equity-stratified India analysis (2023) that does exactly what the candidate says nobody has done.
Already done byCoates MM, Sliwa K, Watkins DA, et al. An investment case for the prevention and management of rheumatic heart disease in the African Union 2021-30: a modelling study. Lancet Glob Health. 2021.
Cohort state-transition model across all 55 African Union member states comparing primary prevention alone, secondary+tertiary care alone, and the full bundle. Reports benefit-cost ratios for every combination (0.2 for primary alone, 4.7 for secondary+tertiary, rising to 8.4 at longer time horizon) and net benefits in 2019 USD. This is a multi-country combination comparison, which is exactly gap sentence 5 in the candidate ('no study has yet compared combinations of primary, secondary, and tertiary strategies').
Already done byDixit J, Prinja S, Jyani G, et al. Evaluating efficiency and equity of prevention and control strategies for rheumatic fever and rheumatic heart disease in India: an extended cost-effectiveness analysis. Lancet Glob Health. 2023.
Markov model for India with an explicit extended cost-effectiveness analysis (ECEA) stratifying costs, QALYs, out-of-pocket expenditure and catastrophic/impoverishing health expenditure by wealth quartile, for every combination of primary/secondary/tertiary prevention. This is literally gap sentence 1 ('no previous study has evaluated the distributional effect of primary, secondary, or tertiary interventions and their combinations') -- closed by the same paper the candidate cites as evidence the gap exists. It is one of the candidate's own three source papers.
Already done byManji RA, Witt J, Tappia PS, Jung Y, Menkis AH, Ramjiawan B. Cost-effectiveness analysis of rheumatic heart disease prevention strategies. Expert Rev Pharmacoecon Outcomes Res. 2013.
Markov model comparing three prevention strategies (primary throat-swab-directed prophylaxis, primary universal benzathine penicillin, secondary prophylaxis) against no prevention, in a developing-world setting. Predates the candidate's obstacle claim by a decade.
Already done bySoudarssanane MB, Karthigeyan M, Mahalakshmy T, et al. Rheumatic fever and rheumatic heart disease: primary prevention is the cost effective option. Indian J Community Med. 2007.
Cost accounting comparing economic output/input ratios of primary (1:1.56) vs secondary (1:1.07) vs tertiary (1:0.12) prevention in Pondicherry, India -- a direct head-to-head economic comparison of the three prevention layers, 16 years before the candidate's obstacle claim.
What it could not killTwo narrow pieces. First, LAC-region economic evaluations specifically remain scarce (jaimesreyes-2022 found only 3), so a Coates/Dixit-style model has not been built for that region. Second, Dixit-2023's own stated limitation -- that intervention coverage was assumed equal across socioeconomic subgroups rather than empirically measured -- has not been resolved by any study found. Neither of these is 'nobody ever built the model'; both are 'extend or refine an existing, working model design.'
The only version still worth doingApply the existing Coates-2021 / Dixit-2023 Markov/state-transition modelling framework to a country or region that currently lacks one (LAC is the one region jaimesreyes-2022 flags as thin) using local epidemiological and cost data, and explicitly test rather than assume whether intervention coverage is equal across socioeconomic subgroups -- Dixit-2023 names this as its own unresolved assumption. That is a worthwhile replication-with-a-real-fix study. A new general-purpose economic-evaluation methodology paper is not worth doing; the methodology already exists and is published.
Objection beyond coverageThe candidate's two strongest gap sentences (1 and 5) are lifted verbatim from the introduction of one of its own three source papers (dixit-2023), where they function as the standard rhetorical setup a paper writes before it fills the gap itself -- and its own results section does fill it, with an equity-stratified model down to wealth quartiles and out-of-pocket expenditure. Nobody checked whether the paper's own body closed the gap its own introduction stated, and nobody checked for an earlier, broader paper (Coates 2021, two years prior, 55 countries) that already ran the same combination comparison at a continental scale. Separately, the 'obstacle' is misdiagnosed: it is not that cost-effectiveness modelling methodology for RHD prevention is missing -- five published Markov/state-transition models exist across 20+ years -- it is that many individual countries lack the local epidemiological and cost inputs to run the existing, transferable model on themselves. That is a data-collection gap, not an economic-evaluation gap, and building yet another model would not fix it.
12 searches run against this one.
Social determinants of ARF and RHD outcomes are asserted but never measuredREFUTED
VerdictSocial determinants of ARF/RHD are one of the most heavily measured relationships in this literature - household crowding, individual SES and post-surgical socioeconomic status all have quantified, adjusted effect sizes across multiple settings - the gap sentences are three different papers' single-study caveats, not a field-wide silence.
Already done byJack S, et al. Risk factors for acute rheumatic fever: A case-control study. Lancet Reg Health West Pac. 2022;doi:10.1016/j.lanwpc.2022.100508
Individual/household-level matched case-control study (124 first-episode ARF cases, 372 population controls, NZ) measuring household crowding, deprivation index and access to primary care as ARF risk factors with adjusted odds ratios (already on our own shelf: aOR crowding 3.88 [1.68-8.98], aOR barriers to primary care 2.07 [1.08-4.00]).
Already done byJaine R, Baker M, Venugopal K. Acute rheumatic fever associated with household crowding in a developed country. Pediatr Infect Dis J. 2011;30(4):315-9
Individual-level case-control measurement of household crowding as an ARF risk factor, New Zealand.
Already done byNCDS 1958 birth cohort analysis - household overcrowding/family size and childhood ARF (on our own shelf, wiki/evidence/CLAIMS-FLAT-rhd.md)
Prospective UK birth-cohort measurement of household crowding/family size against incident ARF (OR 3.3, 95% CI 1.3-9.0) - individual-level, longitudinal, decades before the candidate's cited papers.
Already done byZhang Y, et al. Global burden of rheumatic heart disease and its association with socioeconomic development status, 1990-2019. Eur J Prev Cardiol. 2022;doi:10.1093/eurjpc/zwac044
GBD-derived incidence/mortality/DALY trends for RHD stratified by Socio-Demographic Index (SDI) 1990-2019 across all GBD regions including South Asia - a country-level socioeconomic gradient explicitly measured against the death/DALY trend the candidate says was never tested.
What it could not killOne narrow thread: no single published study has run a formal ecological regression testing average household size AND antibiotic-use/access-to-care rates together, specifically as explanatory variables for the GBD-derived RHD death-RATE TREND in South Asia (as opposed to prevalence levels, or SDI as a proxy). The naseeb-2024 gap sentence is accurately quoted from the source paper's discussion, which genuinely did not test this specific regression. But this is a narrow, single-technique gap inside a topic that is otherwise saturated with individual-level and cross-national evidence pointing the same direction.
The only version still worth doingIf someone insisted on doing it: an ecological time-series/panel regression of GBD South Asia RHD death rates (1990-2021, by country-year) against household-crowding proxies (DHS/census average household size) and antibiotic-access proxies (health system access indices, e.g. UHC service coverage index) would close the narrow surviving thread - but it would be an ecological-level study answering a question the field already has better-designed individual-level answers to (NZ case-control, UK cohort), so it upgrades certainty at the margin rather than answering something genuinely unknown. It is not worth prioritizing over collecting SES/social-determinant fields prospectively in the next RHD surgical outcomes cohort (which actually would close doran-2023's real local gap) or over acting on the individual-level crowding evidence that already exists for the housing-vs-clinics spending decision.
Objection beyond coverageThe three gap sentences are not one field-wide finding - they are three different papers' own methodological caveats, laundered into a single claim by pattern-matching on the phrase 'social determinants.' doran-2023's sentence is a retrospective SURGICAL cohort in the NT Aboriginal population saying it didn't collect SES fields to explain post-surgery SURVIVAL - a local data-collection gap in one dataset, not a claim that SES-and-survival is unstudied (it is well studied in other cardiac-surgery cohorts, see killer papers). pandey-2025's sentence is literally the boilerplate ecological-fallacy caveat that applies to EVERY aggregate meta-analysis by definition ('as with any aggregate meta-analysis, inference is limited to study-level factors') - restating a universal methodological truism as if it were a discovered gap in ARF/RHD research specifically. naseeb-2024's sentence is the only one that is a genuine untested-in-that-exact-form claim, and even it is narrow: it names two specific candidate mediators (household size, antibiotic access) for one specific GBD trend pattern in one specific region, when the causal relationship between crowding/access and ARF risk is already established at the individual level (NZ case-control, UK birth cohort) and at the cross-national level (SDI-stratified GBD analysis). The 'obstacle' claim in the candidate - 'the datasets already exist but were never joined' - is asserted, not evidenced; nobody has shown that household-size-by-year and antibiotic-access-rate-by-year data of adequate quality exist for the five South Asian countries in question, so the obstacle itself may be as fictional as the gap.
16 searches run against this one.
Managing and predicting risk in pregnant women who already have RHDREFUTED
VerdictBoth the screening decision and the risk-prediction decision for RHD in pregnancy are already served by a 2024 WHO-guideline systematic review, an RHD-specific validated risk score (DEVI, replicated twice), a 129-patient Australian outcomes study, and a live Australian screening trial — the source paper's own quote that no Australian outcomes study exists is simply wrong, contradicted by a paper published a year before it.
Already done bySeitler S, Ahmad M, Chu Ahuja SA, et al. Routine Antenatal Echocardiography in High-Prevalence Areas of Rheumatic Heart Disease: A WHO-Guideline Systematic Review. Global Heart 2024;19(1):39.
Systematic review (MEDLINE+Embase, double-blind screening) of 10 studies on antenatal echocardiographic RHD screening, framed as answering exactly the candidate's 'whether to screen' question. Already sits in our own shelf (wiki/evidence/_gaps/seitler-2024-routine-antenatal-echocardiography.md) with its own catalogued residual gaps: RCTs vs standard care, cost-effectiveness, optimal timing, universal-vs-targeted selection. None of those residual gaps match the candidate's framing of 'datasets exist but were never joined.'
Already done byLam CKM, Thorn J, Lyon X, Waugh E, Piper B, Wing-Lun E. Rheumatic heart disease in pregnancy: Maternal and neonatal outcomes in the Top End of Australia. Aust N Z J Obstet Gynaecol. 2023.
9.5-year retrospective study of 129 RHD pregnancies at the Northern Territory's largest obstetric referral hospital, explicitly evaluating maternal/neonatal outcomes AND the current model of care. This is the exact Australian outcomes-under-current-practice study that candidate paper aliyu-2024 claims does not exist — published a full year before Aliyu's claim.
Already done byMarangou J, Ferguson D, Unger HW, Kaethner A, Ilton M, Remenyi B, Ralph AP. Assessing the Role of Echocardiography in Pregnancy in First Nations Australian Women: Is it an Underutilised Resource? Heart Lung Circ. 2024.
322-woman retrospective analysis of echocardiography referral practice and RHD disease patterns in pregnant Northern Territory women, explicitly weighing 'consideration of universal screening during pregnancy' — the same-year, same-country, same-question paper Aliyu 2024 missed.
Already done byJones B, Marangou J, Yan J, et al. NEARER SCAN (LENO BESIK): evaluation of a task-sharing echocardiographic active case finding programme for rheumatic heart disease in Australia and Timor-Leste. BMJ Open 2024.
Live implementation trial (NCT06002243) screening ~1500 children and pregnant women in Australia and Timor-Leste for RHD via task-shared handheld echocardiography, with cost/effectiveness/implementation evaluation — the 'whether to screen' decision is already being operationalised, not merely researched.
What it could not killOne narrow point survives, but it is not the candidate's real claim. Chen 2025 states in its own limitations section that 'studies directly linking age at first birth to RHD risk are limited' — meaning a dedicated epidemiological study isolating age-at-first-birth as a determinant of acquiring RHD (as opposed to a risk factor for adverse outcomes once RHD is already present) does not obviously exist, and my searches found none. This is a footnote-level aside in a burden-of-disease paper's discussion, not a load-bearing part of the 'screen and manage' decision the candidate frames as the gap, and the causal direction is muddled: RHD is caused by antecedent untreated GAS pharyngitis/rheumatic fever in childhood, so 'age at first birth' would at most be a correlate of exposure timing (poverty, early marriage, delayed care-seeking) rather than a mechanistic risk factor — a confound-heavy question, not a clean research design.
The only version still worth doingThe one genuinely open design is the RCT Seitler 2024 itself calls for: a controlled trial of routine antenatal echocardiographic screening versus standard antenatal care in a high-RHD-prevalence LMIC setting, powered on maternal/fetal composite outcomes and paired with a formal cost-effectiveness analysis. That is a real, named gap from an existing systematic review — not the candidate's framing, and not something Lead+D is positioned to run.
Objection beyond coverageThe candidate's stated obstacle ('the datasets already exist but were never joined') describes a data-engineering problem. What actually exists is the opposite: a completed WHO-commissioned systematic review (Seitler 2024) that already did the synthesis work, two independent validations of a purpose-built rheumatic-VHD pregnancy risk score, and a live national screening trial. The gap was never 'nobody joined the datasets' — it is that the source paper (Aliyu 2024) did not cite a directly on-point paper (Lam 2023) published in a mainstream ANZJOG journal a year earlier. That is a citation-completeness failure in one review article, not an unaddressed research question. Chasing it would mean re-running a study whose result and 'current model of care' evaluation were already published, in the same country, under an even larger and more recent dataset (Marangou 2024, n=322) that also explicitly weighed universal screening.
13 searches run against this one.
What actually drives adherence to secondary prophylaxis and what level is enoughREFUTED
VerdictThe 'exact adherence level required' is already quantified with dose-response curves (de Dassel 2018), the candidate's own second source paper states the ≥80% threshold as established fact, and Uganda's own registry team already joined adherence data to RHD-progression outcomes and got a published, if awkward, answer.
Already done byde Dassel JL, de Klerk N, Carapetis JR, Ralph AP. How Many Doses Make a Difference? An Analysis of Secondary Prevention of Rheumatic Fever and Rheumatic Heart Disease. J Am Heart Assoc. 2018.
Nested case-control (97 ARF recurrence cases vs 194 controls) and case-crossover analysis on 7,728 annual adherence estimates from the Northern Territory, Australia. Established the dose-response curve directly: every 10pp increase in adherence = 17-21% lower odds of ARF recurrence; <80% adherence = 4-fold higher odds of recurrence (OR 4.00, 95% CI 1.72-9.29); risk did not measurably drop until ~40% of doses were received; population-attributable-fraction analysis showed 64-69% of recurrences were preventable at >=80% adherence. This is a quantified answer to 'what level of adherence is required,' already sitting in our own shelf (wiki/evidence/_drafts/dassel-2018-many-doses-make.md, 22 extracted claims).
Already done byBeaton A, Aliku T, Dewyer A, et al. Latent Rheumatic Heart Disease: Identifying the Children at Highest Risk of Unfavorable Outcome. Circulation. 2017.
Uses the Ugandan National RHD Registry (n=227 latent RHD, the same registry infrastructure minja-2023 and chang-2020 draw on) to join penicillin prescription/adherence data to echocardiographic progression outcomes via propensity matching, Cox models and Kaplan-Meier analysis. This directly refutes the candidate's stated 'obstacle' that Uganda's adherence and outcome datasets 'were never joined' -- they were joined, in this paper, by the same research group. Result: no ARF recurrences were captured in the registry during follow-up at all (explaining why a Ugandan ARF-recurrence dose-response study does not exist -- there is no event to model, not neglect), and penicillin prophylaxis in borderline RHD showed a paradoxical, non-significant ~2-fold *increase* in odds of progression (OR 2.57, 95% CI 0.81-8.19), almost certainly confounding by indication. Already on our shelf at wiki/evidence/CLAIMS-FLAT-rhd.md lines 620-645.
Already done byMinja NW, Pulle J, Rwebembera J, et al. Evaluating the implementation of a dynamic digital application to enable community-based decentralisation of rheumatic heart disease case management in Uganda: protocol for a hybrid type III effectiveness-implementation study. BMJ Open. 2023.
This is one of the TWO source papers the candidate itself cites as evidence for the gap. Its own background section states as settled fact: 'Secondary antibiotic prophylaxis... has been shown to be effective... contingent on achieving an optimum adherence, at least >=80% coverage of prescribed injections,' explicitly citing 'a Cochrane review and dose-response analyses' as the basis. The candidate's own supporting citation contradicts the candidate's central claim.
Already done byLongenecker CT, Morris SR, Aliku TO, et al. Rheumatic Heart Disease Treatment Cascade in Uganda. Circ Cardiovasc Qual Outcomes. 2017.
Uganda RHD Registry (n=1504) treatment cascade analysis: among patients retained in care, 91.4% already achieved optimal (>80%) BPG adherence. Retention in care, not adherence level once engaged, is identified as the binding constraint on outcomes in Uganda -- reframing what the 'switch the drug/route/person' decision should actually target.
What it could not killA narrow, technically-true residue: nobody has published a graded (continuous %-adherence) dose-response model of ARF recurrence specifically using Ugandan data, the way de Dassel did for the Northern Territory. But this is not because the datasets were 'never joined' (they were, by Beaton 2017, using the same Uganda registry that minja-2023 draws on) -- it is because the Uganda registry recorded zero ARF recurrence events during its follow-up window, so there is nothing to fit a recurrence dose-response curve to yet. The RHD-progression side of the question has been tested (binary prescribed/not, propensity-matched) with a paradoxical result that would need replicating with continuous adherence and better confounding-by-indication control before anyone should trust it either way.
The only version still worth doingNot a fresh 'does adherence level predict outcome' study -- that would just relitigate de Dassel 2018 on a smaller, event-starved cohort. The one design worth funding is a re-analysis of the existing Uganda (and ideally pooled sub-Saharan African, e.g. Ethiopia + Rwanda via bimerew-2024's low-income-country subgroup) registry data using continuous percent-adherence as the exposure and RHD echocardiographic progression/regression as the outcome, with a marginal structural model or target-trial emulation to handle the confounding-by-indication that produced Beaton 2017's paradoxical (penicillin users progressing more) result. That reuses data that already exists, targets the outcome that actually has enough events in Uganda (progression, not ARF recurrence), and would settle whether the counterintuitive 2017 finding was confounding or something real -- which is the one part of this bundle still genuinely unresolved.
Objection beyond coverageThis is two questions welded together, and only one of them is genuinely under-answered. 'What predicts adherence' is one of the most over-studied questions in this literature (bimerew-2024's own meta-analysis pools 33 studies and 7,158 patients on exactly that). 'What level of adherence is protective' has a published, quantified dose-response answer (de Dassel 2018) that the candidate's own second source paper (minja-2023) cites as settled when justifying its own protocol -- so the candidate is asserting as an open question something one of its own two cited papers treats as textbook. The 'Ugandan context' qualifier does the rhetorical work of resurrecting an answered question: Group A streptococcus-triggered autoimmune carditis is not a locally-varying biological process, and the decision at stake (keep the injection vs switch drug/route/person) is a delivery-logistics decision, not one that plausibly turns on whether Uganda's dose-response curve has a different shape than Australia's or Ethiopia's. The 'datasets exist but were never joined' framing is also factually wrong -- Uganda's own registry group joined them in 2017 and found the actual obstacle is an events problem (zero ARF recurrences recorded), not a data-linkage problem.
14 searches run against this one.
Prevalence data simply do not exist for most countriesREFUTED
VerdictCountry-level RHD burden data already exist for nearly every named 'data-free' region (SE Asia, Pacific, Latin America, North Africa, Central Asia, sub-Saharan Africa) as primary surveys, region-wide meta-analyses, a national multi-country mortality registry, and a Global Burden of Disease model that assigns every country including Malawi a numeric estimate -- and that GBD estimate for Malawi (323,422 cases, 2021) is sitting inside the candidate's own source paper.
Already done byGBD RHD Collaborators. Global, Regional, and National Burden of Rheumatic Heart Disease, 1990-2015. N Engl J Med. 2017.
Produces a modeled RHD burden estimate (prevalence, mortality, DALYs) for every one of the ~195 countries in the GBD framework, using covariate-based statistical modeling to fill in countries with no primary survey data, not just the ones with echo screening studies. This is precisely the 'datasets already exist but were never joined' obstacle the candidate frames as unmet -- GBD is the joining exercise, and it has existed since 2017 and is updated on every GBD cycle. It directly contradicts the 'a health ministry does not know how much of this disease exists inside its own borders' decision framing: every ministry already has a number, whether or not their country ran a survey.
Already done byblennerhassett-2025-burden-rheumatic-heart (the candidate's own source paper). Claim #16 on our shelf: 'Malawi, one of the world's poorest nations, had an estimated 323,422 RHD cases in 2021 according to the Global Burden of Disease study.'
The candidate's own cited paper reports a specific GBD-modeled prevalence figure for Malawi, the exact country the candidate's gap sentence #1 names as lacking data. The source material used to build this gap already contains the number the gap claims does not exist -- it just is not a primary survey number, and the gap sentence conflates 'no primary echo survey' with 'no prevalence data.'
Already done byOlsen J, Chimalizeni Y, Carapetis J, et al. Distance from tertiary care, pericardial effusion, and nutritional status predict all-cause mortality among Malawian children with rheumatic heart disease. medRxiv preprint, 2026.
A Malawi-specific, country-level RHD mortality cohort (n=118, 23.7% died during follow-up) with Cox regression on independent mortality predictors. This directly answers gap sentence #1 ('Mortality data attributable to RHD for Malawi could not be obtained from multi-country studies that pooled participants by income group') -- not by extracting Malawi's slice out of a pooled multi-country study, which was never going to work, but by a direct national mortality study that makes the pooled-extraction problem moot. Published 2026, after blennerhassett-2025's 1995-2024 search window, so it is new literature the scoping review could not have seen, not literature it missed.
Already done byNand N, et al. Population-based assessment of cardiovascular complications of rheumatic heart disease in Fiji: a record-linkage analysis. BMJ Open. 2023. / Rheumatic Heart Disease-Attributable Mortality at Ages 5-69 Years in Fiji: A Five-Year, National, Population-Based Record-Linkage Cohort Study. PLoS NTD. 2015.
National, population-based administrative record-linkage cohort studies for RHD mortality and cardiovascular complications, run and published twice (2015, 2023) in the same country. This shows the record-linkage method gap sentence #2 asks for is not a methodological unknown -- it is a proven, replicated design. What has not been done is the same design run simultaneously across multiple countries and across the full Strep A endpoint spectrum (pharyngitis/impetigo/invasive disease/ARF/RHD together, not RHD alone) -- see what_survives below.
What it could not killOne narrow, specific piece: a record-linkage study that runs simultaneously across multiple countries AND covers the full Strep A clinical spectrum in one design (pharyngitis, impetigo, invasive GAS, ARF, and RHD together, not RHD alone) has not turned up in any search. Fiji has done national record-linkage twice, but for one country and mostly RHD/cardiovascular endpoints. The moore-2022 framework paper itself (one of the candidate's own sources) explicitly names this as an open need for vaccine-value calculations, because financing decisions can hinge on high-incidence low-severity endpoints (pharyngitis, impetigo) that a rare-outcome-only design like Fiji's would miss. So gap sentence #2, narrowly read as 'multi-country AND all-endpoint AND record-linkage,' survives -- but gap sentence #1 (Malawi mortality) does not, and the broader subject line ('prevalence data simply do not exist for most countries') does not either.
The only version still worth doingThe one design that would still add real information is a multi-country administrative record-linkage study replicating the Fiji methodology (Nand 2023 / 2015) simultaneously in 3-4 countries that already have the prerequisite linkable health records (a national patient identifier and linkable hospital/death-registry data are the binding constraint, not lack of interest) AND that extends the linkage past RHD to the full Strep A endpoint chain moore-2022 calls for, specifically to generate the endpoint-level cost data that vaccine financing decisions need and that GBD-style modeled estimates cannot supply because GBD does not model pharyngitis/impetigo incidence with the granularity a Gavi-style investment case requires. That is a genuine, still-open contribution. A country-by-country echo-screening prevalence survey, by contrast, is not worth doing anywhere the Watkins/Mutarelli/Noubiap-style meta-analyses and GBD modeling already cover the region -- it would be one more dot on a map that already has enough dots to model from, and it would not change a single ministry's allocation decision that isn't already informed by an existing regional or modeled estimate.
Objection beyond coverageThe candidate's framing conflates three different things that should not be conflated: (1) no primary echo-screening survey has been run in a country, (2) no prevalence estimate exists for that country, and (3) a health ministry cannot make a resource-allocation decision without one. All three are false as stated. GBD has produced a modeled prevalence number for every country since 2017, precisely so that (2) is never true even when (1) is. And moore-2022 -- again, the candidate's own source -- states as established field practice that 'where country-specific estimates of disease burden are lacking, it is important to provide regional estimates to assist with decision making,' meaning (3) is also false: substituting regional/modeled estimates for missing country surveys is not a workaround someone needs to invent, it is the field's documented standard operating procedure. What remains genuinely missing is not 'data' in the sense the subject line uses the word -- it is a specific instrument design (simultaneous multi-country, all-endpoint, administrative-linkage) that would improve precision and endpoint breadth over what modeled estimates give you, which is a different and much smaller claim than 'data simply do not exist for most countries.'
19 searches run against this one.
GBD model estimates cannot be checked against real dataREFUTED
VerdictJoining registry/screening data to GBD RHD estimates to test their validity has already been done and published, including a paper already sitting on our own shelf next to this one.
Already done byKatzenellenbogen JM, et al. Contemporary Incidence and Prevalence of Rheumatic Fever and Rheumatic Heart Disease in Australia Using Linked Data: The Case for Policy Change. J Am Heart Assoc. 2020.
Built a 5-jurisdiction linked administrative/clinical ARF-RHD register covering ~86% of Indigenous Australians, then directly compared the extrapolated national prevalent-case count against the 2017 GBD modelled RHD prevalence for Australia <55 years: 1,518 (GBD) vs 6,156 (linked-data extrapolation), a 4-fold underestimate. This is exactly the registry-plus-population-data-joined-to-GBD comparison the candidate says has not been done, for RHD specifically, and it is already extracted in wiki/evidence/_drafts/katzenellenbogen-2020-contemporary-incidence-prevalence.md claim #4.
Already done byByass P, et al. Cause-specific mortality findings from the Global Burden of Disease project and the INDEPTH Network. Lancet Glob Health. 2016.
Systematically compares GBD's modelled cause-specific mortality estimates against independently collected, verified ground-truth mortality data from INDEPTH Network health-and-demographic-surveillance sites across multiple LMICs. Not RHD-specific, but it establishes that 'compare GBD model output to real, independently collected data in the same place and time' is an established, executed genre of methodological critique, not a blind spot nobody has thought to attack.
Already done byStark differences in cancer epidemiological data between GLOBOCAN and GBD: Emphasis on oral cancer and wider implications. eClinicalMedicine. 2022.
Another disease area (cancer), another instance of the same genre: cross-comparing two global burden-modelling systems and reporting the discordance as a finding worth publishing on its own. Further evidence this class of validation work is a recognized, fundable, publishable line of research, not an unattempted idea.
What it could not killNothing worth calling a literature gap. The one narrow thing left is Shimanda et al.'s own stated limitation that they personally could not trace which local Namibian data sources fed the GBD RHD estimate for that one country -- but that is the primary authors' own documented limitation, already published in 2022, not an undiscovered gap. It is not fixable by 'joining existing datasets' as the candidate's obstacle field claims: Shimanda et al. explicitly say Namibia has no population-based screening data to join -- they call for a NEW study to be conducted, not an existing dataset to be linked. So even the narrow Namibia claim is not the gap the candidate describes.
The only version still worth doingThe only version of this worth doing is not 'join Namibia's existing datasets' (they don't exist) but 'replicate Katzenellenbogen's linked-registry-vs-GBD design in a country that already has usable ground truth' -- e.g. use the REMEDY Global Rheumatic Heart Disease Registry (3,343 patients across 14 LMICs, clinically verified, prospectively followed) as the comparator against GBD's RHD estimates for those same 14 countries and years. That is a real, doable study with existing data. It is also not a gap: it is the same design Katzenellenbogen already ran, just not yet run against REMEDY specifically, which makes it an extension of established method rather than new science.
Objection beyond coverageThe candidate's own 'obstacle' field ('the datasets already exist but were never joined') is empirically false for the specific case it cites. Shimanda et al. 2022 -- one of only two source papers behind this candidate -- say the opposite in their own limitations section: Namibia has no population-based echocardiographic screening data at all, and they explicitly call for one to be conducted in the future. There is nothing sitting on a shelf waiting to be joined; the ground-truth dataset does not exist yet. Separately, the candidate generalizes a single country's data-provenance complaint into a claim about GBD RHD estimates in general being unverifiable against real data -- but Katzenellenbogen et al. 2020, sitting in our own evidence shelf one file away from Shimanda, did precisely this join (linked clinical registry data vs GBD modelled estimate) for Australia and got a clean, publishable, decision-relevant result (GBD undercounts by 4x). The candidate-generation process apparently did not cross-check the two source papers against each other, or against the rest of the shelf, before declaring the comparison undone.
24 searches run against this one.

Doable now, and something is waiting on it

Read this list only through the verdicts above. The top 12 were tested and none survived intact; the rest below have not been tested at all.

The only list here that is arithmetic rather than judgement: questions whose obstacle is cheap to remove and whose answer a real decision is waiting for. Ranked by how many separate research groups raised it.

Why RHD affects women roughly twice as often as mendoable  7 papers
Why nobody has closed itA finding repeated in every paper that nobody has ever chased
What is waiting on the answerWhether to screen pregnant women for RHD and how to manage the ones found, in the population where the disease is most lethal and least explained
“The reasons why male and female RHD prevalence were nearly identical in this pooled analysis have not been determined.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The reasons for the higher RHD prevalence among females remain not well understood, and the contribution of socio-cultural health-seeking differences versus biological factors is unresolved.” — chillo-2023-sub-clinical-rheumatic 2023
“The reasons for the high female-to-male ratio among RHD cases and the absence of prior symptom history in many patients are not fully understood.” — elazrag-2022-handheld-echocardiographic-screening 2022
Cost-effectiveness and sustainability of echo screening programmesdoable  6 papers
Why nobody has closed itNobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantified
What is waiting on the answerWhether a health ministry should run echocardiographic screening at all, and if so on whom, at what age, with which handheld device, read by whom
“The cost-effectiveness of WHO's recommendation to use handheld echocardiography for RHD screening in endemic areas has not been carefully evaluated.” — ali-2026-implementing-world-health 2026
“The cost and cost-effectiveness of primary care echocardiographic screening for RHD have not been measured.” — jiee-2024-heart-community-implementation 2024
“Cost-effectiveness of echocardiographic screening and penicillin prophylaxis for RHD has not been established.” — mutarelli-2025-global-prevalence-sex 2025
The economic value case for a GAS vaccinedoable  4 papers
Why nobody has closed itNobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantified
What is waiting on the answerWhether to license, fund and schedule a Strep A vaccine, and which antigen and age it should be built around
“Estimates of the non-internalized benefits of Strep A vaccination (e.g., population-wide AMR mitigation, herd effects, spillover mental health and quality-of-life effects on household member” — cadarette-2023-full-health-economic 2023
“No empirical VSLY estimates exist for lower-income countries, so VSLY is proxied by per-capita GDP with an unknown elasticitiy.” — cadarette-2023-full-health-economic 2023
“The magnitude of Strep A vaccination's impact on antimicrobial resistance, particularly through reduced antibiotic consumption, has not been quantified within this study.” — cadarette-2023-full-health-economic 2023
Economic evaluation of prevention strategies other than vaccinesdoable  3 papers
Why nobody has closed itNobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantified
What is waiting on the answerWhere a health ministry should put its prevention money, when it does not know how much of this disease exists inside its own borders
“No previous study has evaluated the distributional effect of primary, secondary, or tertiary interventions and their combinations for the prevention and control of rheumatic fever and rheuma” — dixit-2023-evaluating-efficiency-equity 2023
“Whether the same coverage rates for primary, secondary, and tertiary interventions across socioeconomic subgroups hold in reality has not been empirically tested for these interventions.” — dixit-2023-evaluating-efficiency-equity 2023
“Primary prevention using clinical diagnostic criteria and vaccination against group A streptococcal disease were not evaluated because effectiveness and cost data are unavailable.” — dixit-2023-evaluating-efficiency-equity 2023
Social determinants of ARF and RHD outcomes are asserted but never measureddoable  3 papers
Why nobody has closed itThe datasets already exist but were never joined
What is waiting on the answerWhether to spend on housing, crowding and skin health instead of on clinics and drugs
“Information on the social determinants of health was not captured, despite being hypothesised to explain the findings.” — doran-2023-surgery-rheumatic-heart 2023
“It has not been specifically tested whether the average number of people per household or rates of antibiotic use/access to medical care explain the trends in RHD deaths in South Asia.” — naseeb-2024-temporal-trends-burden 2024
“Individual-level determinants such as socioeconomic status and hygiene could not be examined because the analysis was limited to study-level factors.” — pandey-2025-regional-disparities-temporal 2025
At what age a GAS vaccine should be givendoable  3 papers
Why nobody has closed itNobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantified
What is waiting on the answerWhether to license, fund and schedule a Strep A vaccine, and which antigen and age it should be built around
“It is unclear what proportion of RHD cases arising at age 5 could be averted by an infant vaccination schedule.” — giannini-2022-potential-health-impact 2022
“Further epidemiology or immunology studies are required to model the proportion of RHD cases that could be averted by an infant vaccination schedule preventing the infections preceding ARF a” — giannini-2023-modeling-potential-health 2023
“The optimal age for S. pyogenes vaccine introduction in early life is unresolved and requires evaluation in vaccine trials and modelling.” — keeley-2025-early-life-serological 2025
Managing and predicting risk in pregnant women who already have RHDdoable  2 papers
Why nobody has closed itThe datasets already exist but were never joined
What is waiting on the answerWhether to screen pregnant women for RHD and how to manage the ones found, in the population where the disease is most lethal and least explained
“Outcomes of pregnant women with RHD under current clinical practice in Australia (raised as relevant to LMIC contexts) have not been previously evaluated.” — aliyu-2024-rheumatic-heart-disease 2024
“Studies directly linking age at first birth to RHD risk are limited.” — chen-2025-global-burden-trend 2025
An affordable point-of-care test for GAS pharyngitisdoable  2 papers
Why nobody has closed itNobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantified
What is waiting on the answerWhether this sore throat or this skin sore gets an antibiotic, and what test the health worker is allowed to hold while deciding
“The cost-effectiveness of molecular and rapid diagnostics for S. pyogenes pharyngitis in LMIC settings has not been established.” — armitage-2025-evaluating-clinical-decision 2025
“Estimates of number needed to treat (NNT), number needed to harm (NNH), and the antimicrobial resistance implications of using molecular diagnostics for S. pyogenes pharyngitis in LMICs have” — armitage-2025-evaluating-clinical-decision 2025
“The feasibility, acceptability, and affordability of point-of-care Strep A rapid diagnostic tests to enable directed (rather than syndromic) treatment has not been established.” — leong-2025-assessing-evidence-antibiotic 2025

What nobody can decide

Because this is not known, somebody is standing in front of a choice they cannot make.

Blocks no decision: study-internal caveats and unexplained observations that change nothing anyone is choosinglatent  50 papers
Who is stuckNobody. These are the authors themselves, discharging the limitations paragraph
What happens today insteadReaders skip them, and rightly: single-centre sampling, small n, recall and selection bias, English-language search restrictions, unexplained statistical heterogeneity, funnel-plot asymmetry, unblinded image interpretation, a dataset not yet deposited. They constrain the confidence of one paper, not the choice of any clinician, ministry, trialist or developer
Cost of the current default being wrongAlmost none individually. The real cost is collective and it is a finding in itself: roughly a fifth of everything this literature calls a research gap is a caveat about the study that produced it, which means the field systematically over-reports how much is unknown and under-reports which unknowns matter
What would settle itNothing needs to settle them. What would help is journals separating "limitation of this study" from "gap in the field", so that funders reading a discussion section can tell a stuck decision from a small sample
“Non-English-language population-based studies on RHD prevalence in Africa were not captured, which may bias the pooled estimate.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The reasons for heterogeneity in reported RHD prevalence across the included studies have not been fully explained despite subgroup analysis by region, country and sampling unit.” — abdu-2024-prevalence-pattern-rheumatic 2024
“Future RHD task-sharing studies should be conducted by multi- and interdisciplinary teams with expertise in health services and implementation science methods.” — abdullahi-2019-task-sharing-diagnosis 2019
Where a health ministry should put its prevention money, when it does not know how much of this disease exists inside its own bordersactive  27 papers
Who is stuckA national NCD planner writing a budget line, and the donor deciding whether this country qualifies for RHD funding at all
What happens today insteadCites a Global Burden of Disease estimate whose local inputs it cannot identify, extrapolated from a single school study, an urban hospital register, or a pooled low-income-country average; there is no mandatory notification, no register, and most published prevalence comes from schoolchildren in cities while the disease peaks in adults and rural areas
Cost of the current default being wrongMoney goes to the country that published a survey rather than the country with the disease; a ministry drops a program because registered cases fell when in fact detection fell; adults, rural populations and whole regions (Central Africa, Latin America outside Brazil, North Africa, Southeast Asia) stay invisible and therefore unfunded
What would settle itPopulation-based, community (not school) prevalence surveys using standardised 2023 WHF criteria across the ages where disease actually sits, plus mandatory ARF/RHD notification feeding a linked national register, validated against the GBD inputs it is meant to correct
“No population-based data on rheumatic heart disease prevalence from North Africa exists beyond a single Egyptian school study.” — abdu-2024-prevalence-pattern-rheumatic 2024
“Rheumatic heart disease in adults and non-school-attending populations across Africa is largely unmeasured, since most included studies drew samples from school lists.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The true community-level prevalence of RHD in Al Managil is unknown because screening was conducted in schools rather than the community.” — ali-2018-handheld-echocardiography-screening 2018
Whether a health ministry should run echocardiographic screening at all, and if so on whom, at what age, with which handheld device, read by whomactive  24 papers
Who is stuckA ministry of health NCD director deciding whether to buy 40 handheld probes and pull nurses out of clinics for a school screening round
What happens today insteadEither does nothing, or runs a donor-funded school screening campaign on convenience samples with locally invented training and locally chosen criteria, because there is no agreed curriculum, no device-specific cut-off, no accuracy data against the 2023 WHF A-D staging, and no cost-effectiveness estimate anywhere
Cost of the current default being wrongA program that spends the entire NCD budget generating borderline results nobody can act on, refers children to a tertiary echo they never reach, and leaves the ministry unable to defend the spend; or a decision not to screen that misses the only cheap window before valves are surgical
What would settle itA randomised or stepped-wedge comparison of screening strategies (clinical exam, single-view handheld by a trained non-expert, full echo) against 2023 WHF confirmatory criteria in the same population, reporting sensitivity by device and operator training dose, downstream referral completion, harms, and cost per case managed
“No randomised or controlled study has compared different screening strategies (notably clinical evaluation alone versus echocardiography) within African populations, so the yield attributabl” — abdu-2024-prevalence-pattern-rheumatic 2024
“No literature exists on the implementation of handheld echocardiography screening for subclinical RHD in Pakistan.” — ali-2021-detection-subclinical-rheumatic 2021
“The accuracy of a deep learning algorithm trained on DAS waveform data to classify subclinical RHD into definite, borderline, and normal categories has not been tested.” — ali-2021-detection-subclinical-rheumatic 2021
Whether to license, fund and schedule a Strep A vaccine, and which antigen and age it should be built aroundactive  18 papers
Who is stuckA vaccine developer choosing an antigen for a phase 3, and the Gavi/national immunisation committee deciding whether an infant Strep A dose earns a slot in the schedule
What happens today insteadRuns immunogenicity trials against antibody titres that nobody has shown to correlate with protection, in animal models nobody has validated, with antigen coverage inferred from North American emm distribution, and builds the investment case on models that exclude ARF entirely because prevalence data do not exist
Cost of the current default being wrongA century of stalled development repeats: a candidate that raises antibodies and prevents nothing, an antigen that revives the autoimmunity fear and shuts the field down again, or a vaccine that works in Utah and misses the emm types circulating where RHD actually kills
What would settle itA controlled human infection or large field-efficacy trial that fixes a mechanistic correlate of protection, paired with multi-region emm and serology surveillance and a transmission-dynamic impact model that carries the infection-to-ARF-to-RHD pathway rather than skipping it
“No licensed, effective vaccine against GAS exists to control RHD in Africa.” — aliyu-2024-rheumatic-heart-disease 2024
“Large-scale clinical trials are necessary to verify the complete safety of the 6-valent GAS vaccine, since the Phase 1 findings were argued to be biased by the small-scale, open-label (non-b” — azuar-2019-recent-advances-development 2019
“It is unknown whether the 30-valent M-protein-based vaccine, whose Phase 1 trial was planned for 2015, will proceed, since no further information is available regarding this study.” — azuar-2019-recent-advances-development 2019
Whether a decentralised, non-specialist-run RHD service can carry the care, and whether it survives the end of the pilot grantactive  18 papers
Who is stuckA district health manager deciding whether to move RHD follow-up from the referral hospital to health centres staffed by nurses and community health workers, and whether to keep a parallel RHD register or fold it into the national health information system
What happens today insteadRuns care from one or two national referral centres so that patients travel and default; where decentralisation exists it is a donor-funded pilot with an externally secured drug supply, no external control group, no cost data and no evidence it can be absorbed by government
Cost of the current default being wrongPatients drop out of care for the entirely mundane reason that the clinic is far and the queue is long, and a promising model is either scaled without evidence or abandoned when the grant closes, taking the register and the drug supply with it
What would settle itA controlled non-inferiority evaluation of decentralised versus centralised RHD care on prophylaxis adherence and retention, costed from the government and household side, with an explicit test of whether the register can be integrated into routine health information systems and sustained past the funding period
“No evidence exists on task-sharing models for expanding access to RHD prevention and treatment services in limited-resource settings.” — abdullahi-2019-task-sharing-diagnosis 2019
“No studies have addressed prevention of acute rheumatic fever (primary prevention) or passive case-finding of recurrent ARF and RHD through task-sharing approaches.” — abdullahi-2019-task-sharing-diagnosis 2019
“No studies have evaluated task-sharing models for tertiary prevention, including post-operative management of patients after heart valve surgery.” — abdullahi-2019-task-sharing-diagnosis 2019
Which biological target a drug developer should aim at, when nobody can say what actually drives the valve from inflammation to scarlatent  18 papers
Who is stuckA translational researcher or funder choosing between the immune, fibrotic, lymphatic and genetic hypotheses for the first ever disease-modifying RHD drug
What happens today insteadNothing is developed: there is no drug that halts or reverses rheumatic valve fibrosis, targets are picked from single-cohort mechanistic papers, and every candidate is tested in rodent models that cannot reproduce the slow human disease and in cohorts that are almost never African despite the burden being African
Cost of the current default being wrongThe field keeps generating mechanism papers with no therapeutic exit, and RHD stays a surgical disease forever: the only interventions on offer remain a 70-year-old antibiotic and a heart operation most patients cannot reach
What would settle itHuman rheumatic valve tissue and matched plasma at single-cell resolution across disease stages from high-burden populations, feeding patient-derived valve organoid or ex vivo flow models capable of testing candidate antifibrotic and immunomodulatory targets, plus adequately powered multi-ancestry GWAS
“Validated, widely adopted molecular biomarkers (e.g., Tenascin-C, miR-1183/miR-1299, GlcNAc-specific IgG2) for RHD are not yet established in African clinical practice.” — aliyu-2024-rheumatic-heart-disease 2024
“Large-scale validation of molecular biomarkers for RHD in diverse African populations has not been done.” — aliyu-2024-rheumatic-heart-disease 2024
“Molecular biomarkers that stratify RHD patients by risk of disease progression, guide resource allocation, and signal need for intensive monitoring remain unidentified.” — aliyu-2024-rheumatic-heart-disease 2024
Whether to start (and how long to continue) monthly penicillin in a symptom-free child whose only sign of disease is a screening echourgent  16 papers
Who is stuckA district clinician or RHD program nurse holding a screen-positive echo report on a well 10-year-old, and the parents being asked to consent to years of injections
What happens today insteadTreats definite latent RHD as if it were clinical RHD and puts the child on 4-weekly benzathine penicillin indefinitely; leaves borderline cases in limbo with a repeat echo "sometime", because the one randomised trial (GOAL) showed benefit in latent RHD overall but did not settle borderline disease, stopping rules, or who was going to regress anyway
Cost of the current default being wrongEither a decade of painful injections, clinic visits, lost school days and a cardiac label for children whose valves would have normalised on their own, or silent progression to a valve that needs surgery in a country with no cardiac surgeon
What would settle itA multi-country prospective cohort of screen-detected borderline and definite latent RHD followed 10+ years with standardised 2023 WHF staging, nested in a randomised trial of prophylaxis versus surveillance with a pre-specified cessation rule and a progression-risk marker measured at entry
“Long-term follow-up of Africans diagnosed with borderline RHD to determine progression to definite RHD has not been performed in Africa; existing progression data come from South Africa and ” — abdu-2024-prevalence-pattern-rheumatic 2024
“It is unknown whether echocardiography-based active case finding improves outcomes for RHD and how to manage borderline RHD detected through screening.” — abdullahi-2019-task-sharing-diagnosis 2019
“Prospective studies are needed to confirm whether neutrophil count and NLR predict development of RHD in children with ARF.” — akat-2026-neutrophil-count-neutrophil 2026
Whether to keep giving the painful 4-weekly injection, or switch the drug, the route or the person who gives iturgent  15 papers
Who is stuckA nurse in a rural health post with a frightened teenager who has missed three doses, and a prescriber deciding whether the injection is safe enough to keep giving after a death was reported in the next district
What happens today insteadFollows a 1950s regimen: intramuscular benzathine penicillin G every four weeks, indefinitely, on an adherence target of 80% that nobody has validated; some Ethiopian centres have simply stopped giving it after adverse events and switched to oral amoxicillin without evidence either way
Cost of the current default being wrongPatients quietly drop out and come back in heart failure; or a program keeps inflicting a painful injection that a subcutaneous infusion or an oral regimen would have replaced; or a preventable death after injection ends a national program on rumour rather than data
What would settle itA trial of oral versus intramuscular penicillin for progression (already under way), an active pharmacovigilance registry that separates vasovagal collapse from true anaphylaxis and quantifies fatal reactions in severe RHD, and dose-response data linking measured adherence to actual recurrence
“There is currently no ideal alternative drug to benzathine penicillin G for treating GAS or providing secondary prophylaxis for ARF.” — ali-2022-rheumatic-heart-disease 2022
“The causes and risk factors for fatal reactions following benzathine penicillin G injection remain not fully understood.” — ali-2026-implementing-world-health 2026
“An antibiotic substitute for BPG with better administration techniques and a safer profile has not been developed.” — ali-2026-implementing-world-health 2026
Whether to spend on housing, crowding and skin health instead of on clinics and drugsactive  11 papers
Who is stuckA housing or environmental health authority in a remote or informal-settlement community being asked to fund ventilation, washing facilities and reduced crowding on the argument that it will prevent heart disease
What happens today insteadCites crowding as a known ARF risk factor and then does nothing measurable, or installs environmental health hardware whose effect on Strep A transmission has never been tested; even the baseline (how much ventilation, how much crowding, which transmission route dominates) has not been measured
Cost of the current default being wrongCapital spent on the wrong intervention while transmission continues; or the strongest and most durable prevention lever in the disease is left unused because no trial exists to justify the budget line, so the health system keeps paying for injections downstream forever
What would settle itA community-led prospective trial of a defined package of housing and environmental initiatives with Strep A infection and ARF incidence as endpoints and a measured pre-intervention baseline, plus transmission-route studies that establish whether droplet, skin contact or fomite dominates in these homes
“It is not known whether and how RHD outcomes differ by social determinants in African populations.” — aliyu-2024-rheumatic-heart-disease 2024
“Whether reducing sugar-sweetened drink intake could lower ARF/RHD risk in children is unknown and requires investigation.” — baker-2022-risk-factors-acute 2022
“The role of specific housing factors (damp, mould, cold, hot water supply, bed sharing, functional crowding) as ARF risk factors remains to be investigated more rigorously.” — baker-2022-risk-factors-acute 2022
Whether to screen pregnant women for RHD and how to manage the ones found, in the population where the disease is most lethal and least explainedactive  11 papers
Who is stuckAn antenatal clinic midwife deciding whether to put a probe on every pregnant woman, and the obstetric team managing a woman in her third trimester with newly discovered mitral stenosis
What happens today insteadDoes not screen, finds RHD when the woman decompensates in labour, and manages her with risk models built in high-income countries on congenital heart disease; the two-fold female excess in RHD has never been explained, so nobody knows whether to target women at all
Cost of the current default being wrongMaternal deaths that a first-trimester echo and a delivery plan would have prevented; or universal antenatal echo rolled out at scale on no accuracy, timing or cost data, absorbing the antenatal budget of the countries least able to spare it
What would settle itA controlled trial of routine antenatal echocardiographic screening versus standard care in a high-prevalence setting, reporting maternal outcomes, optimal gestational timing, universal versus targeted selection and cost; plus an explanation of the sex difference and RHD-specific pregnancy risk stratification validated in LMIC settings
“The reasons why male and female RHD prevalence were nearly identical in this pooled analysis have not been determined.” — abdu-2024-prevalence-pattern-rheumatic 2024
“Outcomes of pregnant women with RHD under current clinical practice in Australia (raised as relevant to LMIC contexts) have not been previously evaluated.” — aliyu-2024-rheumatic-heart-disease 2024
“Studies directly linking age at first birth to RHD risk are limited.” — chen-2025-global-burden-trend 2025
Whether to call this acute rheumatic fever at the bedside, with no test that can confirm or exclude iturgent  10 papers
Who is stuckA clinical officer in a district hospital looking at a febrile child with sore joints, deciding whether to write a diagnosis that commits the child to ten years of injections
What happens today insteadApplies the Jones criteria, which are a 1944 clinical construct with no gold standard behind them, in a setting where malaria, typhoid and viral arthritis look the same; then either over-calls "probable ARF" and starts prophylaxis on everybody with a sore knee, or under-calls it and the child comes back at 25 with a destroyed mitral valve
Cost of the current default being wrongOver-diagnosis buys years of painful injections and a cardiac label for children who never had ARF; under-diagnosis is the single largest silent feeder of the RHD pipeline, and neither error is currently visible to the person making it
What would settle itProspective validation of a candidate ARF biomarker panel (for example the 5-protein plasma signature) in diverse endemic cohorts against long-term echo outcome rather than against the Jones criteria, plus a TRIPOD-grade multiparametric risk score with category-specific management pathways
“A specified protocol to help frontline health workers discriminate ARF from its endemic mimickers (e.g., malaria, typhoid, viral infections) has not been defined.” — ali-2022-rheumatic-heart-disease 2022
“It has not been determined whether treating any joint symptom as probable ARF in endemic settings causes net harm from overdiagnosis and unnecessary BPG prophylaxis.” — ali-2022-rheumatic-heart-disease 2022
“A definitive diagnostic test for acute rheumatic fever has not been discovered, although biomarker discovery studies (START; ARF Diagnostic Collaborative Network) are still underway.” — ali-2026-implementing-world-health 2026
Whether this sore throat or this skin sore gets an antibiotic, and what test the health worker is allowed to hold while decidingurgent  10 papers
Who is stuckA frontline health worker in an endemic community with a queue of children with sore throats, no culture, and a clinical score built on Western low-prevalence populations
What happens today insteadUses a clinical decision rule of poor accuracy, or treats everyone syndromically, because rapid antigen tests are insensitive, molecular tests are unaffordable, and there is no way to separate true infection from asymptomatic carriage; skin infections are largely ignored as an ARF pathway despite being suspected as a major driver
Cost of the current default being wrongEither mass over-prescription driving resistance and cost with minimal ARF prevented, or missed streptococcal infections in exactly the children whose first ARF episode was preventable for the price of one course of penicillin
What would settle itA trial in a high-ARF setting of point-of-care molecular testing versus syndromic management with ARF incidence as the endpoint, reporting number needed to treat, number needed to harm and antibiotic volume, plus a study that finally quantifies whether treating Strep A impetigo lowers ARF risk
“No reliable and affordable test for GAS pharyngitis currently exists to replace clinical algorithms.” — ali-2022-rheumatic-heart-disease 2022
“A reliable and affordable diagnostic test for group A streptococcal pharyngitis that is available and affordable in limited-resource settings has not been developed.” — ali-2026-implementing-world-health 2026
“The molecular point-of-care test for group A streptococcus (sensitivity 100%, specificity 79%), although documented in Australia, has not been evaluated for availability and affordability in” — ali-2026-implementing-world-health 2026
When to send this patient for valve surgery and which valve to put in, knowing they may never come back for an INR testurgent  10 papers
Who is stuckA cardiac surgeon in a low-income setting choosing between a mechanical valve the patient cannot anticoagulate and a tissue valve that will fail in a 22-year-old, and the cardiologist deciding whether to refer before pulmonary pressures rise
What happens today insteadApplies Western age-based valve-selection guidelines built on degenerative disease in patients with reliable INR services, operates late because referral happens at decompensation, and repairs or replaces on surgeon preference because rheumatic-specific durability data barely exist
Cost of the current default being wrongA young person gets a mechanical valve, defaults warfarin monitoring and dies of valve thrombosis; or gets a tissue valve and needs a redo operation that the country cannot provide; or is referred a year too late and the pulmonary hypertension is already irreversible
What would settle itA risk-stratified comparative trial of tissue versus mechanical replacement in young RHD patients in LMIC settings, incorporating pre-operative adherence history, atrial fibrillation and travel time; plus a prospective study of early versus conventional-timing surgery with survival as the endpoint
“Repair of the rheumatic aortic valve remains in its infancy, with only limited recent reports of leaflet extension or mobilization/shaving techniques.” — antunes-2020-global-burden-rheumatic 2020
“Recurrent RF episodes and ongoing rheumatic activity place mitral valve repair at high risk of recurrent regurgitation or stenosis, but the magnitude and predictors of this failure are not c” — antunes-2020-global-burden-rheumatic 2020
“Whether autologous or heterologous pericardium (fresh or glutaraldehyde-treated) provides durable leaflet extension in very young rheumatic patients has not been established.” — antunes-2020-global-burden-rheumatic 2020

Why it is still open

Not what the gap is about. What is physically stopping anyone from closing it.

No registry or surveillance system was ever built, so there is no denominatorexpensive  44 papers
What this obstacle holds upEvery question of the form how many, where, in whom and how has it changed: national and regional prevalence, ARF incidence, adverse events after penicillin, strain and resistance surveillance, and the primary data underneath global burden estimates.
What would break the obstacleMandatory notification of ARF and RHD, national registers with unique identifiers, sentinel surveillance sites for pharyngitis and impetigo in the countries that carry the burden, and one round of population-based echocardiographic screening in the many countries that have never had one.
Who is positioned to do itA health ministry with donor or WHO backing; this is public infrastructure, not a study, and no research group can build it alone.
“No population-based data on rheumatic heart disease prevalence from North Africa exists beyond a single Egyptian school study.” — abdu-2024-prevalence-pattern-rheumatic 2024
“Rheumatic heart disease in adults and non-school-attending populations across Africa is largely unmeasured, since most included studies drew samples from school lists.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The true community-level prevalence of RHD in Al Managil is unknown because screening was conducted in schools rather than the community.” — ali-2018-handheld-echocardiography-screening 2018
The programme was rolled out with no evaluation attached to itmoderate  40 papers
What this obstacle holds upWhether task-sharing, decentralised penicillin delivery, school screening, registers, training curricula, community education and national control programmes actually work, and whether they survive past the pilot.
What would break the obstacleAttaching an evaluation design before implementation: a comparator or stepped-wedge rollout, pre-specified implementation outcomes and thresholds, incidence and prevalence endpoints rather than adherence proxies, and a published description of the structures and processes used, so the next country can copy something more than an anecdote.
Who is positioned to do itA clinical network or ministry programme office paired with an implementation-science group; most of the work is designing the rollout properly, not extra data collection.
“No evidence exists on task-sharing models for expanding access to RHD prevention and treatment services in limited-resource settings.” — abdullahi-2019-task-sharing-diagnosis 2019
“No studies have addressed prevention of acute rheumatic fever (primary prevention) or passive case-finding of recurrent ARF and RHD through task-sharing approaches.” — abdullahi-2019-task-sharing-diagnosis 2019
“No studies have evaluated task-sharing models for tertiary prevention, including post-operative management of patients after heart valve surgery.” — abdullahi-2019-task-sharing-diagnosis 2019
It was done once, at one site, in a small sample, and never repeatedmoderate  38 papers
What this obstacle holds upAlmost every promising finding: biomarker panels, screening accuracy, adherence patterns, cohort outcomes, all reported from a single hospital, district or school system with no independent replication.
What would break the obstacleA replication network. The cheapest version is not new data but a standing agreement among a handful of endemic-country centres to run the same protocol and pool results, which also fixes the underpowered subgroup analyses each site reports separately.
Who is positioned to do itA multi-centre clinical network; each individual replication is small enough for one site, which is exactly why nobody coordinates them.
“The predictive value of neutrophil count and NLR for RHD appears moderate and should be interpreted with caution.” — akat-2026-neutrophil-count-neutrophil 2026
“Whether platelet indices (MPV, P-LCR) are reliable indicators of inflammatory activity in children with ARF remains uncertain because the study found no significant differences.” — akat-2026-neutrophil-count-neutrophil 2026
“Only 20 of 56 ARF patients had follow-up echocardiographic data, restricting the RHD prediction analysis to a small subgroup.” — akat-2026-neutrophil-count-neutrophil 2026
The datasets already exist but were never joinedcheap  30 papers
What this obstacle holds upQuestions about who progresses, who adheres, what a programme actually cost and which genes matter, where every input has already been collected by somebody and sits in a separate file, register, cohort or supplementary table.
What would break the obstacleIndividual-level pooling and record linkage: a shared patient identifier across a country's inpatient and outpatient RHD registers, an ARF/RHD consortium that pools existing genotype and echo cohorts, meta-analyses done on individual participant data rather than published summaries, deposition of sequencing and screening datasets, and re-analysis of studies whose authors ran no multivariable model on data they already hold.
Who is positioned to do itA single analyst or small unit with data-sharing agreements: a health-ministry data office, a PhD student with an existing registry extract, or a consortium secretariat that does nothing but assemble other people's datasets.
“The reasons for heterogeneity in reported RHD prevalence across the included studies have not been fully explained despite subgroup analysis by region, country and sampling unit.” — abdu-2024-prevalence-pattern-rheumatic 2024
“No study has determined whether ARF/RHD disease registers would be more effective if integrated into general health information systems or maintained as parallel systems.” — abrams-2020-integrating-prevention-control 2020
“Whether the Iyengar India study and the WHO multi-country study included overlapping patient populations is unclear.” — abrams-2020-integrating-prevention-control 2020
Nobody ever built the model: costs, cost-effectiveness and transmission dynamics are unquantifiedcheap  28 papers
What this obstacle holds upWhether screening, point-of-care testing, penicillin delivery, surgery and a future vaccine are worth what they cost, and what a vaccine would actually do to transmission rather than to individuals.
What would break the obstacleHealth-economic and dynamic-transmission modelling, plus the few missing local inputs those models need: country value sets for quality-of-life instruments, willingness-to-pay estimates, procurement and delivery costs, and herd-effect structure in the vaccine models that currently only count direct protection.
Who is positioned to do itA single health economist or modeller with published inputs; most of these gaps are a laptop and six months, and several are explicitly flagged by the authors as things their own model omitted.
“The cost-effectiveness of WHO's recommendation to use handheld echocardiography for RHD screening in endemic areas has not been carefully evaluated.” — ali-2026-implementing-world-health 2026
“The cost-effectiveness of molecular and rapid diagnostics for S. pyogenes pharyngitis in LMIC settings has not been established.” — armitage-2025-evaluating-clinical-decision 2025
“Estimates of number needed to treat (NNT), number needed to harm (NNH), and the antimicrobial resistance implications of using molecular diagnostics for S. pyogenes pharyngitis in LMICs have” — armitage-2025-evaluating-clinical-decision 2025
Two definitions were never reconciled, so everyone measures a slightly different diseasemoderate  28 papers
What this obstacle holds upPrevalence comparisons, screening accuracy, the meaning of borderline disease, and telling acute valvulitis apart from chronic damage, because criteria changed in 2023, devices differ, and there is no gold standard for acute rheumatic fever at all.
What would break the obstacleRe-reading existing echocardiogram archives under the 2023 World Heart Federation staging, device-specific screening thresholds, agreed criteria for acute versus chronic valve morphology, and a reference standard for acute rheumatic fever that is not the criteria being tested. Most of the images already exist.
Who is positioned to do itA guideline body plus a few centres willing to re-score their stored studies; the expensive part is agreement, not data collection.
“A specified protocol to help frontline health workers discriminate ARF from its endemic mimickers (e.g., malaria, typhoid, viral infections) has not been defined.” — ali-2022-rheumatic-heart-disease 2022
“It has not been determined whether treating any joint symptom as probable ARF in endemic settings causes net harm from overdiagnosis and unnecessary BPG prophylaxis.” — ali-2022-rheumatic-heart-disease 2022
“A simplified algorithm for ARF diagnosis that incorporates handheld echocardiography has not been tested at the primary care level.” — ali-2026-implementing-world-health 2026
Nobody is paying for it, because the patients are poorexpensive  28 papers
What this obstacle holds upEverything that needs a commercial sponsor or a purchaser: a vaccine, a better or less painful penicillin, any drug that slows valve fibrosis, transcatheter valves designed for rheumatic anatomy, surgery capacity, and even a billing code for a point-of-care throat test.
What would break the obstacleMoney with a name on it: push funding or advance market commitments for Strep A products, a product-development partnership for rheumatic-specific valves and long-acting penicillin, national reimbursement listings, and explicit correction of the funding inequity the authors themselves flag, where the research follows the wealthy country rather than the disease.
Who is positioned to do itFunders and industry: a global health financing body, a product-development partnership, or a ministry willing to reimburse. Researchers can document the gap but cannot close it.
“There is currently no ideal alternative drug to benzathine penicillin G for treating GAS or providing secondary prophylaxis for ARF.” — ali-2022-rheumatic-heart-disease 2022
“An antibiotic substitute for BPG with better administration techniques and a safer profile has not been developed.” — ali-2026-implementing-world-health 2026
“The molecular point-of-care test for group A streptococcus (sensitivity 100%, specificity 79%), although documented in Australia, has not been evaluated for availability and affordability in” — ali-2026-implementing-world-health 2026
The thing cannot be observed: there is no testexpensive  25 papers
What this obstacle holds upDiagnosing acute rheumatic fever, staging or monitoring subclinical disease, deciding who progresses, and judging whether any vaccine works, because there is no biomarker, no correlate of protection and no affordable accurate throat test.
What would break the obstacleBiomarker discovery carried through to a deployable assay: external validation of candidate proteomic signatures in diverse cohorts, internationally standardised and calibrated immunoassays, an affordable molecular or antigen test that performs in a rural clinic, and validated artificial-intelligence reading of echocardiograms.
Who is positioned to do itA diagnostics laboratory or industry partner working with endemic-country cohorts; discovery is done, translation is what is missing.
“Evidence regarding the relationship between platelet indices, NLR, and the development of RHD in children with ARF remains limited.” — akat-2026-neutrophil-count-neutrophil 2026
“The accuracy of a deep learning algorithm trained on DAS waveform data to classify subclinical RHD into definite, borderline, and normal categories has not been tested.” — ali-2021-detection-subclinical-rheumatic 2021
“It is not yet known whether a DAS-based DL algorithm can detect subclinical RHD without the need for echocardiography.” — ali-2021-detection-subclinical-rheumatic 2021
The long follow-up was never startedneeds-a-decade  23 papers
What this obstacle holds upNatural history: what happens to borderline and latent disease over ten years, which acute cases become chronic, how repaired and replaced valves last, and what happens to women with RHD after pregnancy.
What would break the obstacleCohorts that must be opened now and read much later, or the cheaper substitute of retrospectively assembling follow-up from registers that have been running for decades in Australia, New Zealand, Uganda and Nepal.
Who is positioned to do itA long-lived clinical network with stable funding, or an analyst willing to mine a registry that has already accumulated twenty years of follow-up nobody has looked at.
“Long-term follow-up of Africans diagnosed with borderline RHD to determine progression to definite RHD has not been performed in Africa; existing progression data come from South Africa and ” — abdu-2024-prevalence-pattern-rheumatic 2024
“Prospective studies are needed to confirm whether neutrophil count and NLR predict development of RHD in children with ARF.” — akat-2026-neutrophil-count-neutrophil 2026
“Recurrent RF episodes and ongoing rheumatic activity place mitral valve repair at high risk of recurrent regurgitation or stenosis, but the magnitude and predictors of this failure are not c” — antunes-2020-global-burden-rheumatic 2020
The biology is genuinely unknown and no shortcut existsexpensive  16 papers
What this obstacle holds upWhy a throat infection becomes autoimmune valve disease, which antigens and post-translational modifications drive cross-reactivity, and why the mitral valve and female sex are preferentially hit.
What would break the obstacleSustained immunology and matrix-biology programmes, single-cell and multi-omic work on human valve material, and characterisation of the autoantibodies and modified self-epitopes that are still hypothesised rather than demonstrated.
Who is positioned to do itA funded immunology or cardiovascular research laboratory, ideally paired with a surgical centre that can supply fresh human valve tissue.
“The pathophysiology of rheumatic heart disease, including why some cases evolve rapidly with multivalve involvement while others remain chronic, is still not well understood.” — antunes-2020-global-burden-rheumatic 2020
“It remains unknown why rheumatic valve disease progresses predominantly to insufficiency in some patients and to stenosis in others.” — antunes-2020-global-burden-rheumatic 2020
“Why pure mitral stenosis caused by commissural fusion with nearly normal leaflets can appear in patients in the second and third decade of life is unexplained, despite the dogma that mitral ” — antunes-2020-global-burden-rheumatic 2020
You cannot run the experiment: it means withholding penicillin, or risking autoimmunityexpensive  16 papers
What this obstacle holds upWhether prophylaxis actually works and for how long, whether screening changes outcomes, whether antibiotics for skin infection prevent rheumatic fever, and whether immune-modulating drugs or a vaccine can be given to children at all.
What would break the obstacleDesigns that dodge equipoise: non-inferiority and de-escalation trials such as oral versus injected penicillin or stopping rules after normalisation, stepped-wedge rollout of screening so nobody is denied it, natural experiments where a service already changed, and a settled safety framework for vaccine-induced cross-reactivity so trials can be designed rather than feared.
Who is positioned to do itA trial group with ethics-committee and community backing, ideally embedded in a programme that is expanding anyway so the comparison arises from rollout order rather than denial.
“It is unknown whether echocardiography-based active case finding improves outcomes for RHD and how to manage borderline RHD detected through screening.” — abdullahi-2019-task-sharing-diagnosis 2019
“The role of antibiotic prophylaxis in halting the progression of subclinical RHD had not been confirmed until the 2022 Ugandan RCT.” — ali-2022-rheumatic-heart-disease 2022
“Whether treating all PCR-positive cases of pharyngitis in high-RHD-risk LMIC settings would be beneficial has not been evaluated.” — armitage-2025-evaluating-clinical-decision 2025
Nobody asked the patients, families, providers or decision-makersmoderate  15 papers
What this obstacle holds upWhy people present late, why they stop coming, what an injection every month actually costs a family in pain and dignity, what providers fear, and what makes a ministry adopt a vaccine.
What would break the obstacleQualitative and mixed-methods work at a single site: interviews, focus groups, acceptability studies, and community-governed research design. Several papers explicitly collected the provider view and never the patient view, or vice versa.
Who is positioned to do itA single social-science or nursing researcher with local partners; one of the few themes where an individual can close a gap in a year.
“Local socio-cultural contexts have not been seriously considered in the design of most RHD prevention and control interventions.” — abrams-2020-integrating-prevention-control 2020
“Indigenous-led governance models ensuring cultural safety and community ownership of Strep A POCT research have not yet been established.” — barth-2022-roadmap-incorporating-group 2022
“Whether the informal written consent procedures used by HCPs in some Ethiopian centres reduce HCP anxiety without further undermining patient adherence to BPG.” — birru-2025-acceptability-implementation-challenges 2025
Nothing was ever compared head to headmoderate  15 papers
What this obstacle holds upWhich handheld device, which penicillin formulation, which valve, which vaccine candidate, which surgical timing, which education format: options proliferate and are each tested alone against nothing.
What would break the obstacleComparative studies: device-versus-device and assay-versus-assay evaluations on the same subjects, randomised comparisons of delivery formats, and risk-stratified comparisons of tissue against mechanical valves in young patients, which the field keeps calling for and never funds.
Who is positioned to do itA clinical network or an evaluation body with no stake in the products; some comparisons need only stored samples or images.
“No randomised or controlled study has compared different screening strategies (notably clinical evaluation alone versus echocardiography) within African populations, so the yield attributabl” — abdu-2024-prevalence-pattern-rheumatic 2024
“Whether complete rigid annuloplasty rings are essential, or whether flexible/incomplete bands give comparable durability in rheumatic mitral regurgitation, remains unsettled.” — antunes-2020-global-burden-rheumatic 2020
“The performance of Strep A POCT by trained operators has not yet been shown to have high concordance with gold standard conventional laboratory techniques (in Australia).” — barth-2022-roadmap-incorporating-group 2022
A finding repeated in every paper that nobody has ever chasedcheap  14 papers
What this obstacle holds upThe two-to-one female excess, unexplained ethnic and regional differences in prevalence, prophylaxis and anticoagulation control, and the heterogeneity that every meta-analysis reports and then leaves as a limitation.
What would break the obstacleDeliberate secondary analysis aimed at the anomaly instead of around it: sex-stratified pooling of existing screening datasets, meta-regression on data already extracted, and hypothesis-driven work on sex-dependent immune mechanisms nobody has tested.
Who is positioned to do itA single analyst with published or pooled datasets; this is the clearest opportunity in the whole corpus because the observation is decades old and the follow-up has simply never been done.
“The reasons why male and female RHD prevalence were nearly identical in this pooled analysis have not been determined.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The reasons for the higher RHD prevalence among females remain not well understood, and the contribution of socio-cultural health-seeking differences versus biological factors is unresolved.” — chillo-2023-sub-clinical-rheumatic 2023
“Why socio-economic factors (maternal education, distance to facility, household size, rural vs semi-urban residence) previously associated with RHD showed no association in this population i” — chillo-2023-sub-clinical-rheumatic 2023
The evidence exists but the reviews cannot see itcheap  13 papers
What this obstacle holds upEvery pooled prevalence estimate and systematic review conclusion in the corpus, because the searches excluded non-English publications, grey literature and unpublished work in exactly the Francophone, Lusophone, Arabophone and Latin American regions where the disease is concentrated.
What would break the obstacleRedoing the reviews without the language and indexing filters: multilingual search and translation, regional and grey-literature databases, thesis repositories, and formal small-study and publication-bias assessment where enough studies exist.
Who is positioned to do itA single reviewer with translation support, or a regional research network that knows where its own literature is published. Machine translation makes this cheaper now than when the reviews were written.
“Non-English-language population-based studies on RHD prevalence in Africa were not captured, which may bias the pooled estimate.” — abdu-2024-prevalence-pattern-rheumatic 2024
“English-language and date restrictions on searches may have excluded relevant publications, and the true evidence base on integrated RHD programmes is therefore uncertain.” — abrams-2020-integrating-prevention-control 2020
“Only English-language studies with ≥80% adherence as the definition of good adherence were included, so evidence in other languages or using alternative cut-offs is missing.” — bimerew-2024-adherence-secondary-antibiotic 2024
No laboratory or animal model reproduces the diseaseexpensive  13 papers
What this obstacle holds upVaccine efficacy prediction, antifibrotic drug testing, immunomodulation, and any mechanistic claim that needs to move from a mouse or rat to a human.
What would break the obstacleBetter preclinical systems: validated and standardised animal models, non-human primate infection models, patient-derived valve organoids and ex vivo systems that recreate valve haemodynamics, and a controlled human infection model where ethics permit.
Who is positioned to do itA specialist laboratory with animal-facility or organoid capability, typically funded as vaccine-development infrastructure rather than as RHD research.
“Functional consequences of the PITX2 exon 4 rare missense variants identified remain unestablished.” — amath-2026-genetic-variability-amp 2026
“Whether the rare coding variants in PITX2 exon 4 contribute mechanistically to atrial fibrillation pathophysiology has not been experimentally tested.” — amath-2026-genetic-variability-amp 2026
“It remains unclear whether anti-SodA antibodies truly fail to confer protection during natural infection, since this conclusion was drawn solely from rodent studies.” — azuar-2019-recent-advances-development 2019
The household and social causes are asserted in every paper and measured in nonemoderate  9 papers
What this obstacle holds upCrowding, damp, ventilation, dust, cleaning, skin infection and sugary drinks as causes of Strep A transmission and rheumatic fever, and whether primordial prevention does anything.
What would break the obstacleActual measurement in homes: baseline ventilation and humidity, surface and air sampling with genomics, and community-led trials of housing and hygiene interventions with Strep A outcomes rather than proxy indicators. Also an evaluation tool that counts what affected communities count as health.
Who is positioned to do itAn environmental-health team working with an Indigenous or community-controlled organisation; the measurements are cheap, the partnership is what takes time.
“Whether reducing sugar-sweetened drink intake could lower ARF/RHD risk in children is unknown and requires investigation.” — baker-2022-risk-factors-acute 2022
“The role of specific housing factors (damp, mould, cold, hot water supply, bed sharing, functional crowding) as ARF risk factors remains to be investigated more rigorously.” — baker-2022-risk-factors-acute 2022
“The contribution of impetigo (as distinct from pharyngitis) to ARF risk in Australia has not been quantified, despite skin infection being a noted driver.” — barth-2022-roadmap-incorporating-group 2022
Obstacle not identifiablemoderate  8 papers
What this obstacle holds upGeneric bias caveats: recall bias, social desirability, residual confounding, unblinded interpretation, stated as limitations without a specific obstacle behind them.
What would break the obstacleNothing specific. These are the boilerplate limitations sections rather than named unsolved problems.
Who is positioned to do itNot applicable.
“Possible residual confounding from unmeasured or inaccurately measured exposures cannot be ruled out.” — baker-2022-risk-factors-acute 2022
“Recall bias is possible because cases were interviewed shortly after illness onset while controls did not have a similar recall stimulus.” — baker-2022-risk-factors-acute 2022
“Social desirability bias may have caused differential reporting between cases and controls on sensitive exposures such as smoking indoors.” — baker-2022-risk-factors-acute 2022
The diseased tissue is out of reachexpensive  5 papers
What this obstacle holds upEverything about what the valve is actually doing: the transcriptional programmes of human rheumatic valve cells, calcium distribution, bacterial persistence, and whether blood markers reflect the valve at all.
What would break the obstacleTissue biobanking attached to cardiac surgery in endemic countries, with standardised sampling so pieces can be compared, plus paediatric valve material which is scarce because children rarely reach surgery.
Who is positioned to do itA surgical centre in a high-burden country partnered with a molecular pathology laboratory; the material passes through operating theatres every week and is mostly discarded.
“Whether the epithelial-mesenchymal transition signature observed in peripheral plasma reflects actual valvular tissue damage and repair in ARF has not been directly demonstrated.” — carapetis-2025-plasma-protein-biomarker 2025
“Adequate in-depth analysis of heart valve tissue from ARF/RHD patients has not been performed.” — middleton-2026-cxcr-associated-cell 2026
“Only a small number of rheumatic valves were analysed in this study, and broader sampling is limited by the rarity of paediatric valve surgery.” — middleton-2026-cxcr-associated-cell 2026

What it is about

The thing in the world that is not understood.

Study-design limits that recur across this whole literaturecrowded  27 papers
What nobody has establishedA large share of the field's stated gaps are not about the disease at all: single-centre retrospective designs, samples of a few dozen, convenience and non-random sampling, recall and self-report bias, non-concurrent controls, unblinded reading, and results from one population that nobody has shown transfer to another.
Why it is still openSmall single-site studies are what one clinician with no research budget can produce, and in a neglected disease that is most of the output.
The study that would close itNot a study: this is the field's structural condition, and it is what multi-site prospective designs with pre-registered analysis would fix.
“Only 20 of 56 ARF patients had follow-up echocardiographic data, restricting the RHD prediction analysis to a small subgroup.” — akat-2026-neutrophil-count-neutrophil 2026
“The retrospective, single-center design limits validation of admission neutrophil count and NLR as predictors of RHD.” — akat-2026-neutrophil-count-neutrophil 2026
“Whether the moderate predictive performance of neutrophil count (AUC 0.72) and NLR (AUC 0.68) holds in different populations and healthcare settings is unknown.” — akat-2026-neutrophil-count-neutrophil 2026
Prevalence data simply do not exist for most countriescrowded  26 papers
What nobody has establishedFor most of the world there is no population-based estimate of ARF incidence or RHD prevalence: whole regions (North Africa, Central Africa, Southeast Asia, Latin America outside Brazil, Europe outside Turkey) rest on one or two studies or none, and only about ten countries have prospective ARF incidence data.
Why it is still openPrevalence surveys are individually fundable but nobody coordinates them, so effort piles onto countries that already have data and skips the ones that do not.
The study that would close itStandardised population-based echo prevalence surveys with linked ARF case-finding in a set of deliberately chosen data-void countries, using one protocol.
“No population-based data on rheumatic heart disease prevalence from North Africa exists beyond a single Egyptian school study.” — abdu-2024-prevalence-pattern-rheumatic 2024
“No literature exists on the implementation of handheld echocardiography screening for subclinical RHD in Pakistan.” — ali-2021-detection-subclinical-rheumatic 2021
“The full burden of GAS on Africa is unclear because surveillance is limited.” — aliyu-2024-rheumatic-heart-disease 2024
Language, grey-literature and publication-bias limits on the whole RHD evidence basecrowded  12 papers
What nobody has establishedSystematic reviews in this field almost uniformly exclude non-English and grey literature, find funnel-plot asymmetry, cannot do subgroup analyses, and pool countries into income groups so country data cannot be recovered.
Why it is still openTranslation costs money and reviewers write the limitation instead; the excluded literature is from precisely the highest-burden Francophone and Lusophone countries.
The study that would close itA multilingual systematic review with grey-literature searching in French, Portuguese, Arabic and Spanish for RHD prevalence in Africa and Latin America.
“Non-English-language population-based studies on RHD prevalence in Africa were not captured, which may bias the pooled estimate.” — abdu-2024-prevalence-pattern-rheumatic 2024
“Whether the Iyengar India study and the WHO multi-country study included overlapping patient populations is unclear.” — abrams-2020-integrating-prevention-control 2020
“English-language and date restrictions on searches may have excluded relevant publications, and the true evidence base on integrated RHD programmes is therefore uncertain.” — abrams-2020-integrating-prevention-control 2020
What echocardiographic threshold should define a positive RHD screencrowded  10 papers
What nobody has establishedThe criteria have no diagnostic gold standard behind them, jet-length cut-offs have never been validated as pathology-versus-physiology, and simplified protocols may miss morphological disease while confirmatory criteria cannot be applied where the disease is.
Why it is still openThere is no tissue or outcome reference standard available in living screened children, so every validation is circular; the field has been refining criteria against each other since 2012.
The study that would close itA pathology- or outcome-anchored validation study that fixes each WHF criterion against long-term progression rather than against another echo reading.
“Whether simplified echocardiographic scoring systems can be used to risk-stratify and target BPG treatment in Malawi's resource-limited settings has not been validated locally.” — blennerhassett-2025-burden-rheumatic-heart 2025
“It is worth noting that the criteria used to define mild MR are not specified in the Kotit guideline.” — fernandezdetoro-2026-comparison-different-criteria 2026
“There is no consensus on the criteria that define a positive RHD screen, and no device-specific criteria for handheld ultrasound.” — kaltenborn-2022-assessing-performance-characteristics 2022
A drug that halts or reverses RHD once it has startedcrowded  9 papers
What nobody has establishedThere is no proven anti-inflammatory, immunosuppressive or antifibrotic therapy for RHD; candidate targets (TNF, IL-6, CXCR3, Treg restoration via low-dose IL-2, VEGFR3, TGF-beta/STAT3) have never entered a trial, and approved antifibrotics cannot reach avascular valve tissue.
Why it is still openTrialling a biologic in children in low-income settings runs into safety, cost and regulatory walls at once, and the outcome measure (echo progression) is itself contested.
The study that would close itA phase 2 trial of an existing licensed immunomodulator in children with progressive latent or early RHD, with echocardiographic progression at 18 months as the endpoint.
“Can tumor necrosis factor-alpha blockade safely prevent progression of rheumatic heart disease?” — branco-2026-elevated-interleukin-levels 2026
“Are anti-interleukin-6 therapies safe and effective in children with subclinical rheumatic heart disease?” — branco-2026-elevated-interleukin-levels 2026
“What is the long-term safety of cytokine-targeted biological interventions in children?” — branco-2026-elevated-interleukin-levels 2026
What predicts which ARF cases go on to chronic RHDcrowded  9 papers
What nobody has establishedThe transition from acute episode to permanent valve disease has no established predictor, no measured time-dependency, and is left out of every impact model because the natural history is not quantified.
Why it is still openIt requires catching people at their first ARF episode, which is precisely the event that goes undiagnosed in the settings where it is common.
The study that would close itA prospective inception cohort of first-episode ARF with baseline immune and echo phenotyping, followed for five years to valve outcome.
“Prospective studies are needed to confirm whether neutrophil count and NLR predict development of RHD in children with ARF.” — akat-2026-neutrophil-count-neutrophil 2026
“The predictive value of neutrophil count and NLR for RHD appears moderate and should be interpreted with caution.” — akat-2026-neutrophil-count-neutrophil 2026
“Whether platelet indices (MPV, P-LCR) are reliable indicators of inflammatory activity in children with ARF remains uncertain because the study found no significant differences.” — akat-2026-neutrophil-count-neutrophil 2026
GAS vaccine candidates stalled between preclinical promise and clinical datacrowded  8 papers
What nobody has establishedMultiple candidates have phase 1 safety data and then nothing; no head-to-head comparison exists, no non-M candidate has reached the clinic, and no glycoconjugate has been tested in humans.
Why it is still openThe market is in countries that cannot pay, so candidates park after phase 1; no commercial sponsor will fund a comparison that could rank their own candidate second.
The study that would close itA funder-convened head-to-head immunogenicity comparison of the leading candidates in a single trial with a common assay and a common endpoint.
“No licensed, effective vaccine against GAS exists to control RHD in Africa.” — aliyu-2024-rheumatic-heart-disease 2024
“Large-scale clinical trials are necessary to verify the complete safety of the 6-valent GAS vaccine, since the Phase 1 findings were argued to be biased by the small-scale, open-label (non-b” — azuar-2019-recent-advances-development 2019
“It is unknown whether the 30-valent M-protein-based vaccine, whose Phase 1 trial was planned for 2015, will proceed, since no further information is available regarding this study.” — azuar-2019-recent-advances-development 2019
What training a non-expert needs before their echo screen can be trustedcrowded  8 papers
What nobody has establishedThere is no consensus curriculum, no minimum number of supervised scans, and no measured sensitivity for screeners trained under the one-to-two-week regimens actually used.
Why it is still openProgrammes train to whatever their timeline allows and publish the yield, not the operator's error rate; measuring the operator requires an expert gold standard on every scan.
The study that would close itA trial randomising nurses or clinical officers to short versus extended structured training, with each trainee's scans compared against expert-read standard echo on the same children.
“The very low sensitivity (6.3%) of auscultation performed by medical students versus 16.4% (Uganda) and 10% (Mozambique) reported by others has not been studied; whether this reflects operat” — chillo-2023-sub-clinical-rheumatic 2023
“The impact of longer, specific, and practical training for nonspecialist screening personnel on screening outcomes remains unclear, especially at larger scale.” — diniz-2024-agreement-between-handheld 2024
“The sensitivity and specificity of echo screening performed by medical students and fresh graduates after only one to two weeks of training have not been formally evaluated.” — elazrag-2022-handheld-echocardiographic-screening 2022
RHD in adults and in people who are not in school is essentially unmeasuredcrowded  8 papers
What nobody has establishedAlmost all screening evidence comes from schoolchildren; adult prevalence, adult-specific echo criteria, performance of handheld echo and of diagnostic algorithms in adults, older adults and pregnant women are all unestablished, even though adults carry the clinical burden.
Why it is still openSchools are a free sampling frame and adults are not; the criteria were built in children and revalidating them in adults means starting the whole diagnostic argument again.
The study that would close itA community-based, all-age echo prevalence survey with adult-specific criteria development and validation against clinical outcomes.
“Rheumatic heart disease in adults and non-school-attending populations across Africa is largely unmeasured, since most included studies drew samples from school lists.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The proteomic signature was identified in a pediatric Ugandan cohort (median age 9 years) and has not been tested in adults or other age groups.” — carapetis-2025-plasma-protein-biomarker 2025
“Epidemiological data on RHD in young adults and adults in LAC are essentially missing, despite this being the peak age group worldwide.” — jaimesreyes-2022-current-situation-acute 2022
The immune mechanism that turns a strep throat into acute rheumatic fevercrowded  7 papers
What nobody has establishedMolecular mimicry is assumed rather than demonstrated: the pathogenic autoantibodies are uncharacterised, the striking IgG3 polarisation is unexplained, and whether tolerance is lost centrally or peripherally is unknown.
Why it is still openThe disease is rare where the labs are and common where they are not; the acute window is short and the patients are children.
The study that would close itDeep immune phenotyping (B-cell repertoire, autoantibody profiling, T-cell specificity) of paired first-episode ARF cases and GAS-pharyngitis controls sampled concurrently.
“What immunological mechanisms drive the development and progression of rheumatic heart disease?” — branco-2026-elevated-interleukin-levels 2026
“The functional source and role of elevated plasma PACSIN2 in ARF is not understood.” — carapetis-2025-plasma-protein-biomarker 2025
“Whether IgG boosting may be greater during toxin-mediated or invasive S. pyogenes infections and in acute rheumatic fever has not been examined.” — keeley-2025-early-life-serological 2025
National ARF/RHD control programmes: whether they work, what they cost, whether they lastcrowded  7 papers
What nobody has establishedThe comprehensive programmes people cite (Nepal, Sudan) were never rigorously evaluated, their effect cannot be separated from concurrent economic development, implementation costs are unmeasured, and transferability is unproven.
Why it is still openProgrammes are political commitments; evaluating one properly means being able to report that it failed, which no ministry commissions.
The study that would close itA prospective evaluation of a national programme with pre-specified implementation and effectiveness indicators, an economic evaluation, and a comparison district.
“The current national ARF/RHD control strategy and its implementation status in Ethiopia is not well documented.” — birru-2025-acceptability-implementation-challenges 2025
“Echocardiographic screening programs for RHD similar to PROVAR have not been launched in most LAC countries.” — jaimesreyes-2022-current-situation-acute 2022
“The clinical effectiveness of ARF case-finding and its impact on RHD outcomes has not been evaluated.” — klassen-2026-implementing-national-acute 2026
Whether borderline RHD seen on echo in a child with no ARF history actually progressescontested  7 papers
What nobody has establishedNobody has established what fraction of screen-detected borderline RHD progresses to definite disease, over what timescale, or which children are the ones who progress.
Why it is still openIt needs a decade of follow-up on healthy-looking children in exactly the places with the least capacity to keep a cohort intact; funders buy prevalence surveys because they finish.
The study that would close itA multi-country cohort of borderline-RHD children identified by school screening, followed with annual echo for ten years against a matched screen-negative arm.
“Long-term follow-up of Africans diagnosed with borderline RHD to determine progression to definite RHD has not been performed in Africa; existing progression data come from South Africa and ” — abdu-2024-prevalence-pattern-rheumatic 2024
“The other factors (besides prophylaxis) that determine regression or progression of sub-clinical RHD are unknown and require further research including genetic studies.” — chillo-2023-sub-clinical-rheumatic 2023
“The echocardiographic manifestations of subclinical RHD, the progression of disease and possible regional differences remain poorly defined.” — fernandezdetoro-2026-comparison-different-criteria 2026
Cost-effectiveness and sustainability of echo screening programmesthin  7 papers
What nobody has establishedNobody has costed a screening programme end to end, including confirmatory echo, referral, prophylaxis and the cost of false positives, in the settings the WHO recommends it for.
Why it is still openHealth economists need programme cost data that programmes do not collect, and the effectiveness denominator is itself unresolved (see the screening-outcomes theme).
The study that would close itA prospective micro-costing of an existing national screening programme with a decision-analytic model calibrated to its own referral and adherence data.
“The cost-effectiveness of WHO's recommendation to use handheld echocardiography for RHD screening in endemic areas has not been carefully evaluated.” — ali-2026-implementing-world-health 2026
“The cost and cost-effectiveness of primary care echocardiographic screening for RHD have not been measured.” — jiee-2024-heart-community-implementation 2024
“Whether echocardiographic screening for RHD by primary care doctors is sustainable beyond the pilot period is untested.” — jiee-2024-heart-community-implementation 2024
School-based sampling versus true community prevalencecrowded  7 papers
What nobody has establishedPrevalence is measured where people are easy to find (schools, urban hospitals) and it is unknown how far that departs from the community rate, including for the rural majority and for children who do not attend school or do not return for confirmatory echo.
Why it is still openHousehold sampling costs several times more per case found, and every programme is judged on cases found.
The study that would close itA paired study measuring RHD prevalence in the same district by school sampling and by random household sampling.
“Whether the children who did not attend for standard echocardiography had RHD, and the effect on the final echo prevalence, is not known.” — ali-2018-handheld-echocardiography-screening 2018
“The true community-level prevalence of RHD in Al Managil is unknown because screening was conducted in schools rather than the community.” — ali-2018-handheld-echocardiography-screening 2018
“The rural prevalence of RHD in Malawi could not be determined because country-specific rural data were pooled across low-income countries.” — blennerhassett-2025-burden-rheumatic-heart 2025
Why RHD affects women roughly twice as often as mencontested  7 papers
What nobody has establishedThe female excess is reported everywhere and explained nowhere; whether it is care-seeking, exposure, pregnancy, or sex-linked biology is undetermined, and even the age at which the gap opens is unknown.
Why it is still openIt falls between epidemiology and biology, and every paper that notices it does so in a limitations paragraph rather than as a research question.
The study that would close itA prospective screening cohort stratified by sex and pubertal stage with parallel measurement of exposure, care-seeking and sex-hormone/X-linked candidate pathways.
“The reasons why male and female RHD prevalence were nearly identical in this pooled analysis have not been determined.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The reasons for the higher RHD prevalence among females remain not well understood, and the contribution of socio-cultural health-seeking differences versus biological factors is unresolved.” — chillo-2023-sub-clinical-rheumatic 2023
“The reasons for the high female-to-male ratio among RHD cases and the absence of prior symptom history in many patients are not fully understood.” — elazrag-2022-handheld-echocardiographic-screening 2022
Social determinants of ARF and RHD outcomes are asserted but never measuredcontested  7 papers
What nobody has establishedPoverty, crowding, education, ethnicity and access are invoked as explanations and then not collected: studies repeatedly report that social determinant data were unavailable, or that expected associations did not appear and cannot be explained.
Why it is still openSocial variables are collected badly or not at all in clinical datasets, and by the time the analysis notices, the cohort is closed.
The study that would close itA prospective cohort with individual-level socioeconomic, housing and health-access measurement linked to RHD incidence and outcome.
“It is not known whether and how RHD outcomes differ by social determinants in African populations.” — aliyu-2024-rheumatic-heart-disease 2024
“Whether high-risk children (particularly Māori and Pacific) were actually treated with recommended first-line antibiotics remains unclear because of the supply-order data gap and known ethni” — cannon-2023-trends-penicillin-dispensing 2023
“Why socio-economic factors (maternal education, distance to facility, household size, rural vs semi-urban residence) previously associated with RHD showed no association in this population i” — chillo-2023-sub-clinical-rheumatic 2023
Whether echo screening programmes change outcomes rather than just find casesthin  7 papers
What nobody has establishedNo study has compared a screened population with an unscreened one on death, heart failure, surgery or ARF recurrence.
Why it is still openThe outcome takes 10-15 years to accrue, the control arm looks like withholding a known-good test, and screening programmes are funded as service delivery rather than as trials.
The study that would close itA cluster-randomised trial of school echo screening plus treat-on-detect versus usual care in an endemic district, followed for RHD-related hospitalisation and death.
“No randomised or controlled study has compared different screening strategies (notably clinical evaluation alone versus echocardiography) within African populations, so the yield attributabl” — abdu-2024-prevalence-pattern-rheumatic 2024
“It is unknown whether echocardiography-based active case finding improves outcomes for RHD and how to manage borderline RHD detected through screening.” — abdullahi-2019-task-sharing-diagnosis 2019
“Whether a screen-to-treat policy implemented by the Ministry of Health in Kordofan will reduce morbidity and mortality from RHD has not been evaluated.” — elazrag-2022-handheld-echocardiographic-screening 2022
How accurate handheld echo is against standard echo in a real screening settingcrowded  6 papers
What nobody has establishedThere is no properly designed diagnostic-accuracy study of handheld versus standard echo for RHD, no head-to-head device comparison, and no evidence that criteria validated on one device transfer to another.
Why it is still openIt requires bringing full-size machines to every child including the screen-negatives, which doubles the cost of the field work and yields a negative-heavy dataset nobody wants to fund.
The study that would close itA paired diagnostic-accuracy study in which every screened child receives handheld and standard echo, blinded, across at least three device models and two operator tiers.
“The feasibility of implementing handheld echo screening in rural endemic areas, including machine availability, training of rural health workers, and verification of results by specialists, ” — ali-2026-implementing-world-health 2026
“The hand-held echocardiogram cannot perform continuous Doppler, which may have caused up to ~13% false-positive sub-clinical RHD diagnoses, and the true prevalence using this device is uncer” — chillo-2023-sub-clinical-rheumatic 2023
“False-negative screening studies were not verified, so accuracy conclusions remain limited because only screen-positive children were referred for standard echocardiography.” — diniz-2024-agreement-between-handheld 2024
What drives progressive valve fibrosis long after the acute episodecontested  6 papers
What nobody has establishedThe mechanisms of chronic fibrosis are unresolved: whether immune activity continues, why the mitral valve is preferentially hit, why some valves go to stenosis and others to regurgitation, and what role EndMT, VIC subtypes and shear stress play.
Why it is still openHuman valve tissue is only available at surgery, which means late-stage disease in patients who reached a surgical centre, so the early mechanism is never sampled.
The study that would close itSingle-cell and spatial transcriptomics on explanted rheumatic valves stratified by lesion type (stenotic versus regurgitant), correlated with regional mechanical loading.
“The pathophysiology of rheumatic heart disease, including why some cases evolve rapidly with multivalve involvement while others remain chronic, is still not well understood.” — antunes-2020-global-burden-rheumatic 2020
“It remains unknown why rheumatic valve disease progresses predominantly to insufficiency in some patients and to stenosis in others.” — antunes-2020-global-burden-rheumatic 2020
“Why pure mitral stenosis caused by commissural fusion with nearly normal leaflets can appear in patients in the second and third decade of life is unexplained, despite the dogma that mitral ” — antunes-2020-global-burden-rheumatic 2020
A blood biomarker that diagnoses acute rheumatic fevercrowded  6 papers
What nobody has establishedCandidate signatures (a five-protein SomaScan panel, IgG3, a CXCR3/CCL5 phenotype) exist from single cohorts, none has been externally validated, none discriminates first from recurrent episodes, and none has been built into a usable assay.
Why it is still openValidation is blocked by the same missing gold standard the biomarker is meant to replace, and the discovery cohorts are small and geographically narrow.
The study that would close itA prospective multi-site validation of the existing candidate panels in children presenting with possible ARF across at least four endemic countries, with a point-of-care prototype as the endpoint.
“The 5-protein ARF biomarker signature has not been replicated in a diverse global cohort outside Uganda.” — carapetis-2025-plasma-protein-biomarker 2025
“The small sample size precluded investigation of biomarker signatures unique to each clinical manifestation of ARF, such as carditis versus chorea.” — carapetis-2025-plasma-protein-biomarker 2025
“A wider spectrum of ARF and RHD patients with differential outcomes over time and across the RHD severity spectrum has not been recruited and analyzed.” — carapetis-2025-plasma-protein-biomarker 2025
RHD registers and surveillance systems that do not exist, do not link, or do not workcrowded  6 papers
What nobody has establishedRegisters are absent in most countries, notification is not mandatory anywhere in Latin America, existing registers cannot be linked for lack of unique identifiers, patient-level fields are largely empty, and it is unknown whether registers should sit inside or beside general health information systems.
Why it is still openRegistries are infrastructure, not research; they get built as project deliverables and decay when the project ends.
The study that would close itA feasibility and data-quality evaluation of a linked electronic ARF/RHD register at sub-national scale, measuring completeness against an independent case-ascertainment audit.
“No study has determined whether ARF/RHD disease registers would be more effective if integrated into general health information systems or maintained as parallel systems.” — abrams-2020-integrating-prevention-control 2020
“It is not known how RHD information systems should interact with the rest of the health system.” — abrams-2020-integrating-prevention-control 2020
“Registers aimed at systematically collecting morbidity and mortality outcomes of RHD patients have not been developed in LAC.” — jaimesreyes-2022-current-situation-acute 2022
Which antigens give a GAS vaccine broad enough strain coveragecrowded  6 papers
What nobody has establishedNo M protein is universally conserved, multivalent candidates were built on North American and European strain distributions, cross-protection within emm clusters is unquantified, and the biology of the carbohydrate antigens is only partly characterised.
Why it is still openStrain surveillance in the highest-burden countries barely exists, so vaccine coverage is designed against the epidemiology of the places least likely to use it.
The study that would close itGlobal emm and emm-cluster surveillance in high-burden regions paired with cross-opsonisation testing of candidate vaccine sera against locally circulating strains.
“Extensive epidemiological investigations are still needed to identify emerging emm types and the factors that affect their distribution in order to design future multivalent vaccines with be” — azuar-2019-recent-advances-development 2019
“It is not known whether the multivalent N-terminal vaccines developed from US/Canadian epidemiology will be effective in other parts of the world with very different GAS strain distributions” — azuar-2019-recent-advances-development 2019
“The biological function of the GAS cell wall carbohydrate (GAC) antigen is still not clear, despite its consideration as a universal vaccine antigen.” — azuar-2019-recent-advances-development 2019
An affordable point-of-care test for GAS pharyngitiscrowded  6 papers
What nobody has establishedNo reliable and affordable GAS test exists for the settings that need it; rapid antigen sensitivity is constrained, molecular tests are unaffordable, and it is unresolved whether PCR-positive culture-negative cases should be treated at all.
Why it is still openDiagnostics developers price for high-income respiratory markets; the carriage-versus-infection question makes the accuracy target itself ambiguous.
The study that would close itA field evaluation of low-cost molecular and next-generation antigen tests in an endemic primary-care setting, with serology to separate true infection from carriage and an NNT/NNH and AMR analysis attached.
“No reliable and affordable test for GAS pharyngitis currently exists to replace clinical algorithms.” — ali-2022-rheumatic-heart-disease 2022
“A reliable and affordable diagnostic test for group A streptococcal pharyngitis that is available and affordable in limited-resource settings has not been developed.” — ali-2026-implementing-world-health 2026
“The molecular point-of-care test for group A streptococcus (sensitivity 100%, specificity 79%), although documented in Australia, has not been evaluated for availability and affordability in” — ali-2026-implementing-world-health 2026
Whether penicillin prophylaxis changes the course of latent or subclinical RHDcontested  6 papers
What nobody has establishedOne trial (GOAL) showed benefit for definite latent RHD, but the effect in borderline disease, the harms of years of injections in well children, and whether guidelines should follow are all unsettled.
Why it is still openRandomising well children to no treatment after GOAL is now ethically contested, and the borderline group is exactly the one where equipoise is loudest.
The study that would close itA randomised trial of BPG versus no prophylaxis restricted to borderline latent RHD, powered for echocardiographic progression at three years, with adverse events and psychosocial harm as co-primary outcomes.
“The role of antibiotic prophylaxis in halting the progression of subclinical RHD had not been confirmed until the 2022 Ugandan RCT.” — ali-2022-rheumatic-heart-disease 2022
“The GOAL trial showed no significant difference in regression rate between BPG-treated and untreated groups, and the efficacy of BPG in causing regression of definite and borderline latent R” — lamichhane-2023-detection-management-latent 2023
“There is a lack of abundant evidence on the safety profile of antibiotic prophylaxis in latent RHD.” — lamichhane-2023-detection-management-latent 2023
Whether the 2023 WHF A-D staging performs as claimed once applied in the fieldthin  6 papers
What nobody has establishedThe abbreviated 2023 screening criteria and the new A-D stages have not been tested for sensitivity, specificity, prognostic separation, or whether the stage assigned changes anything a clinician does.
Why it is still openThe criteria are two years old and the cohorts predate them; re-reading archived images is unglamorous work with no new fieldwork to justify a grant.
The study that would close itA re-read of existing screening cohorts under both 2012 and 2023 criteria with linkage to their subsequent progression and management.
“The study dataset was not adequate to apply the newly available 2023 WHF criteria, so agreement of handheld screening with standard echocardiography under the updated classification is unkno” — diniz-2024-agreement-between-handheld 2024
“The accuracy of screening findings under the 2023 WHF risk stratification, including for cases in which secondary prophylaxis might be initiated while awaiting confirmatory echo, has not bee” — diniz-2024-agreement-between-handheld 2024
“No included studies have been performed after the adoption of the 2023 WHF criteria, so prevalence under the new four-stage (A–D) classification is unknown.” — mutarelli-2025-global-prevalence-sex 2025
What actually drives adherence to secondary prophylaxis and what level is enoughthin  6 papers
What nobody has establishedThe adherence level required for protection has never been established, existing measurement is self-report or register-based and unvalidated, adult data are almost absent, and patient-side acceptability has usually been inferred from providers.
Why it is still openThe eighty-percent threshold is inherited convention nobody has tested; the study needs registry linkage that most programmes cannot produce.
The study that would close itA cohort with objectively verified injection receipt linked to ARF recurrence, powered to estimate a dose-response between adherence and recurrence.
“The exact level of adherence to prophylactic antibiotics required for effective prevention of recurrent rheumatic attack is not known.” — bimerew-2024-adherence-secondary-antibiotic 2024
“Studies conducted purely on adults with RHD/ARF are limited, with only one Ugandan study (30% adherence) identified.” — bimerew-2024-adherence-secondary-antibiotic 2024
“It remains unknown whether the included studies measured adherence reliably, since most were observational (cross-sectional, retrospective or prospective) and self-report or register-based.” — bimerew-2024-adherence-secondary-antibiotic 2024
An alternative to benzathine penicillin Gthin  6 papers
What nobody has establishedNo better drug exists: long-acting alternatives, oral non-inferiority, subcutaneous infusion and less frequent regimens are all unproven, and one Ethiopian hospital has already switched to oral amoxicillin without evidence.
Why it is still openBPG is off-patent and cheap, so there is no commercial sponsor; the trial has to be publicly funded and run for years.
The study that would close itA non-inferiority trial of oral penicillin versus intramuscular BPG for preventing progression in mild RHD (the trial the field says is running but has not reported).
“There is currently no ideal alternative drug to benzathine penicillin G for treating GAS or providing secondary prophylaxis for ARF.” — ali-2022-rheumatic-heart-disease 2022
“An antibiotic substitute for BPG with better administration techniques and a safer profile has not been developed.” — ali-2026-implementing-world-health 2026
“Whether the cessation of BPG administration at the University of Gondar Hospital and replacement with oral amoxicillin has changed patient treatment outcomes and adherence.” — birru-2025-acceptability-implementation-challenges 2025
Unexplained heterogeneity between prevalence studiesthin  6 papers
What nobody has establishedPooled prevalence estimates carry extreme residual heterogeneity that subgroup analysis by region, year and sampling unit does not explain, and country-level differences (Ethiopia versus its neighbours, South Asia's death burden) are attributed to speculation.
Why it is still openIt needs raw data from dozens of groups who have no incentive to share, so every review re-pools the published summaries and re-reports the heterogeneity.
The study that would close itAn individual participant data meta-analysis of screening studies with harmonised criteria and measured setting-level covariates.
“The reasons for heterogeneity in reported RHD prevalence across the included studies have not been fully explained despite subgroup analysis by region, country and sampling unit.” — abdu-2024-prevalence-pattern-rheumatic 2024
“The markedly higher RHD pooled prevalence in Ethiopia compared with Sudan, Uganda and Tanzania is not explained beyond speculation about socioeconomic, awareness and healthcare-access differ” — mebrahtom-2023-rheumatic-heart-disease 2023
“The actual reasons for the disproportionately high contribution of South Asia to global RHD deaths despite declining age-standardized death rates have not been definitively established.” — naseeb-2024-temporal-trends-burden 2024
The economic value case for a GAS vaccinecrowded  5 papers
What nobody has establishedNo benefit-cost ratio exists because development, manufacturing and delivery costs have never been estimated; AMR benefit, herd effects and household spillovers are unvalued, and there are no willingness-to-pay or VSLY estimates for low-income countries.
Why it is still openManufacturers will not share cost structures for a product with no market, and the health-economics inputs for low-income countries simply do not exist.
The study that would close itA costed target-product-profile exercise with manufacturers paired with a full societal-perspective economic evaluation including AMR and indirect effects.
“Estimates of the non-internalized benefits of Strep A vaccination (e.g., population-wide AMR mitigation, herd effects, spillover mental health and quality-of-life effects on household member” — cadarette-2023-full-health-economic 2023
“No empirical VSLY estimates exist for lower-income countries, so VSLY is proxied by per-capita GDP with an unknown elasticitiy.” — cadarette-2023-full-health-economic 2023
“The magnitude of Strep A vaccination's impact on antimicrobial resistance, particularly through reduced antibiotic consumption, has not been quantified within this study.” — cadarette-2023-full-health-economic 2023
Whether housing and environmental-health interventions cut Strep A transmissionthin  5 papers
What nobody has establishedNine proposed environmental initiatives (ventilation, cleaning, dust reduction, wound care) have no published efficacy evidence for Strep A, baseline household ventilation has never been measured, and there is no evaluation tool that would capture Aboriginal definitions of health or cost-benefit.
Why it is still openIt is expensive, slow, politically loaded, and crosses housing and health budgets that do not share a funder; the causal chain to RHD takes years to show.
The study that would close itA community-led stepped-wedge trial of a housing and environmental-health package in remote communities, with Strep A skin and throat infection incidence as the primary outcome.
“Whether reducing sugar-sweetened drink intake could lower ARF/RHD risk in children is unknown and requires investigation.” — baker-2022-risk-factors-acute 2022
“The role of specific housing factors (damp, mould, cold, hot water supply, bed sharing, functional crowding) as ARF risk factors remains to be investigated more rigorously.” — baker-2022-risk-factors-acute 2022
“No study has evaluated the impact of primordial prevention strategies on GAS, ARF and RHD in Malawi's resource-poor communities.” — blennerhassett-2025-burden-rheumatic-heart 2025
Task-sharing and decentralising RHD care to non-specialiststhin  5 papers
What nobody has establishedThere is no evidence on task-sharing models for RHD across the whole pathway: community health workers giving BPG, school nurses screening throats, mid-level providers doing post-operative follow-up, and how training standards are maintained over time.
Why it is still openTask-sharing evaluations need health-services and implementation expertise that cardiology groups do not have, and ministries implement without a control arm.
The study that would close itA stepped-wedge implementation trial of a defined task-sharing package at primary-care level, evaluated with implementation-science outcomes alongside adherence and retention.
“No evidence exists on task-sharing models for expanding access to RHD prevention and treatment services in limited-resource settings.” — abdullahi-2019-task-sharing-diagnosis 2019
“No studies have addressed prevention of acute rheumatic fever (primary prevention) or passive case-finding of recurrent ARF and RHD through task-sharing approaches.” — abdullahi-2019-task-sharing-diagnosis 2019
“No studies have evaluated task-sharing models for tertiary prevention, including post-operative management of patients after heart valve surgery.” — abdullahi-2019-task-sharing-diagnosis 2019
ARF has no diagnostic gold standard and no field-usable algorithmcrowded  5 papers
What nobody has establishedThe Jones criteria have never had their accuracy quantified against a true reference, no simplified algorithm has been prospectively tested at primary-care level in an endemic setting, and there is no protocol for separating ARF from malaria, typhoid and viral illness.
Why it is still openThere is no reference standard to validate against, so the study has to construct one from follow-up, which makes it long, expensive and methodologically arguable before it starts.
The study that would close itA prospective multi-country study enrolling every child presenting with fever plus joint symptoms, applying Jones, a simplified algorithm and handheld echo in parallel, with adjudicated outcome at 12 months as the reference.
“A specified protocol to help frontline health workers discriminate ARF from its endemic mimickers (e.g., malaria, typhoid, viral infections) has not been defined.” — ali-2022-rheumatic-heart-disease 2022
“It has not been determined whether treating any joint symptom as probable ARF in endemic settings causes net harm from overdiagnosis and unnecessary BPG prophylaxis.” — ali-2022-rheumatic-heart-disease 2022
“A definitive diagnostic test for acute rheumatic fever has not been discovered, although biomarker discovery studies (START; ARF Diagnostic Collaborative Network) are still underway.” — ali-2026-implementing-world-health 2026
A biomarker that stages RHD and flags who will deterioratethin  5 papers
What nobody has establishedCandidate molecular markers (Tenascin-C, miRNAs, TNF-alpha, GlcNAc-specific IgG2) have never been validated at scale, and nothing exists to stratify a diagnosed patient by risk of progression.
Why it is still openDiscovery studies are cheap and validation cohorts are not; the African populations where validation matters most have the least biobanking infrastructure.
The study that would close itA biobanked longitudinal RHD cohort with serial sampling, testing candidate panels against echocardiographic progression and clinical events over five years.
“Validated, widely adopted molecular biomarkers (e.g., Tenascin-C, miR-1183/miR-1299, GlcNAc-specific IgG2) for RHD are not yet established in African clinical practice.” — aliyu-2024-rheumatic-heart-disease 2024
“Large-scale validation of molecular biomarkers for RHD in diverse African populations has not been done.” — aliyu-2024-rheumatic-heart-disease 2024
“Molecular biomarkers that stratify RHD patients by risk of disease progression, guide resource allocation, and signal need for intensive monitoring remain unidentified.” — aliyu-2024-rheumatic-heart-disease 2024
Why patients delay care, drop out of follow-up, and feel stigmatisedthin  5 papers
What nobody has establishedThe determinants of initial care-seeking for sore throat, of retention in lifelong follow-up, of default after a positive screen, and of stigma are largely unstudied, and socio-cultural context has not been designed into interventions.
Why it is still openQualitative work is under-funded in cardiovascular research and is treated as a preliminary to the real study rather than the study.
The study that would close itA multi-site qualitative and mixed-methods study of the care pathway from first sore throat to lifelong prophylaxis, in at least three culturally distinct endemic settings.
“Local socio-cultural contexts have not been seriously considered in the design of most RHD prevention and control interventions.” — abrams-2020-integrating-prevention-control 2020
“The factors influencing an individual's initial decision to seek care for RHD-related illness in Malawi are not well understood.” — blennerhassett-2025-burden-rheumatic-heart 2025
“The factors influencing retention in long-term RHD care in Malawi have not been elucidated.” — blennerhassett-2025-burden-rheumatic-heart 2025
Benzathine penicillin injections: severe reactions, pain, and safer formulationscontested  5 papers
What nobody has establishedThe true incidence and mechanism of fatal reactions after BPG are not established (vasovagal versus anaphylaxis is unresolved), injection pain is unmanaged, and pre-filled syringes, subcutaneous infusion and product-brand effects on needle blockage are untested in high-burden settings.
Why it is still openThe events are rare enough to need enormous denominators, and a hospital that has had a death stops giving BPG rather than studying it.
The study that would close itA prospective pharmacovigilance registry of BPG administrations in a high-burden country with adjudicated adverse events, plus a randomised comparison of pre-filled versus reconstituted BPG on pain and completion.
“The causes and risk factors for fatal reactions following benzathine penicillin G injection remain not fully understood.” — ali-2026-implementing-world-health 2026
“The true incidence of fatal and severe reactions to BPG in patients with severe RHD from highly endemic countries is not well quantified for programmatic risk assessment.” — ali-2026-implementing-world-health 2026
“The true incidence, nature and risk factors of severe adverse reactions following BPG injection for ARF/RHD prophylaxis remain insufficiently characterised.” — birru-2025-acceptability-implementation-challenges 2025
Interventions to improve prophylaxis adherence have never been trialledthin  5 papers
What nobody has establishedHealth-worker training, patient education, reminder systems, support groups, home-based delivery and group versus individual counselling are all recommended and none has been tested against a control.
Why it is still openEveryone is confident these work, which is exactly why nobody randomises them; programmes implement and report before-after.
The study that would close itA randomised trial of a bundled adherence intervention (reminders, education, home delivery) versus standard clinic-based BPG, with injection completion at twelve months as the outcome.
“Whether the implementation of comprehensive HCP education and training programs on RHD management and BPG administration will improve BPG delivery and uptake in Ethiopia.” — birru-2025-acceptability-implementation-challenges 2025
“Whether patient and community awareness programs will shift the paradigm of poor BPG acceptability and improve secondary prophylaxis adherence.” — birru-2025-acceptability-implementation-challenges 2025
“Whether the informal written consent procedures used by HCPs in some Ethiopian centres reduce HCP anxiety without further undermining patient adherence to BPG.” — birru-2025-acceptability-implementation-challenges 2025
AI and algorithmic detection of RHDthin  5 papers
What nobody has establishedDeep-learning classification of echo, of ECG/PCG waveforms, automated abnormality flagging on portable devices, and AI separation of acute from chronic carditis are all proposed with a sparse published evidence base and no prospective validation.
Why it is still openModel papers are cheap and validation is not; the labelled datasets are small, single-site, and not shared.
The study that would close itA prospective external validation of an echo-based RHD classifier against expert WHF reads in a screening population it was not trained on.
“The accuracy of a deep learning algorithm trained on DAS waveform data to classify subclinical RHD into definite, borderline, and normal categories has not been tested.” — ali-2021-detection-subclinical-rheumatic 2021
“It is not yet known whether a DAS-based DL algorithm can detect subclinical RHD without the need for echocardiography.” — ali-2021-detection-subclinical-rheumatic 2021
“The sensitivity and specificity of the proposed FC-LSTM and Conv-LSTM classifiers for ECG, PCG, and hybrid signals in distinguishing RHD from normal have not been measured.” — ali-2021-detection-subclinical-rheumatic 2021
Whether a GAS vaccine can be made without triggering autoimmunitycontested  5 papers
What nobody has establishedThe autoimmunity risk that killed early candidates has never been resolved, only worked around; whether M-protein or GAC-based antigens provoke cross-reactive responses in previously exposed people is undetermined.
Why it is still openThe risk is regulatory and reputational rather than scientific: nobody wants to be the trial that revives the 1970s ban, so the question stays theoretical.
The study that would close itA carefully monitored dose-escalation trial in GAS-experienced adults from an endemic population with systematic autoantibody and cardiac surveillance.
“Whether M-protein-only vaccines cause autoimmune reactions in individuals with pre-existing GAS exposure has not been determined in the general population.” — azuar-2019-recent-advances-development 2019
“The precise variability of GAC cross-reactivity with human tissues remains unresolved, as a recent study showed GlcNAc is not a universal GAS virulence factor.” — azuar-2019-recent-advances-development 2019
“The role of the GlcNAc sidechain in the autoimmune sequelae of Strep A infection remains controversial and unresolved.” — burns-2023-progress-glycoconjugate-vaccine 2023
Animal models cannot predict human GAS vaccine efficacythin  5 papers
What nobody has establishedNo validated, standardised animal model exists for GAS vaccine evaluation, non-human primates cannot be reliably infected, and mouse data have repeatedly failed to translate.
Why it is still openModel development is infrastructure work that no single vaccine developer will pay for and no funder frames as a deliverable.
The study that would close itA head-to-head evaluation of candidate models (murine, NHP, controlled human infection) against the same set of vaccine candidates and the same immunological readouts.
“It is unknown whether the peptide/lipopeptide/saccharide delivery systems described will work in advanced preclinical settings using non-human or human primates, which has not yet been done.” — azuar-2019-recent-advances-development 2019
“It is not known whether non-human primate models can establish GAS infection beyond non-invasive colonisation, limiting their utility for GAS vaccine testing.” — azuar-2019-recent-advances-development 2019
“Available animal models are insufficient to predict efficacy of GAS vaccine candidates in humans.” — fan-2024-recent-scientific-advancements 2024
Cardiac surgery access and capacity in endemic settingsthin  5 papers
What nobody has establishedWhole countries have no cardiac surgery capacity, the humanitarian mission model is judged to transfer no skills without a demonstrated alternative, postoperative follow-up is uncontrolled, and there is no robust cost comparison against high-income settings.
Why it is still openSurgery is the most expensive part of the pathway and the least funded by global health, and mission organisations are not set up to be evaluated.
The study that would close itA costed evaluation of surgical capacity-building models (mission-based versus embedded training partnership) measured on locally performed case volume and outcomes over five years.
“Long-term postoperative follow-up of RHD patients in LMICs is essentially uncontrolled because of poor compliance and lack of services, so the true benefit of repair versus replacement canno” — antunes-2020-global-burden-rheumatic 2020
“The current humanitarian mission model for cardiac surgery in sub-Saharan Africa has been judged suboptimal for skill transfer and needs to be reshaped, but an effective alternative has not ” — antunes-2020-global-burden-rheumatic 2020
“Whether Malawi has any in-country capacity for cardiac surgery for advanced RHD has not been established and is described as virtually absent.” — blennerhassett-2025-burden-rheumatic-heart 2025
Health-system and supply-chain failures that break RHD carethin  5 papers
What nobody has establishedStaff turnover, waiting times, drug stockouts, BPG supply chains dependent on external donation, and absent echo infrastructure are named as causes of failure and never measured.
Why it is still openIt is unglamorous operational research, and the findings implicate the ministry that would have to fund it.
The study that would close itA health-system readiness and supply-chain audit across facilities delivering RHD care in a high-burden country, linked to patient-level continuity outcomes.
“Health service and system-level factors that may have affected outcomes, including access to primary care, Aboriginal Health workers, health information in Aboriginal languages, and cultural” — doran-2023-surgery-rheumatic-heart 2023
“The impact of staff turnover, variation in infrastructure, and high patient volume on adoption and fidelity is recognized but not quantified.” — klassen-2026-implementing-national-acute 2026
“The influence of system-level factors such as long waiting times, staff shortages and drug availability on decentralised RHD implementation has not been measured.” — minja-2023-evaluating-implementation-dynamic 2023
unclusteredorphan  5 papers
What nobody has establishedTrue miscellany: whether chest pain in screen-detected children is cardiac, whether infection-prevention interventions work in multi-valvular disease, malnutrition in paediatric RHD, the pandemic's effect on RHD patients, and inequity in global health research funding.
Why it is still openEach is a single author's aside rather than a question the field has taken up.
The study that would close itEach would need its own study; none has a second paper behind it in this corpus.
“Whether chest pain in RVD-positive children is attributable to RHD or to other causes remains undetermined, as no pericardial damage was found on ultrasound.” — nganougnindjio-2025-rheumatic-valvulopathy-sub 2025
“The effectiveness of specific infection prevention interventions cannot be assessed in this study.” — stan-2026-clinical-characteristics-management 2026
“What is the magnitude of malnutrition in pediatric patients with rheumatic heart disease?” — tsega-2023-descriptive-rheumatic-heart 2023
Host genetic susceptibility, and the absence of African and Oceanian cohorts from itthin  4 papers
What nobody has establishedThe causal variants behind the HLA signals are unresolved, published GWAS are underpowered, most reported variants come from single studies, and the populations with the highest burden are the least represented in reference panels.
Why it is still openGenomics money follows existing biobanks, and the biobanks are European; building an African RHD cohort at GWAS scale is a decade-long infrastructure project, not a study.
The study that would close itA large multi-ancestry GWAS and HLA fine-mapping study built primarily on sub-Saharan African and Oceanian cohorts with echo-confirmed phenotypes.
“Functional consequences of the PITX2 exon 4 rare missense variants identified remain unestablished.” — amath-2026-genetic-variability-amp 2026
“Whether the rare coding variants in PITX2 exon 4 contribute mechanistically to atrial fibrillation pathophysiology has not been experimentally tested.” — amath-2026-genetic-variability-amp 2026
“The role of PITX2 variants in AF susceptibility cannot be isolated because the control group lacked rheumatic mitral valve disease.” — amath-2026-genetic-variability-amp 2026
GAS vaccine formulation, adjuvant and delivery routethin  4 papers
What nobody has establishedAdjuvant choice, mucosal versus systemic route, dose number, carrier protein and bioconjugation chemistry are all unsettled, and mucosal candidates still need five doses and toxic adjuvants to work in mice.
Why it is still openIt is a large combinatorial preclinical programme with no single hypothesis, so it gets done piecemeal by individual labs on individual formulations.
The study that would close itA systematic mouse factorial study varying route, adjuvant and dose count with matched serum IgG, mucosal IgA, functional activity and challenge protection readouts.
“Whether combined C- and N-terminal peptide vaccines can be reformulated without the toxic Freund's complete adjuvant while retaining efficacy is unresolved.” — azuar-2019-recent-advances-development 2019
“It is unclear whether CjPglB can efficiently transfer a rhamnose polymer whose GlcNAc is linked to rhamnose via the Strep A-specific β-1,4 linkage, requiring alternative strategies for PGCT-” — burns-2023-progress-glycoconjugate-vaccine 2023
“Future studies are required to identify the best protein carrier to induce T-cell help and enable GAC-directed memory responses.” — burns-2023-progress-glycoconjugate-vaccine 2023
Immune correlates of protection for a GAS vaccinethin  4 papers
What nobody has establishedNo correlate of protection exists, so no vaccine candidate can be licensed on immunogenicity, and it is not known whether antibodies to conserved antigens are mechanistic or merely associated.
Why it is still openThe field openly expects this will not be resolved until a large efficacy trial happens, which is circular: the trial needs the correlate to be affordable.
The study that would close itA controlled human infection model or large natural-history cohort with pre-exposure serology across a defined antigen panel, powered to identify protective thresholds separately for carriage, pharyngitis and pyoderma.
“No defined human immunity correlate of protection exists to guide clinical development of GAS vaccines.” — fan-2024-recent-scientific-advancements 2024
“The contribution of conserved antigens such as SpyCEP, SpyAD, SLO and GAC to natural protection in humans was unknown.” — keeley-2025-early-life-serological 2025
“The relative roles of humoral immunity to conserved antigens versus emm type-specific M protein in protection against S. pyogenes remain uncharacterised.” — keeley-2025-early-life-serological 2025
Economic evaluation of prevention strategies other than vaccinesthin  4 papers
What nobody has establishedNo study has compared combinations of primary, secondary and tertiary prevention, nor their distributional effects; models are blocked by absent surgery-complication data, absent local utility value sets, and single-source transition probabilities.
Why it is still openThe model inputs do not exist locally, so analysts import them from elsewhere and then note that the result may not transfer.
The study that would close itAn extended cost-effectiveness analysis of prevention-strategy combinations in one high-burden country, built on locally collected transition, cost and utility data.
“No previous study has evaluated the distributional effect of primary, secondary, or tertiary interventions and their combinations for the prevention and control of rheumatic fever and rheuma” — dixit-2023-evaluating-efficiency-equity 2023
“Whether the same coverage rates for primary, secondary, and tertiary interventions across socioeconomic subgroups hold in reality has not been empirically tested for these interventions.” — dixit-2023-evaluating-efficiency-equity 2023
“The specific complications of rheumatic heart disease surgery, such as valve failure and the need for repeat valve replacement surgery, have not been modelled because sound epidemiological d” — dixit-2023-evaluating-efficiency-equity 2023
Standardised immunoassays and safety surveillance for GAS vaccine trialsthin  4 papers
What nobody has establishedThere are no international antibody standards, no inter-laboratory reproducibility data for the platforms in use, unquantified cross-reactivity with S. dysgalactiae, and no standardised safety-surveillance framework.
Why it is still openAssay harmonisation is thankless coordination work that no single group is rewarded for, and it only becomes urgent when trials are already running.
The study that would close itA multi-laboratory ring trial of the Luminex and MSD platforms on a shared sample panel, with an international reference serum established.
“A specific SLO neutralising assay was not conducted for the SLO-GACPR conjugate, leaving retained functional SLO antibody activity unconfirmed.” — burns-2023-progress-glycoconjugate-vaccine 2023
“Calibration of the S. pyogenes antibody assays to international standards when these are developed is needed to allow wider comparison and validation.” — keeley-2025-early-life-serological 2025
“The degree to which cross-reactive immune responses to Streptococcus dysgalactiae subsp. equisimilis may have contributed to the measured anti-SLO and anti-SpyAD responses has not been quant” — keeley-2025-early-life-serological 2025
GBD model estimates cannot be checked against real datathin  4 papers
What nobody has establishedGlobal Burden of Disease figures for RHD rest on sparse, miscoded inputs, assume log-linear trends, cannot be traced to their local sources, and cannot distinguish real epidemiological change from changing diagnostic practice.
Why it is still openGBD is the number everyone quotes precisely because nothing better exists; auditing it requires the primary data whose absence created it.
The study that would close itA validation study comparing GBD RHD estimates against primary registry and screening data in several countries, with the input sources made explicit.
“The Bayesian age-period-cohort projections of RHD incidence to 2040 rely on assumptions about future healthcare access, prevention and socioeconomic change that cannot be verified.” — chen-2025-global-burden-trend 2025
“The EAPC method assumes a log-linear trend over time and may miss nonlinear patterns or sudden shifts in RHD burden.” — chen-2025-global-burden-trend 2025
“Integrating clinical registry data and population-based screening with GBD estimates to improve precision and validity of burden estimates has not been done.” — chen-2025-global-burden-trend 2025
Whether penicillin resistance is emerging and what antibiotic use these programmes causethin  4 papers
What nobody has establishedPenicillin-binding-protein mutations are unmonitored in most endemic populations, the antibiotic consumption attributable to sore-throat programmes is unquantified, and whether reduced susceptibility translates into treatment failure is unknown.
Why it is still openPenicillin has worked for seventy years so surveillance feels optional, and the labs that would do it are not in the endemic countries.
The study that would close itLongitudinal genomic surveillance of GAS isolates in a sore-throat programme region linked to dispensing data and clinical failure rates.
“Whether penicillin resistance will emerge in African GAS strains remains uncertain and unmonitored.” — aliyu-2024-rheumatic-heart-disease 2024
“It is not known whether pbp mutations are emerging in the local (New Zealand) GAS population, and further surveillance of local GAS isolates is warranted.” — cannon-2023-trends-penicillin-dispensing 2023
“The potential overuse of amoxicillin during and after the RFPP has not been assessed.” — cannon-2023-trends-penicillin-dispensing 2023
Whether ARF cases are being counted at allthin  4 papers
What nobody has establishedUndiagnosed ARF may have different risk factors from diagnosed ARF, symptom-triggered case-finding misses subclinical carditis, recalled sore-throat history is unreliable, and apparent first episodes may be undocumented recurrences.
Why it is still openActive surveillance for a rare acute illness is expensive per case, and the cases it finds make the passive system look bad.
The study that would close itA community-based active ARF surveillance study with echo on all suspected cases, compared against passive facility-based case-finding in the same population.
“It is unknown whether unrecognised (undiagnosed) ARF cases have different risk factors from those who are diagnosed.” — baker-2022-risk-factors-acute 2022
“The extent to which a symptom-triggered approach misses children with subclinical carditis detectable only by echocardiography is not measured.” — klassen-2026-implementing-national-acute 2026
“The reasons for the low (66.9%) reported history of sore throat among RVD-positive students—compared with 100% in other African studies—are not established, with possible memory bias and the” — nganougnindjio-2025-rheumatic-valvulopathy-sub 2025
Prosthesis choice and anticoagulation failure in young LMIC patientscontested  3 papers
What nobody has establishedIt is not established whether tissue or mechanical valves give better survival in young patients who cannot maintain INR control, whether poor time-in-range is adherence or surveillance, whether DOACs could be used, or how to identify in advance the patients who will fail anticoagulation.
Why it is still openWestern guidelines are age-based and get applied unmodified; the trial needs long follow-up in exactly the patients who are hardest to retain.
The study that would close itA risk-stratified prospective comparison of mechanical versus tissue valve replacement in young LMIC patients, with pre-operative adherence and access variables used to build and validate a failure-prediction tool.
“No study has assessed the relationship between anticoagulation use and clinical outcomes in severe multi-valvular RHD using these data.” — stan-2026-clinical-characteristics-management 2026
“Whether INR monitoring can be feasibly improved to allow safe mechanical valve use in low-income regions remains unresolved.” — zhang-2025-recent-advances-prevention 2025
“No clinical trials have assessed whether the promising disrupted tri-leaflet mechanical and polymer leaflet valve designs actually improve outcomes in young patients with RHD and poor INR co” — zilla-2024-mechanical-valve-replacement 2024
Whether skin infection, not just sore throat, causes ARF and whether treating it helpscontested  3 papers
What nobody has establishedNo trial has tested whether treating laboratory-confirmed Strep A skin infection reduces ARF or RHD, the impetigo contribution to ARF risk is unquantified, the scabies link is uninvestigated, and no antibiotic is preferred on evidence.
Why it is still openThe whole prevention edifice is built on the throat, so the skin pathway has been an aside for fifty years; the trial needs very large numbers to see an ARF signal.
The study that would close itA cluster-randomised trial of active impetigo detection and treatment (or MDA) in a high-ARF community, with ARF incidence as the primary outcome.
“The contribution of impetigo (as distinct from pharyngitis) to ARF risk in Australia has not been quantified, despite skin infection being a noted driver.” — barth-2022-roadmap-incorporating-group 2022
“No trial has determined whether antibiotic treatment of laboratory-confirmed Strep A skin infections reduces the subsequent risk of ARF or RHD.” — leong-2025-assessing-evidence-antibiotic 2025
“No placebo-controlled trial of penicillin for laboratory-confirmed Strep A skin infections exists.” — leong-2025-assessing-evidence-antibiotic 2025
Managing and predicting risk in pregnant women who already have RHDthin  3 papers
What nobody has establishedRisk-stratification models were built on congenital heart disease in high-income settings and may not work where RHD predominates; cardiac medication and prophylaxis use in pregnancy is low for unexplained reasons, postpartum follow-up is short, and there are no data at all from LMIC Asia or Latin America.
Why it is still openObstetric and cardiology services do not share datasets, and the countries with the most cases contribute the least research.
The study that would close itA prospective multi-country RHD-in-pregnancy registry with twelve-month postpartum follow-up, used to derive and validate an RHD-specific risk model.
“Outcomes of pregnant women with RHD under current clinical practice in Australia (raised as relevant to LMIC contexts) have not been previously evaluated.” — aliyu-2024-rheumatic-heart-disease 2024
“Studies directly linking age at first birth to RHD risk are limited.” — chen-2025-global-burden-trend 2025
“The long-term reproductive outcomes of women with RHD are insufficiently documented.” — chen-2025-global-burden-trend 2025
No animal or ex vivo model reproduces human rheumatic valve diseasethin  3 papers
What nobody has establishedRodent models cannot reproduce slow progressive fibrosis and calcification, no patient-derived valve organoid exists, and the rat work has been done in females only.
Why it is still openIt is methods development with no disease-specific funding line, and the field keeps publishing in the flawed model because it is what exists.
The study that would close itDevelopment and characterisation of a patient-derived 3D valvular organoid combining valve endothelial and interstitial cells on an ECM scaffold under physiological shear.
“Ex vivo models that recapitulate the hemodynamic and pathophysiological features of human rheumatic mitral valve disease are needed to correlate regional shear stress with histopathological ” — liu-2025-chronic-mitral-valve 2025
“There remains a lack of reliable models that accurately recapitulate the pathophysiological features of RHD, hindering mechanistic investigations.” — liu-2025-chronic-mitral-valve 2025
“Rodent models cannot replicate the slow, progressive fibrosis and calcification characteristic of human mitral valve disease, limiting translational relevance.” — liu-2025-chronic-mitral-valve 2025
How Strep A actually spreads between people and in householdsthin  3 papers
What nobody has establishedThe relative contribution of respiratory droplets, aerosols, skin-to-skin contact and fomites has never been studied holistically, and no household-level transmission study exists.
Why it is still openIt requires longitudinal household sampling with genomics in remote settings, expensive and logistically brutal, and transmission research went out of fashion decades ago.
The study that would close itA household transmission study with genomic linkage of isolates plus environmental and air sampling, in a high-incidence community.
“The specific dynamics of Strep A transmission pathways are unknown.” — giannini-2022-potential-health-impact 2022
“A transmission dynamic model is needed to estimate the total (direct and indirect) effects of Strep A vaccination.” — giannini-2022-potential-health-impact 2022
“The relative impact of individual Strep A transmission modes (large respiratory droplets, skin-to-skin contact, and fomites) has not been studied holistically.” — giannini-2023-modeling-potential-health 2023
Whether RHD programmes should be integrated into general health systems or run in parallelthin  3 papers
What nobody has establishedWhich programme features make integration work, whether integrated models outperform parallel ones, and how RHD care should attach to UHC and to other NCD programmes are all unestablished.
Why it is still openIntegration is a policy preference asserted in guidance documents; comparing models requires two ministries willing to run different systems side by side.
The study that would close itA prospective quasi-experimental comparison of integrated versus vertical RHD programme models across matched districts.
“Which programme factors facilitate integration into health systems while still delivering good outcomes for RHD is not known.” — abrams-2020-integrating-prevention-control 2020
“No prospective quasi-experimental or experimental study has compared the effectiveness of RHD-related health technologies within fully integrated versus partially integrated programme models” — abrams-2020-integrating-prevention-control 2020
“The association between the degree of programme integration and population health outcomes could not be quantified because of the small number of studies and outcome heterogeneity.” — abrams-2020-integrating-prevention-control 2020
At what age a GAS vaccine should be giventhin  3 papers
What nobody has establishedIt is unknown what proportion of RHD could be averted by an infant schedule, because the interval between early-life infection and later ARF is unmodelled.
Why it is still openIt depends on the natural-history parameters that nobody has measured, so the modellers keep flagging it rather than solving it.
The study that would close itAn age-structured transmission and progression model calibrated to cohort data on early-life GAS exposure and subsequent ARF incidence.
“It is unclear what proportion of RHD cases arising at age 5 could be averted by an infant vaccination schedule.” — giannini-2022-potential-health-impact 2022
“Further epidemiology or immunology studies are required to model the proportion of RHD cases that could be averted by an infant vaccination schedule preventing the infections preceding ARF a” — giannini-2023-modeling-potential-health 2023
“The optimal age for S. pyogenes vaccine introduction in early life is unresolved and requires evaluation in vaccine trials and modelling.” — keeley-2025-early-life-serological 2025
Outcomes of RHD managed without surgerythin  3 papers
What nobody has establishedLong-term survival, cause of death and independent predictors of mortality in adults and children with RHD who never receive an intervention are scarcely documented anywhere in sub-Saharan Africa.
Why it is still openCohorts of unoperated patients are cohorts of people who fell out of the health system, which is exactly what makes them hard to follow.
The study that would close itA prospective cohort of medically managed RHD patients with active follow-up and verbal-autopsy-verified cause of death.
“Long-term outcome data for adults with RHD managed without surgical or percutaneous interventions in sub-Saharan Africa are scarce.” — nasir-2026-survival-mortality-predictors 2026
“Independent predictors of complications and mortality in severe multi-valvular RHD have not been identified, because no multivariable regression was performed.” — stan-2026-clinical-characteristics-management 2026
“What is the mortality rate and main cause of death in pediatric RHD patients in Ethiopia?” — tsega-2023-descriptive-rheumatic-heart 2023
At what age children should first be screened, and how often to repeat itthin  3 papers
What nobody has establishedThe ideal age to start echo screening, and when to re-screen someone who screened negative, have never been determined.
Why it is still openIt needs the same children screened more than once, and screening programmes are funded for coverage rather than for repeat rounds.
The study that would close itA repeat-screening cohort study across age bands measuring incident RHD detection yield by age at first screen and by re-screen interval.
“The optimal age at which to initiate RHD screening remains uncertain.” — mutarelli-2025-global-prevalence-sex 2025
“Studies are needed to determine the ideal age at which to screen children for RHD.” — rwebembera-2023-world-heart-federation 2023
“The optimal timing of repeat echocardiographic screening in individuals with a previous negative screening result has not been determined.” — topcu-2024-echocardiographic-screening-rheumatic 2024
Valve lymphangiogenesis as a driver of chronic valve damagethin  2 papers
What nobody has establishedA mouse model shows VEGFR3-driven lymphangiogenesis in inflamed valves, but nothing is confirmed in human tissue, the cells of origin are unidentified, and whether blocking it after disease is established helps is untested.
Why it is still openOne group owns the finding and the human validation needs valve tissue at single-cell resolution; the murine work is far ahead of anything human.
The study that would close itSingle-cell profiling of human rheumatic valve endothelium paired with a late-treatment VEGFR3-blockade arm in the K/B.g7 model.
“The role of lymphatic vessels in the progression of valvular disease is not yet well established.” — lupieri-2025-rheumatic-heart-valve 2025
“It is unknown whether the PROX1+LYVE1neg VECs on the ventricular surface of the mitral valve are mechanosensing cells that respond to oscillatory shear stress.” — osinski-2022-pathogenic-lymphangiogenesis-promotes 2022
“It is unknown how the emergent lymphatics, infiltrating immune cells, and resident fibroblasts interact to promote chronic valve fibrosis.” — osinski-2022-pathogenic-lymphangiogenesis-promotes 2022
Antenatal echo screening for RHD in pregnancythin  2 papers
What nobody has establishedNo controlled trial has compared antenatal echo screening with standard antenatal care, the optimal timing is unknown, universal versus targeted screening is unresolved, and diagnostic accuracy, cost-effectiveness, acceptability and false-positive harms are all unreported.
Why it is still openPregnancy trials carry extra regulatory weight, and antenatal services in these settings are already stretched, so screening gets added as service delivery without a control arm.
The study that would close itA randomised trial of routine antenatal handheld echo versus standard antenatal care in a high-prevalence setting, with maternal cardiac events as the primary outcome.
“The prevalence of RHD in pregnant women in Sudan, and the contribution of undiagnosed RHD to maternal mortality, is unknown.” — elazrag-2022-handheld-echocardiographic-screening 2022
“No controlled trials have compared routine antenatal echocardiographic screening versus standard antenatal care for RHD in pregnant women in high-prevalence areas.” — seitler-2024-routine-antenatal-echocardiography 2024
“The optimal timing of echocardiographic screening during pregnancy has not been established.” — seitler-2024-routine-antenatal-echocardiography 2024
Vaccine impact models exclude herd effects and the whole ARF pathwaythin  2 papers
What nobody has establishedEvery published impact projection uses a static cohort model that leaves out indirect effects, omits ARF and APSGN entirely for lack of prevalence data, and assumes vaccine characteristics that no product has.
Why it is still openIt is blocked upstream by missing transmission and natural-history data, so modellers publish the static version and list the omission.
The study that would close itA dynamic transmission model of Strep A including the ARF-to-RHD pathway, calibrated to a setting with both infection and sequelae surveillance.
“The static cohort model only estimates direct effects of Strep A vaccination and excludes indirect herd effects.” — giannini-2022-potential-health-impact 2022
“The model excludes acute rheumatic fever (ARF) and acute post-streptococcal glomerulonephritis (APSGN) because of limited prevalent burden data.” — giannini-2022-potential-health-impact 2022
“The model does not capture the benefit of reduced antibiotic use from fewer Strep A infections because quality data are unavailable.” — giannini-2022-potential-health-impact 2022
Transcatheter valve therapy designed for rheumatic anatomythin  2 papers
What nobody has establishedCalcium distribution in rheumatic valves has never been described in the detail device design needs, no transcatheter mitral valve has been designed for rheumatic disease, and whether existing TAVI prostheses behave the same in a fused rheumatic landing zone is unknown.
Why it is still openDevice companies design for degenerative aortic stenosis in wealthy elderly markets; rheumatic anatomy is a small, poor, young market.
The study that would close itA CT-based morphological characterisation of rheumatic valve anatomy across a surgical cohort, paired with bench and early clinical evaluation of a prosthesis adapted to it.
“Detailed descriptions of calcium distribution in rheumatic valves are lacking, which is needed to inform transcatheter valve development.” — weich-2023-transcatheter-heart-valve 2023
“The true requirement for transcatheter aortic valve intervention after prior MIAV/PMBV or mitral valve replacement in RHD is not known.” — weich-2023-transcatheter-heart-valve 2023
“The number of redo RHD valve operations performed for a different valve (as opposed to the originally operated valve) is not known.” — weich-2023-transcatheter-heart-valve 2023
How prevention-programme evaluations should be designed and reportedthin  2 papers
What nobody has establishedEvaluations report adherence, awareness and detection rather than incidence, do not describe the programme structures they tested, and never attribute effect to individual components.
Why it is still openMethodological standard-setting has no natural home in this field, and each programme reports in whatever terms it collected.
The study that would close itA reporting-standard development exercise (Delphi plus retrospective application to published programmes) for ARF/RHD intervention evaluations.
“The effectiveness of school-based sore throat management programmes for preventing ARF remains unproven.” — baker-2022-risk-factors-acute 2022
“The feasibility and advisability of school-based primary prevention programs in low- and middle-income countries have not been established.” — shimanda-2024-preventive-interventions-reduce 2024
“Future evaluations of rheumatic heart disease interventions should report effects on disease incidence or prevalence rather than only adherence, awareness, disease progression, recurrence, p” — shimanda-2024-preventive-interventions-reduce 2024
Whether echo can distinguish acute valvulitis from established chronic RHDthin  2 papers
What nobody has establishedThe WHF guideline contains no criteria for telling reversible acute valve change from permanent chronic change, and no echocardiographic feature has been shown to predict which valves normalise.
Why it is still openIt needs incident ARF cases with imaging at presentation, which only a handful of registry settings can assemble, and the question was only sharpened by the 2023 criteria.
The study that would close itA prospective ARF cohort with echo at diagnosis and serial follow-up, testing candidate morphological features (beading, nodularity, leaflet thickness pattern) against valve normalisation at two years.
“No evidence exists for the use of echocardiography to differentiate acute-on-chronic valvulitis from chronic RHD.” — rwebembera-2023-world-heart-federation 2023
“Whether the 2023 WHF criteria can reliably distinguish acute valvulitis from chronic RHD with ARF recurrence has not been determined.” — williamson-2025-application-new-world 2025
“It is not known whether there are echocardiographic features at diagnosis that predict normalisation of the valves following ARF, particularly morphological features that distinguish acute f” — williamson-2025-application-new-world 2025
When to operate in rheumatic valve diseasethin  2 papers
What nobody has establishedWhether operating earlier, before pulmonary pressures rise and while ejection fraction is preserved, reduces death has never been tested prospectively, and no comparative treatment strategy analysis exists.
Why it is still openThe patients present late by definition, so the early-surgery arm is hard to fill, and surgical timing trials are rare everywhere.
The study that would close itA randomised or prospective comparative study of early versus guideline-timed intervention in severe rheumatic valve disease with preserved ventricular function.
“Whether earlier surgery, before pulmonary artery systolic pressure becomes dangerously elevated, reduces RHD-related death has not been tested prospectively.” — doran-2023-surgery-rheumatic-heart 2023
“It is unknown whether earlier intervention alters pulmonary vascular outcomes in severe multi-valvular RHD, as the available data are cross-sectional.” — stan-2026-clinical-characteristics-management 2026
“Optimal surgical timing for severe multi-valvular RHD cannot be concluded from these data, as there was no comparator.” — stan-2026-clinical-characteristics-management 2026
Rheumatic valve repair technique and durabilitythin  2 papers
What nobody has establishedAortic valve repair for rheumatic disease is in its infancy, leaflet-extension materials are unvalidated in young patients, rigid versus flexible annuloplasty is unsettled, and guidelines give no basis for choosing repair over replacement.
Why it is still openRepair depends on surgeon skill, which makes randomisation contentious, and the recurrent-rheumatic-activity confounder means results may not transfer between settings.
The study that would close itA multicentre randomised comparison of repair versus replacement in young rheumatic mitral disease, with freedom from reoperation at ten years.
“Repair of the rheumatic aortic valve remains in its infancy, with only limited recent reports of leaflet extension or mobilization/shaving techniques.” — antunes-2020-global-burden-rheumatic 2020
“Recurrent RF episodes and ongoing rheumatic activity place mitral valve repair at high risk of recurrent regurgitation or stenosis, but the magnitude and predictors of this failure are not c” — antunes-2020-global-burden-rheumatic 2020
“Whether autologous or heterologous pericardium (fresh or glutaraldehyde-treated) provides durable leaflet extension in very young rheumatic patients has not been established.” — antunes-2020-global-burden-rheumatic 2020
Clinical decision rules for sore throat in high-ARF settingsthin  2 papers
What nobody has establishedThe clinical presentation of Strep A in LMIC settings is poorly characterised, existing decision rules have been validated in a couple of sites, and inter-rater variation in the subjective signs they depend on is unmeasured.
Why it is still openDecision rules were developed for high-income antibiotic stewardship, a different objective from ARF prevention, so nobody has rebuilt them for this purpose.
The study that would close itA multi-site prospective study of sore-throat presentations across several endemic countries with duplicate independent clinical examination and NAAT plus serology reference.
“There is a critical gap in understanding the clinical presentation of Strep A infections in LMIC settings.” — armitage-2025-evaluating-clinical-decision 2025
“Whether the accuracy of CDRs for S. pyogenes sore throat diagnosis holds in high-ARF settings beyond Fiji, as the study was conducted at mainly two health centres in Suva.” — hart-2025-clinical-decision-rules 2025
“The true disease status (true infection vs asymptomatic carriage) of NAAT-positive sore throat cases, because streptococcal serology was not done.” — hart-2025-clinical-decision-rules 2025
Harms of screening and of labelling a well child with latent RHDthin  2 papers
What nobody has establishedThe psychological cost to children and families of a latent-RHD label plus years of injections has been noted repeatedly and never measured.
Why it is still openHarms research does not attract funding in a field arguing for more screening, and the people harmed are asymptomatic children who never sought care.
The study that would close itA prospective quality-of-life and anxiety study in screen-detected borderline children and their caregivers versus screen-negative controls.
“The psychological trauma experienced by patients and caregivers from long-term antibiotic prophylaxis has been noted but not quantified.” — lamichhane-2023-detection-management-latent 2023
“Further studies are needed to quantify the potential harms of echocardiographic RHD screening for the screened population.” — rwebembera-2023-world-heart-federation 2023
How long secondary prophylaxis should continuethin  2 papers
What nobody has establishedThe optimal stopping point is undefined, including whether prophylaxis can be safely ceased in children whose echo has normalised.
Why it is still openStopping an established preventive therapy is a hard trial to get past an ethics committee and an even harder one to recruit into.
The study that would close itA randomised discontinuation trial in patients with normalised echocardiography after a defined prophylaxis duration.
“The optimal duration of secondary antibiotic prophylaxis administration, including whether it can safely be ceased after echocardiographic normalization in children, is undefined.” — rwebembera-2023-world-heart-federation 2023
“The optimal time point at which to stop secondary penicillin prophylaxis has not been established.” — topcu-2024-echocardiographic-screening-rheumatic 2024
The evidence base for secondary prophylaxis itself is old and weakthin  2 papers
What nobody has establishedThe efficacy of secondary prophylaxis rests on mid-twentieth-century trials of low methodological quality and on assumption rather than contemporary evidence.
Why it is still openNobody wants to destabilise the one intervention the whole control strategy stands on, and re-running it is not ethical.
The study that would close itA systematic re-appraisal with GRADE of the original prophylaxis trials, paired with a contemporary registry-based effectiveness estimate.
“Few studies have compared the efficacy of penicillin for RHD with a control group, and the methodology of existing RCTs is of low quality.” — topcu-2024-echocardiographic-screening-rheumatic 2024
“The efficacy of most secondary antibiotic prophylaxis measures is derived from assumptions and low-quality historical data rather than robust contemporary evidence.” — zhang-2025-recent-advances-prevention 2025
What a working model of RHD care actually looks likethin  2 papers
What nobody has establishedThe elements of a care model that improves long-term outcomes remain uncertain, and short descriptive pilots cannot show whether a dedicated clinic slows progression.
Why it is still openThe question is asked at the end of papers that could not answer it; it needs comparative long-term follow-up that no single centre can mount.
The study that would close itA comparative evaluation of RHD care models (dedicated clinic, integrated NCD clinic, decentralised primary care) on progression and survival over five years.
“The crucial elements of a model of RHD care that improves long-term outcomes remain uncertain.” — doran-2023-surgery-rheumatic-heart 2023
“Whether the clinic reduces progression of RHD or delays need for surgery cannot be determined from a 6-month descriptive study.” — njedock-2023-initiating-first-rheumatic 2023
Household economic burden of RHD carethin  2 papers
What nobody has establishedOut-of-pocket costs to families of lifelong RHD care, and whether decentralised delivery reduces them, are unquantified.
Why it is still openCost-to-patient is outside the outcome set most trials register, and self-reported costs are known to be biased so people avoid measuring them.
The study that would close itA household cost survey nested in a decentralised-care trial, measuring direct and indirect costs and catastrophic health expenditure.
“Whether decentralised RHD care reduces out-of-pocket expenditures for households has not been quantified.” — minja-2023-evaluating-implementation-dynamic 2023
“Patient-level disparities in healthcare costs associated with RHD secondary prevention remain insufficiently understood, particularly in low-income countries such as Uganda.” — xu-2024-disparity-secondary-prevention 2024
Reproducibility of echocardiographic reading between operatorsthin  2 papers
What nobody has establishedInterobserver variability in applying WHF criteria, and inter-equipment variability over a study period, are routinely unassessed even in studies whose conclusions depend on them.
Why it is still openIt is a boring appendix study that no one gets credit for, so it is listed as a limitation instead of being done.
The study that would close itA formal interobserver and inter-device reliability study of WHF criteria application across readers of differing expertise.
“Interobserver variability in echocardiographic diagnosis was not formally assessed, so the reproducibility of WHF-criteria RHD classification is unmeasured.” — nasir-2026-survival-mortality-predictors 2026
“Inter-operator and inter-equipment echocardiographic measurement variability over the two-year study period cannot be excluded.” — stan-2026-clinical-characteristics-management 2026
Rolling out Strep A point-of-care testing: uptake, funding and governanceorphan  1 papers
What nobody has establishedReal-world uptake, operator performance against laboratory standards, workforce demands, comparative cost-effectiveness, reimbursement pathways and Indigenous-led governance for Strep A POCT are all still to be established.
Why it is still openThe technology is ready and the funding mechanism is not; without a reimbursement item there is no service to evaluate.
The study that would close itAn implementation trial of multi-pathogen POCT across remote primary-care services with concordance testing, workload measurement and a community-governed evaluation framework.
“The uptake and utilisation of Strep A POCT in remote regions still need to be evaluated to confirm the potential benefits described.” — barth-2022-roadmap-incorporating-group 2022
“The impact and benefits of multi-pathogen POCT for populations and health services, particularly workload and workforce requirements to sustain POCT, are not yet understood.” — barth-2022-roadmap-incorporating-group 2022
“Comparative evidence demonstrating the clinical benefits and cost-effectiveness of Strep A POCT versus current standard of care in Australia has not yet been generated.” — barth-2022-roadmap-incorporating-group 2022
Which post-translational modifications create the cross-reactive epitopesorphan  1 papers
What nobody has establishedNobody has mapped which citrullination, glycosylation or collagen modifications generate the neo-epitopes that break tolerance in RHD.
Why it is still openIt needs explanted human rheumatic valve tissue plus a proteomics core, a combination that exists in almost no single institution; one review has raised it and nobody has picked it up.
The study that would close itMass-spectrometry PTM mapping of explanted rheumatic valve tissue against non-rheumatic valve controls, with the candidate neo-epitopes tested for patient-serum reactivity.
“Does the altered quantitative expression of SLRPs in RHD tissues induce further post-translational modifications beyond collagen cross-linking during ECM remodelling?” — lumngwena-2022-pathophysiology-rhd-outstanding 2022
“What post-translational modifications on collagen peptides (bound by the PARF motive of M protein) enhance antigenicity in RHD?” — lumngwena-2022-pathophysiology-rhd-outstanding 2022
“Do citrullination and other post-translational modifications of vimentin, laminin and collagen contribute to their cross-reactivity in RHD?” — lumngwena-2022-pathophysiology-rhd-outstanding 2022
The cytokine profile of subclinical RHD and whether the inflammation persistsorphan  1 papers
What nobody has establishedWhether children with borderline echo findings carry an ongoing inflammatory signature, and whether it persists or resolves, rests on a single Brazilian cohort.
Why it is still openIt sits between the screening world and the immunology world and belongs administratively to neither.
The study that would close itSerial cytokine profiling of a screen-detected borderline cohort against screen-negative controls over three years, tied to echocardiographic outcome.
“What is the cytokine profile of borderline subclinical rheumatic heart disease?” — branco-2026-elevated-interleukin-levels 2026
“Do inflammatory markers persist longitudinally in subclinical rheumatic heart disease?” — branco-2026-elevated-interleukin-levels 2026
“Do interleukin-17A and interferon-gamma play different roles across phases of disease progression, or are they insensitive early markers?” — branco-2026-elevated-interleukin-levels 2026
Epigenetic and co-infection modifiers of RHD progression in Africaorphan  1 papers
What nobody has establishedWhether other endemic infections leave epigenetic marks that explain why African RHD progresses differently has never been tested.
Why it is still openIt is a speculative hypothesis from one review, needs African cohorts with infection histories, and would be expensive to run on a hunch.
The study that would close itGenome-wide methylation and histone-PTM profiling in an African RHD cohort stratified by co-infection history.
“Do epigenetic modifiers from other endemic infections contribute to the differential RHD progression rates observed in Africa?” — lumngwena-2022-pathophysiology-rhd-outstanding 2022
“What role do histone modifications play in RHD, and could global profiling of histone PTMs reveal mechanisms linking African tropical infections to RHD phenotypic variation?” — lumngwena-2022-pathophysiology-rhd-outstanding 2022
“How do non-coding transcripts influence RHD progression in the context of other tropical African infections?” — lumngwena-2022-pathophysiology-rhd-outstanding 2022
Whether streptococcal material persists in the rheumatic valveorphan  1 papers
What nobody has establishedOne Indian study looked for streptococcal DNA in valve tissue and did not find it; whether that means absence, clearance, or sampling too late is unresolved.
Why it is still openIt needs surgical tissue from patients at varied disease durations, and a negative result is hard to publish and harder to fund.
The study that would close itA multicentre valve-tissue study with sequencing confirmation, stratified by time since ARF and by prophylaxis history.
“Whether streptococcal DNA is present in valvular tissue of RHD patients cannot be determined from this study because the modest sample size limits generalizability and the absence of detecta” — parashar-2025-detection-streptococcal-deoxyribonucleic 2025
“It is unknown whether streptococcal DNA, if present, is detectable only in the early stages of valvular disease before host defenses and ongoing prophylaxis clear it, because the cohort repr” — parashar-2025-detection-streptococcal-deoxyribonucleic 2025
“Larger, multicenter studies with standardized methodologies, including sequencing confirmation, are needed to determine whether streptococcal persistence plays a role in RHD pathogenesis.” — parashar-2025-detection-streptococcal-deoxyribonucleic 2025
Digital tools for RHD care deliveryorphan  1 papers
What nobody has establishedWhether a case-management application functions on unstable rural connectivity, and whether it can be absorbed into national electronic medical records rather than living as a project silo, is untested.
Why it is still openDigital health projects are procured, not evaluated, and the app usually outlives neither its grant nor its connectivity.
The study that would close itA pragmatic evaluation of an RHD case-management app deployed through the national EMR in rural facilities, measuring uptime, use and follow-up completion.
“Whether the ACT application functions adequately under unstable internet connectivity in rural settings remains uncertain.” — minja-2023-evaluating-implementation-dynamic 2023
“Whether the ACT platform can be absorbed into broader electronic medical record expansion in Uganda is untested.” — minja-2023-evaluating-implementation-dynamic 2023
Which GAS strains are actually rheumatogenicorphan  1 papers
What nobody has establishedNo definitive identification of rheumatogenic strains has ever been made, so the concept underpins vaccine design without an evidence base.
Why it is still openIt requires isolates from the antecedent infection, which is exactly the event that is missed before ARF is diagnosed.
The study that would close itProspective emm-typing and whole-genome sequencing of GAS isolates from ARF cases versus uncomplicated pharyngitis in the same endemic population.
“A definitive identification of rheumatogenic GAS strains has not been established.” — lupieri-2025-rheumatic-heart-valve 2025

Grouped three separate times by three independent readers, each shown the same questions and asked to organise them differently. Where they agree, the signal is stronger than any one of them.