FALSIFYR ATTACK

Claim card

Claim: treatment -> score
Formula: score ~ treatment + age + baseline_score
Estimate: 0.443
p-value: 0.041
Confidence interval: [0.0251, 0.861]

Survival verdict

COLLAPSES 0/100

The survival score is a heuristic summary, not a formal probability that the claim is true.

Smallest kill

Influential row deletion: remove 1 row -> p = 0.069
Rows: 43
Method: ranked

Attack leaderboard

RankFamilyAttackStatusp-valueExplanation
1row_deletionInfluential row deletionkilled0.069remove 1 row -> p = 0.069
2measurement_errorBinary label flipkilled0.06450% kill rate after flipping 2.5% of treatment labels
3missingnessMedian imputationkilled0.109Median imputation -> p = 0.109
4missingnessMean imputationkilled0.097Mean imputation -> p = 0.097
5splitRandom train/test split stabilitykilled0.07550% kill rate after holding out 20.0% of rows
6splitStratified train/test split stabilitykilled0.05750% kill rate after holding out 20.0% of rows
7missingnessMode/explicit-missing imputationkilled0.097Mode/explicit-missing imputation -> p = 0.097
8measurement_errorOutcome noisekilled0.05750% kill rate at noise SD = 0.35 x outcome SD
9standard_errorHC3 robust standard errorskilled0.051HC3 p = 0.051
10placeboFake predictor placebosurvived0.5790.0% of random fake predictors were at least as strong as original
11placeboTreatment-label permutationsurvived0.5222.5% of placebo labels were at least as strong as original
12measurement_errorPredictor noise: agesurvived0.041no 50% kill rate up to noise SD = 0.35 x age predictor SD
13standard_errorHC2 robust standard errorssurvived0.045HC2 p = 0.045
14standard_errorHC1 robust standard errorssurvived0.044HC1 p = 0.044
15specificationBounded specification searchsurvived0.039claim survived 3 bounded specifications; max p = 0.039
16covariateDrop agesurvived0.039drop age -> p = 0.039
17standard_errorHC0 robust standard errorssurvived0.039HC0 p = 0.039
18covariateDrop baseline_scoresurvived0.033drop baseline_score -> p = 0.033
19standard_errorBootstrap uncertaintysurvived<0.001bootstrap p approx = <0.001

Family sections

Row deletion / influence

1 killed, 0 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
1row_deletionInfluential row deletionkilled0.069remove 1 row -> p = 0.069

Measurement error

2 killed, 1 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
2measurement_errorBinary label flipkilled0.06450% kill rate after flipping 2.5% of treatment labels
8measurement_errorOutcome noisekilled0.05750% kill rate at noise SD = 0.35 x outcome SD
12measurement_errorPredictor noise: agesurvived0.041no 50% kill rate up to noise SD = 0.35 x age predictor SD

Missing data

3 killed, 0 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
3missingnessMedian imputationkilled0.109Median imputation -> p = 0.109
4missingnessMean imputationkilled0.097Mean imputation -> p = 0.097
7missingnessMode/explicit-missing imputationkilled0.097Mode/explicit-missing imputation -> p = 0.097

Split stability

2 killed, 0 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
5splitRandom train/test split stabilitykilled0.07550% kill rate after holding out 20.0% of rows
6splitStratified train/test split stabilitykilled0.05750% kill rate after holding out 20.0% of rows

Standard-error / uncertainty

1 killed, 4 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
9standard_errorHC3 robust standard errorskilled0.051HC3 p = 0.051
13standard_errorHC2 robust standard errorssurvived0.045HC2 p = 0.045
14standard_errorHC1 robust standard errorssurvived0.044HC1 p = 0.044
17standard_errorHC0 robust standard errorssurvived0.039HC0 p = 0.039
19standard_errorBootstrap uncertaintysurvived<0.001bootstrap p approx = <0.001

Placebo

0 killed, 2 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
10placeboFake predictor placebosurvived0.5790.0% of random fake predictors were at least as strong as original
11placeboTreatment-label permutationsurvived0.5222.5% of placebo labels were at least as strong as original

Specification

0 killed, 1 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
15specificationBounded specification searchsurvived0.039claim survived 3 bounded specifications; max p = 0.039

Covariate dependence

0 killed, 2 survived, 0 unavailable.

RankFamilyAttackStatusp-valueExplanation
16covariateDrop agesurvived0.039drop age -> p = 0.039
18covariateDrop baseline_scoresurvived0.033drop baseline_score -> p = 0.033

Fragility curves

Influential row deletion

StepRows removedEstimatep-valueKilledRows
110.3550.069yes43

Binary label flip

Perturbation levelKill ratep-value
0.0140.0%0.043
0.02563.3%0.064
0.0566.7%0.071
0.156.7%0.071

Random train/test split stability

Perturbation levelTraining rowsHoldout rowsKill rateEstimatep-value
0.2601473.3%0.4330.075
0.3522283.3%0.4160.121

Stratified train/test split stability

Perturbation levelTraining rowsHoldout rowsKill rateEstimatep-value
0.2601460.0%0.4670.057
0.3522270.0%0.4160.098

Outcome noise

Perturbation levelKill ratep-value
0.0523.3%0.043
0.113.3%0.044
0.236.7%0.039
0.3556.7%0.057

Fake predictor placebo

ReplicateEstimatep-valueKilled
10.03310.882no
20.2460.266no
3-0.3280.131no
4-0.3040.164no
50.2790.200no
6-0.03250.883no
7-0.2510.278no
8-0.001140.996no
9-0.06760.757no
10-0.2450.260no
11-0.006860.976no
120.1630.459no

Showing first 12 of 80 curve rows.

Treatment-label permutation

ReplicateEstimatep-valueKilled
10.1130.609no
2-0.1430.523no
30.2480.262no
40.2120.331no
5-0.0590.787no
6-0.3990.066no
70.09650.662no
8-0.1670.444no
9-0.1720.450no
100.03710.866no
11-0.2870.192no
12-0.008020.971no

Showing first 12 of 80 curve rows.

Predictor noise: age

Perturbation levelKill ratep-value
0.050.0%0.041
0.10.0%0.041
0.20.0%0.041
0.350.0%0.041

Limitations and caveats

A killed claim is fragile under the named attack. It does not prove the original claim is false, and it does not say the perturbed analysis is preferable.

This report summarizes attacks that falsifyr could run from the fitted object and recoverable data. Unsupported or unavailable attacks are listed below so absence of evidence is not mistaken for robustness.

No requested attack returned an unavailable status.

Reproducibility appendix

Call
attack(model = fit, term = "treatment", attacks = c("row_deletion", "standard_error", "covariate_drop", "missingness", "measurement_error", "placebo", "specification", "split"), intensity = "fast", seed = 11)
Model class
lm
Attack settings
seed = 11; intensity = fast; profile = default; alpha = 0.05; kill rule = p_over_alpha; alternative = two.sided
Attack families
row_deletion, standard_error, covariate_drop, missingness, measurement_error, placebo, specification, split
Unknown requested families
NA
Package version
0.1.0
R version
4.4.1
Platform
x86_64-w64-mingw32
Session info
Show session info
R version 4.4.1 (2024-06-14 ucrt)
Platform: x86_64-w64-mingw32/x64
Running under: Windows 11 x64 (build 26200)

Matrix products: default


locale:
[1] C
system code page: 65001

time zone: America/New_York
tzcode source: internal

attached base packages:
[1] stats     graphics  grDevices utils     datasets  methods   base

other attached packages:
[1] falsifyr_0.1.0

loaded via a namespace (and not attached):
 [1] utf8_1.2.4         RColorBrewer_1.1-3 R6_2.6.1           tidyselect_1.2.1
 [5] lattice_0.22-6     farver_2.1.2       magrittr_2.0.3     gtable_0.3.6
 [9] zoo_1.8-14         glue_1.8.0         tibble_3.2.1       pkgconfig_2.0.3
[13] dplyr_1.1.4        generics_0.1.3     lifecycle_1.0.4    ggplot2_4.0.0
[17] cli_3.6.5          S7_0.2.0           fansi_1.0.6        scales_1.4.0
[21] sandwich_3.1-1     grid_4.4.1         vctrs_0.6.5        compiler_4.4.1
[25] lmtest_0.9-40      pillar_1.9.0       rlang_1.1.7       
Started
2026-07-10 22:45:10 UTC
Finished
2026-07-10 22:45:11 UTC
Elapsed seconds
0.978