GeoDemand 1.2 / Research summary

Event-linked search demand:
findings and limits

← Back to the interactive case study

GeoDemand provides a reproducible reported-flood-episode and search-interest analysis pipeline. This descriptive release keeps its provenance-checked weather subset separate from an exploratory flood archive. Their episode counts must not be added together.

Results

Standardized peak lift, measured in baseline population SD
Analysis Episodes Mean peak lift Median peak lift Median peak lag
Verified weather context 5 3.305 +19 days
Exploratory flood, partial-preserving 22 10.250 4.556 +3 days
Shared, partial-preserving 19 11.160 5.307 +4 days
Shared, alternative 19 11.576 4.919 +4 days

Dash: statistic not included in this summary. Shared rows are the same episodes, not additional observations.

The 22 exploratory peak lifts have quartiles of 2.975 and 14.365; peak lags range from −7 to +24 days. A few large normalized peaks influence the mean. In the 19 shared episodes, the mean partial-preserving-minus-alternative peak-lift difference is −0.416, the median difference is zero, and the mean absolute difference is 2.310 baseline SD. Every peak date agrees.

Sample and provenance

The official-only plan includes 20 episodes. Five weather-context units (six eligible repeat records) pass admission and numerical checks; flood-specific series fail sparse-signal or baseline-normalization rules. No acquired control-state records support valid contrasts. “Verified” refers to pipeline checks of hashes, sidecar lineage and values, not independent event confirmation or service authentication.

The exploratory archive covers 40 episodes. Twenty-two pass in the partial-preserving version; 19 pass in the alternative and are shared. Duplicate copies are counted once. Only 10 of the 22 belong to the original frozen 20-episode cohort. The archive likely came from PyTrends, despite an inherited manual-export label; retrieval times and attempt identities remain unresolved. All recorded partial flags in the primary version are False. The alternative omits those flags, so its findings are conditional on accepting unknown partial status. Version selection followed the audit and was not preregistered.

Method

Baseline: onset −28 to −8 inclusive, with at least 14 valid days. Lead: −7 to −1. Immediate response: onset through event end +2. Early recovery: end +3 to +7. Extended recovery: end +8 to +28. Each response phase requires 80% valid coverage. Zero fraction ≥0.5 and baseline population SD ≤0.001 exclude normalized responses. Thresholds are unchanged; no denominator epsilon or outcome imputation is used.

Peak lift equals the response-window maximum minus baseline mean, divided by baseline population SD. Earliest tied peaks define timing. The official analysis respects partial flags and averages eligible repeats within requests. The interactive curves are equal-weight daily means standardized per episode and displayed from onset −28 to +28. Peak metrics use the full event-end-anchored window.

Robustness and validation

In the verified subset, moving the baseline end to −1 admits six units across five episodes. Overall mean peak lift becomes 6.980, but this includes a changed sample; the paired mean change on shared units is +0.206. Requiring complete two-repeat groups leaves one unit, which cannot establish general repeat reliability.

Two executions of each local analysis produced byte-identical outputs. The release checks passed 250 tests with two skips on Python 3.11, 3.12 and 3.13, alongside formatting, lint and type checks. Computational reproducibility does not resolve incomplete acquisition provenance.

Interpretation

These findings are descriptive. They do not establish causal effects, representative public demand, physical event severity or predictive skill. Google Trends indices are relative rather than search counts. Sparse-series selection, variable window duration and uncertain archive history constrain interpretation. Missing or excluded series are not evidence of absent public response. Identical peak dates across related archive versions are not independent replication.

The fixed prediction model and training-mean comparator remain gated off because the verified analysis has fewer than 20 independent event groups and insufficient concept coverage. Exploratory eligibility does not substitute for the frozen prediction requirements.

Scope and access

The bounded descriptive study is complete: five verified weather-context units, 22 exploratory flood units and a 19-episode version comparison. Additional acquisition or prediction would be a new study. This public-facing summary contains aggregate findings; raw acquisitions, local paths and data-bearing release packages remain outside the website.

GeoDemand project source · Explore the methodology