Skip to main content
Research and methods

Evaluation and limitations

We distinguish implementation checks from evidence that a system works in an operational setting.

Implemented evaluation protocol

Versioned source fixtures cover valid reports, non-matching interests, duplicates, updated reports, malformed inputs, future timestamps and source outages. Database and API tests check tenant isolation, review versioning and exports. These checks assess software behaviour, not operational detection accuracy.

Measured results

No independently validated incident precision, recall, early-warning advantage or campaign uplift is claimed. A public live observation is evidence of source retrieval, not a benchmark. Release-specific software test results and reproducible procedures are recorded in the repository evaluation report.

Next operational evaluation

Build an appropriately licensed event and non-event dataset with independent labels. Split by event and time, then measure false alerts, precision and recall where labels support them, citation correctness, source lag, latency and resource use against a frozen baseline.

Known limits

Coverage is limited to configured sources. Missing alerts are not evidence of safety. Publisher updates, outages, duplicates and incomplete geography affect relevance. Human reviews can be mistaken. The system is decision support and has not been validated for critical operational decisions.