News Intelligence — Forensic Press Archive
A press archive of our own, with a forensic dossier for every story: the text, the images, a full-page screenshot, the HTML and a structured version — each file carrying a SHA-256 hash.
Overview
There are 76 registered sources (40 Brazilian, the rest international and Latin American), collected in accordance with robots.txt and crawl-delay; when an outlet edits a story, the previous version stays on file together with the screenshot and the HTML captured at that moment. On top of this archive runs a disinformation triage that produces a prioritized queue with the evidence attached — not a true-or-false verdict. When the signal is too strong to ignore and too weak to conclude, the story goes to human review with the reason stated; and the interface says so whenever the AI layer did not run on a given story, so a partial analysis never looks like a complete one.
Capabilities
- A forensic dossier per story: text, images, screenshot and HTML with SHA-256 and a perceptual hash
- Deduplication by URL and edit history, with proof of the previous version
- Three-layer triage: corroboration across outlets, matching against published fact-checks, and an AI reading
- Human review queue with a declared abstention zone
- Search and filtering panel that explains where each figure comes from and what it does not cover
Use Cases
- Preserve proof of what an outlet published, ahead of a silent correction or a takedown
- Follow watchlist terms across the press, with a daily sweep of the public Google News feed
- Prioritize reporting work by the stories that match corrections already published by fact-checking organizations
- Bring press context — with source, date and hash — into an ongoing investigation
Integrations
- OpenSearch and MinIO/S3 (archive and preserved files)
- OODA Reports / Investigate
- REST API: ingestion, search, statistics and analysis triggering, with per-organization permissions
SLA & Guarantees
Hourly triage · evidence with SHA-256 · no automatic verdict