Data standards and guidance
How we choose, document and publish open data so it stays trustworthy and lightweight.
Our role
We are an independent compiler and guide — not KNBS, IEBC or any ministry. Official producers remain the source of truth for formal statistics. Our open-data pages use GDS-inspired structure for clarity; we are not a government open-data portal.
Publicly useful public information should be findable, understandable and reusable.
What we publish
- Structured extracts we maintain (Supabase / Sanity)
- Plain-language summaries and small visual aggregates
- CSV/JSON downloads where an export route exists
- Links to official portals we do not host
What we never publish as open data
- Contact messages, feedback or usefulness votes
- Page-view or search analytics
- User accounts or admin content
- Draft or unpublished editorial material
- Personal data beyond public office-holder directories
Metadata on every dataset
Each dataset page should state:
- Title and plain-language description
- Publisher (original) and compiler (CitizenGuide.KE)
- Temporal and geographic coverage
- Update frequency
- Licence / reuse terms
- Formats (CSV, JSON) when available
- Field dictionary and known limitations
- Source links where we have them
Formats and access
- HTML summary — figures and accessible bar tables (small payloads)
- CSV — preferred for spreadsheets and analysis
- JSON — preferred for software
- Catalogue API —
/api/data/datasets
Licence and credit
Free to reuse for any purpose with credit to CitizenGuide.KE as compiler and to the original publisher where known. Not an official Government of Kenya statistics release.
Quality principles
- Prefer stable public sources over scraping private systems
- Mark snapshots (for example 2022 polling stations) clearly
- Do not imply KNBS, IEBC or other endorsement without agreement
- Keep interactive pages light; put bulk data on downloads
- Follow open-data maturity ideas (documentation, reuse, accessibility) without claiming formal certification
Versioning
Large electoral or census files should be treated as dated releases. When we replace a major extract, dataset notes should say what changed. Prefer citing the temporal coverage field on the dataset page.