Skip to content

Add an organization-level /status/ view and incident directory #33

Description

@szmyty

Outcome

Add a first-class organization-level /status/ experience that truthfully presents public Ego Hygiene service/component availability, active incidents, maintenance, freshness, and monitoring coverage while preserving drill-down to canonical repository/provider evidence.

This is the organization-facing status surface missing from the visual-control-plane program in #30. It complements the broader /intelligence/ overview; it does not turn all repository, release, conformance, or audit evidence into an uptime score.

Primary user question

Which public Ego Hygiene services or components are currently affected, what is known versus unknown or stale, who owns the response, and where is the canonical incident/status evidence?

Architecture

Hygiene status contract (#62)
        +
Relay status workflow (relay#41)
        +
Observatory normalized fleet snapshot (observatory#23)
        ↓
egohygiene.io /status/
        ↔
Organization Intelligence navigation

The page is a projection. Canonical service intent, observations, incidents, and maintenance records remain with their owning sources.

Default experience

Show, in order:

  1. active incidents and planned maintenance;
  2. compact fleet/component posture with timestamp and freshness;
  3. unknown, stale, unreachable, unsupported, and unmonitored components;
  4. recently resolved incidents;
  5. ownership/contact and canonical evidence links;
  6. historical detail behind progressive disclosure.

Do not hide uncertainty behind a green banner.

Required behavior

  • Support healthy, degraded, partial outage, major outage, maintenance, unknown, stale, unreachable, intentionally unmonitored, unsupported, incompatible, and not-applicable states.
  • Keep runtime/service availability separate from CI, release/distribution, Hygiene, audit, security, and compliance posture.
  • Show represented snapshot revision, generated/observed times, stale-after behavior, and provider/source provenance.
  • Isolate malformed or unavailable repositories/components.
  • Preserve last-known context when current evidence is unavailable, but label it stale/provider-unavailable.
  • Deep-link to repository status pages, public incident records, owners, and authorized evidence.
  • Provide a static/degraded fallback that remains understandable if live/provider assets fail.
  • Integrate with the shared Organization Intelligence navigation and state vocabulary.
  • Treat all snapshot content as untrusted text and allowlist navigable URLs.

Privacy and disclosure

  • Public output must not reveal private repository/service existence, counts, internal hostnames, private topology, credentials, raw logs, or exploit-relevant diagnostics.
  • Visibility filtering occurs before aggregation.
  • A private/local organization build may consume separately authorized snapshots.
  • Missing/denied private sources must not be distinguishable in ways that leak their existence.

Accessibility and resilience

  • Mobile-first incident summary.
  • Keyboard-accessible filters, disclosure controls, and timelines.
  • Screen-reader meaningful state/impact text.
  • No color-only semantics.
  • Reduced-motion behavior.
  • Text/table alternative for any timeline or fleet visualization.
  • Useful no-JavaScript/static fallback.
  • One failed source cannot make the whole page unavailable.

Acceptance criteria

  • A canonical organization /status/ route is registered through egohygiene/hygiene#25.
  • The page consumes the versioned Observatory Reconcile the legacy issue backlog with repository ownership and native dependencies #23 snapshot rather than polling providers or repositories independently in the browser.
  • Active incidents, maintenance, current posture, monitoring applicability, freshness, and recent history are available.
  • Unknown/stale/unreachable/unmonitored/unsupported states never appear healthy.
  • Runtime status remains semantically separate from build, release, distribution, Hygiene, audits, security, and compliance.
  • Every material state preserves provenance and canonical drill-down where public.
  • One malformed/unavailable source cannot break the organization page.
  • Public output cannot leak private repository/service metadata.
  • Provider-unavailable and static fallback behavior is tested.
  • Mobile, keyboard, screen-reader, reduced-motion, no-color-only, and text-equivalent behavior is verified.
  • At least one monitored canary, one intentionally unmonitored fixture, one active incident, and one stale/provider-down fixture are proven.

Dependencies / related

Non-goals

  • Running monitors or editing incidents from the public page.
  • A universal organization health score.
  • Claiming an SLA or compliance certification.
  • Replacing repository/provider status sources.
  • Exposing private topology or raw telemetry.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions