Incidents and on-call

kanman can be on call for the services your team runs: it opens an incident when something degrades, investigates, acts within your rules and closes the loop with a summary and follow-up stories.

Running what you ship is part of your team’s work. kanman can be on call for the services a team owns. When a service degrades, kanman opens an incident, tells the people on call, finds out what changed, and does what your settings allow: start a fix with a regression test, propose a revert, or recommend the next steps to a person. It never changes production outside the code.

To set this up, see Put kanman on call. All settings are listed in On-call settings.

Services and signal sources

A service is something your team runs, for example the checkout API. Each service lists the repositories it is built from and its environments (production, staging), so kanman knows where to look and what it may fix.

A signal source is where kanman hears that something degrades or reads what went wrong:

Kind What it does
Uptime check kanman calls a URL or API on a schedule (method, headers, expected status, text in the response, timeout) and reports when it fails several times in a row.
Alerting tools Prometheus Alertmanager, Grafana alerting, Datadog, New Relic, PagerDuty, Opsgenie, Sentry, Amazon CloudWatch, Azure Monitor and Atlassian Statuspage send their alerts to kanman. Anything else can use the signed generic format.
Logs, errors and traces kanman reads errors and warnings from Grafana Loki, Datadog Logs, Amazon CloudWatch Logs and Sentry while it investigates. Kubernetes logs are read by your self-hosted runner, inside your network.

One source can feed several services. Every source has its own address and secret; kanman refuses alerts that do not carry the right secret.

What happens when something degrades

  1. An incident opens. The first alert at or above the service’s severity threshold opens an incident with the severity the tool reported. Further alerts with the same identity, and other alerts of the same service within 30 minutes, join the same incident instead of opening new ones.

  2. kanman tells people. It posts the incident to the team’s incident channel in Slack or Microsoft Teams and notifies the people you chose. Everything it learns afterwards goes into the same thread.

  3. kanman investigates and posts each finding as soon as it has it:

    • merges and deploys before the first alert, and whether one of the merges was kanman’s own run
    • whether CI on the main branch is green or red
    • the most frequent error in the logs and where in the code it happens
    • configuration files changed by recent merges
    • stories linked to the recent changes

    When nothing explains the failure, kanman says so: “I don’t know the cause yet.”

  4. kanman acts within your settings:

    What kanman found What it does
    The error points into your code and a recent merge changed that file It opens a revert pull request and asks you in Decisions whether to merge it. Reverts are never merged without you.
    The error points into your code It files an expedite story and starts a fix, or asks first if the service does not allow it to start fixes on its own. The fix run reproduces the error with a regression acceptance spec before it changes anything, then goes through the outcome gate like any other story.
    A configuration or infrastructure problem It asks the person on call in Decisions, with its recommendation. Changes to production outside the code are always a decision for a person.
    An outside provider or an unknown cause It recommends next steps to the person on call.

    Merging and deploying the fix follow your team’s merge policy. kanman does not deploy.

  5. kanman closes the loop. When the alerts recover (the alerting tool reports them resolved, or the uptime check passes again), kanman closes the incident, writes a short summary with a timeline, lists it in the team’s weekly report, and drafts follow-up stories in Intake, for example “Catch this earlier” or “Harden the failing code path”. You review and approve them like any other story.

Expedite stories are never stopped by budgets or throttles; see Budgets.

The incident page

Open Incidents on a team to see open and resolved incidents. An incident’s page shows:

  • severity, status, cause, how many alerts joined, how long it lasted
  • the timeline: every alert, finding, action and decision, in order, with links to the fix story, its run, and the decisions
  • the summary and the follow-up stories once it is resolved

Anyone in the workspace can add a note, resolve the incident by hand, or close it as a false alarm.

Safety

  • Observe only is a hard switch per service: kanman watches, investigates and reports, but never starts a fix, proposes a revert or touches the alert.
  • kanman never silences, acknowledges or resolves alerts in your alerting tool unless you allow it for the service (PagerDuty and Opsgenie).
  • Outside the on-call hours you set for kanman, it records alerts but opens no incident; your people on call take over.
  • Everything kanman does is in the audit log: received alerts, opened and resolved incidents, changed settings and sources, and every action in an alerting tool.
  • Secrets and credentials of sources are stored encrypted and never shown again after you save them.

Last updated: January 1, 0001

Open kanman