Put kanman on call: services and signal sources

Add the services your team runs, connect uptime checks, alerting tools and log sources, and choose what kanman may do when something degrades.

This guide puts kanman on call for a service: you describe the service, connect where alerts come from and where errors can be read, and decide what kanman may do. How incidents work is explained in Incidents and on-call.

You need to be an owner or admin of the workspace to add services and sources. Every member can see them, test a connection and work on incidents.

1. Add a service

  1. Open the team and click the Services tab.
  2. Click New service.
  3. Enter a name, for example “Checkout API”, and an optional description.
  4. Tick the repositories the service is built from. Only repositories the team is allowed to work on are offered; add more under Settings > Repositories of the team.
  5. List the environments, for example production with its URL.
  6. Click Save. The service page opens.

2. Choose what kanman may do

On the service page, under On-call settings:

Setting Default Meaning
Observe only Off kanman watches, investigates and reports, but never acts.
Start expedite fixes on its own Off For code-caused incidents kanman starts a fix right away. Off: it asks in Decisions first.
Acknowledge and resolve alerts in the source Off PagerDuty and Opsgenie only. Off: kanman never changes the alert in your tool.
Lowest severity that opens an incident High Lower alerts are recorded only.
On-call hours for kanman Always Outside these hours kanman records alerts but opens no incident.
Who gets notified Nobody Workspace members, notified in kanman when an incident opens and when it is resolved.
Incident places All channels of the team The team’s Slack or Microsoft Teams places where incidents are posted.

Click Save. All settings are described in On-call settings.

Incident places in Slack or Microsoft Teams

kanman posts incidents to the places your team talks to it in. Connect Slack or Microsoft Teams and add the team’s channels under the team’s Settings > Communication, as described in Slack and Microsoft Teams. Then, under Incident places in the on-call settings, tick the channels that should get this service’s incidents, for example a dedicated incidents channel. With nothing ticked, every channel of the team gets them.

Each incident is one thread per place: kanman opens it with the title, severity and service, adds every finding and action as a reply, and posts the resolution at the end. Incident posts appear in the team’s chat history in the app like every other message kanman sends.

3. Connect signal sources

On the service page, click Add source, choose the kind, give it a name and fill in the settings. After you save an alerting tool, kanman shows the webhook URL and the secret once. The URL has this form:

https://api.kanman.ai/functions/v1/signal-webhook/<source token>

Copy both into your tool right away. kanman stores the secret encrypted and never shows it again; Rotate secret creates a new one (the old one stops working at once). If a source feeds several services, tick them in the source’s settings.

Use Test connection on any source: an uptime check runs once, a log source reads the last 15 minutes, an alerting tool shows whether alerts have arrived. Testing never opens an incident.

Uptime check

Settings: URL, method, interval (60 seconds or more), timeout, expected status codes (for example 200 or 2xx), text the response must contain, request body, how many failures in a row open an incident (default 2), and the severity. Add request headers for protected endpoints, for example Authorization; their values are stored encrypted. kanman only checks public addresses. The first passing check after a failure resolves the incident.

Prometheus Alertmanager

Add a receiver and route your alerts to it:

receivers:
  - name: kanman
    webhook_configs:
      - url: https://api.kanman.ai/functions/v1/signal-webhook/<source token>
        send_resolved: true
        http_config:
          authorization:
            type: Bearer
            credentials: <secret>

kanman reads severity and service (or job, app) from the alert labels and summary and description from the annotations.

Grafana alerting

Create a contact point of type Webhook with the URL. Under Optional Webhook settings, set the authorization header scheme to Bearer and the credentials to the secret. Alternatively, enter the secret as HMAC signature secret. Use the contact point in a notification policy.

Datadog

  1. In Integrations > Webhooks, add a webhook named kanman with the URL.
  2. Set Custom Headers to {"X-Kanman-Token": "<secret>"}.
  3. Use this payload:
{"id":"$ID","alert_id":"$ALERT_ID","aggregate":"$AGGREG_KEY","title":"$EVENT_TITLE","transition":"$ALERT_TRANSITION","priority":"$ALERT_PRIORITY","alert_type":"$ALERT_TYPE","body":"$EVENT_MSG","link":"$LINK","tags":"$TAGS","date":"$DATE"}
  1. Mention @webhook-kanman in the monitors that should reach kanman. Recoveries resolve the incident.

New Relic

Create a webhook destination with the URL and the secret as Bearer token, then add it to a workflow with the default webhook payload. Closed issues resolve the incident.

PagerDuty

  1. In Integrations > Generic Webhooks (v3), add a subscription for the service with the URL.
  2. Select the events incident.triggered, incident.resolved, incident.reopened, incident.escalated and incident.priority_updated.
  3. PagerDuty shows its own signing secret: edit the source in kanman and paste it into Signing secret from the tool.
  4. Optional, to let kanman acknowledge and resolve incidents: add a REST API key and the email of the PagerDuty user kanman acts as, then turn on Acknowledge and resolve alerts in the source for the service.

Opsgenie

Add a Webhook integration with the URL and a custom header X-Kanman-Token with the secret. Created and closed alerts reach kanman. Optional: add an API key and choose the region to let kanman acknowledge and close alerts.

Sentry

  1. In Settings > Developer Settings, create an internal integration with the URL as webhook URL, turn on Alert Rule Action and the issue events.
  2. Paste the integration’s Client Secret into Signing secret from the tool of the source.
  3. Add the action “Send a notification via kanman” to the alert rules that should reach kanman. Issue alerts and metric alerts are supported; resolving the issue in Sentry resolves the incident.

Amazon CloudWatch

  1. Create an HTTPS subscription on the SNS topic your alarms publish to, with the URL plus ?key=<secret>: https://api.kanman.ai/functions/v1/signal-webhook/<source token>?key=<secret>
  2. kanman confirms the subscription by itself and checks the AWS signature of every message. Optionally enter the topic ARN in the source to accept only that topic.
  3. Set the severity in the alarm description, for example [critical] Checkout 5xx; untagged alarms count as high. ALARM opens, OK resolves.

Azure Monitor

In the action group, add a Webhook action with the URL plus ?key=<secret> and turn on the common alert schema. Fired alerts open, resolved alerts resolve. Sev0 is critical, Sev1 high, Sev2 medium, Sev3 low.

Atlassian Statuspage

Add a webhook subscriber with the URL plus ?key=<secret>. Component outages and incidents on your status page reach kanman; “operational” or a resolved status page incident resolves it.

Other tools

Send alerts in kanman’s generic format, signed with the secret. The format and the signature are described in On-call settings.

Logs, errors and traces

Log sources do not send anything; kanman reads them while it investigates an incident.

Kind Settings Credentials
Grafana Loki API URL, query (LogQL selector, for example {app="checkout"}), tenant Bearer token, or user name and password
Datadog Logs API URL of your site (for example https://api.datadoghq.eu), search query API key and application key with read access to logs
Amazon CloudWatch Logs Region, log group, filter pattern Access key with permission to filter the log group’s events (read only)
Sentry Organization, project, API URL for self-hosted Sentry Auth token with read access to issues and events
Kubernetes Namespace, label selector, container None: your self-hosted runner reads the logs with its own access, inside your network

kanman only reads errors and warnings around the start of the incident and keeps the error lines it reports in the incident’s timeline.

4. Try it

Send a test alert from your tool, or let an uptime check point at a page that answers an error. Within a minute the incident appears under Incidents and in your channel, and kanman starts posting what it finds.

Last updated: January 1, 0001

Open kanman