A (host, service) pair is flapping once it exceeds flap_count warning or critical notifications within flap_interval minutes. The notification that trips the state carries "Now flapping!! No more messages!" and every later one is dropped, including RECOVER. The state ends silently flap_interval minutes after a RECOVER, provided no further alert arrived meanwhile. Hooked into notify.send_notification, the single choke point for channel delivery, so only outbound notifications are suppressed — eventlog keeps recording, leaving the journal and /log with the full history of the flap. Threshold alerts key on their metric path, so a flapping disk check cannot silence a CPU alert; connectivity, boot and shutdown events key on the host itself. State lives at module level in flap.py and is never pickled. Flapping pairs surface in Host.stateinfo() and render as an amber badge on the live dashboard. Config: flap_count (5), flap_interval (10 minutes), 0 in either disables; both editable on the settings page. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
11 KiB
Notification System
Overview
Notifications are dispatched to the owner and managers of a host, each via their own configured notification channels. Channel definitions are global; users reference them by name. No users configured → no notifications sent.
Architecture
Alert event (udp.py / threshold.py)
└─ notify.send_notification(host_name, Notification)
├─ flap.observe(host, service, level) → pass | trip | suppress
├─ look up host.owner + host.managers
├─ for each user → user.notification_channels
└─ for each channel → _dispatch_to_channel (filtered by min_level)
Every notification carries:
- title —
[LEVEL] hostname(e.g.[CRITICAL] webserver01) - body — detail message (metric value, threshold, duration)
- url — link to the plugin metrics page (
{base_url}/plugins#{hostname}) - level —
RECOVER | WARNING | CRITICAL | INFO - service — flap-detection key within the host (empty = the host itself)
Configuration
Base URL
Set base_url so notification links point to your hbd instance:
base_url: https://hbd.example.com
Channel definitions
Channels are defined under notification_channels. Each entry specifies a delivery type and its credentials. Ownership is the single visibility signal:
| Field | Default | Description |
|---|---|---|
owner |
(absent) | Owning username. Present = private to that user; absent = global. |
min_level |
WARNING |
Minimum alert level this channel receives. |
(The former private flag is retired; leftover private keys are ignored and dropped on the next edit.)
Global channels (no owner; set in the config file or by an admin) can be selected by all users but edited only by admins:
notification_channels:
pushover_ops:
type: pushover
token: your-app-token
user: your-user-key
min_level: WARNING
email_ops:
type: email
recipients: [ops@example.com]
sender: hbd@example.com
smtp_server: smtp.example.com
smtp_port: 587
smtp_user: hbd@example.com
smtp_password: secret
min_level: WARNING
matrix_oncall:
type: matrix
homeserver: https://matrix.example.org
access_token: syt_xxx
room_id: "!abc:matrix.example.org"
min_level: CRITICAL
sms_oncall:
type: sms_voipms
api_user: me@example.com
api_password: secret
did: "5551234567"
dst: "5559876543"
min_level: CRITICAL
signal_ops:
type: signal
cli_path: /usr/local/bin/signal-cli
user: +12025551234
recipient: +12025559999
mattermost_devops:
type: mattermost
host: mattermost.example.com
token: webhook-token
channel: devops-alerts
username: heartbeat-bot
User-created channels are written by authenticated users through the API, their profile page, or the settings page. They carry an owner field and are private to that user:
notification_channels:
alice_personal:
type: pushover
token: personal-token
user: personal-key
owner: alice # private to alice
Channel visibility
| Channel | Who can see / select it | Who can edit it |
|---|---|---|
No owner (global) |
All users | Admins |
owner set (private) |
Only the owner |
The owner |
| Any channel | Admins always see everything | Admins |
Admins can promote a private channel to global by clearing its owner on the settings page, or demote a global channel by assigning an owner.
Users with notification channels
Each user lists which channels they receive notifications on. Users can manage their own selection from the profile page:
users:
alice:
full_name: Alice Smith
password: pbkdf2:sha256:...
admin: true
notification_channels: [pushover_ops, email_ops]
bob:
full_name: Bob Jones
password: pbkdf2:sha256:...
notification_channels: [sms_oncall, matrix_oncall]
Host access — owner and managers
Notifications for a host go to its owner and all managers:
hosts:
webserver01:
owner: alice # receives all notifications for this host
managers: [bob] # also receives notifications
threshold_config: default
watch: true # bold in dashboard (cosmetic only)
dyndns: false
dbserver01:
owner: alice
managers: [bob]
threshold_config: database
dyndns: false
watch: true only affects display (bold name in the live dashboard). Notifications are now controlled entirely by owner/managers.
Channel Types
min_level filtering
Every channel accepts an optional min_level field:
| Value | Channels receive |
|---|---|
WARNING (default) |
WARNING, CRITICAL, RECOVER |
CRITICAL |
CRITICAL only (and RECOVER) |
RECOVER is always passed through — you don't want to miss a recovery.
pushover
Sends push notifications via Pushover. Includes title, body, and a clickable URL.
type: pushover
token: your-app-token # Required: Pushover application token
user: your-user-key # Required: Recipient's user key
min_level: WARNING
Sends via SMTP. Subject = title, body = message + URL on final line.
type: email
recipients: [ops@example.com, oncall@example.com]
sender: hbd@example.com
smtp_server: smtp.example.com
smtp_port: 587 # 587 = STARTTLS (default), 465 = SSL
smtp_user: hbd@example.com
smtp_password: secret
min_level: WARNING
matrix
Sends a formatted HTML message to a Matrix room via matrix-nio.
type: matrix
homeserver: https://matrix.example.org
access_token: syt_xxx # Bot account access token
room_id: "!abc:matrix.example.org"
min_level: WARNING
Setup:
- Create a bot Matrix account
- Obtain its access token (Element → Settings → Help & About → Access Token)
- Invite the bot to the target room and note the room ID
sms_voipms
Sends SMS via the voip.ms REST API. Message is truncated to 160 characters.
type: sms_voipms
api_user: me@example.com # voip.ms account email
api_password: secret # voip.ms API password
did: "5551234567" # Your voip.ms DID (sending number)
dst: "5559876543" # Destination number
min_level: CRITICAL
signal
Sends via signal-cli.
type: signal
cli_path: /usr/local/bin/signal-cli
user: +12025551234 # Your registered Signal number
recipient: +12025559999 # Recipient number
min_level: WARNING
Setup:
signal-cli -u +12025551234 register
signal-cli -u +12025551234 verify CODE
mattermost
Sends via Mattermost incoming webhook. Message is formatted as Markdown.
type: mattermost
host: mattermost.example.com
token: your-webhook-token
channel: devops-alerts
username: heartbeat-bot # Optional: display name
icon: https://…/icon.png # Optional: bot icon URL
min_level: WARNING
Notification events
| Source | Level | Title example | Body example |
|---|---|---|---|
| Host overdue | CRITICAL | [CRITICAL] webserver01 |
IPv4 overdue |
| Host recover | RECOVER | [RECOVER] webserver01 |
IPv4 back after being overdue for 5:23 |
| Host boot | INFO | [INFO] webserver01 |
webserver01 booted |
| Host shutdown | INFO | [INFO] webserver01 |
IPv4 shutdown |
| Threshold breach | WARNING/CRITICAL | [CRITICAL] webserver01 |
cpu_percent = 95.2 (threshold: > 90.0) |
| Threshold reminder | CRITICAL | [REMINDER/CRITICAL] webserver01 |
REMINDER (CRITICAL): … ongoing for 3600s |
| Connection issue | WARNING | [WARNING] webserver01 |
new address detected … |
Reminder notifications (re-notify) are sent only for CRITICAL level alerts.
Flap detection
A check that toggles between OK and alerting produces a notification per swing. Flap detection silences it after the first few.
A (host, service) pair is flapping once it exceeds flap_count WARNING/CRITICAL
notifications within flap_interval minutes. Threshold alerts key on their metric path,
so a flapping disk check does not silence an unrelated CPU alert; connectivity, boot and
shutdown events key on the host itself.
flap_count: 5 # notifications within the window that trip flapping (0 disables)
flap_interval: 10 # minutes — both the counting window and the quiet window
Lifecycle:
| Event | Effect |
|---|---|
Alerts 1..flap_count within the window |
Delivered normally |
Alert flap_count + 1 |
Delivered with Now flapping!! No more messages! appended to the body |
| Every notification after that | Dropped — including RECOVER and INFO |
| RECOVER while flapping | Dropped, and starts the flap_interval quiet window |
| WARNING/CRITICAL during the quiet window | Restarts the quiet window; still flapping |
| Quiet window elapses | Flapping ends silently — no notification |
Only outbound notifications are suppressed. notify.eventlog keeps recording every event,
so the journal and the /log page retain the full history of the flap.
Flapping pairs appear in each host's stateinfo() under flapping (a list of service
keys; "" means the host itself) and render as an amber flapping badge next to the
host name on the live dashboard, with the affected services in its tooltip.
State lives at module level in hbd/server/flap.py and is never pickled — a server restart
starts every check with a clean slate. Hosts with watch: false never reach the
notification path, so they never flap.
API reference
send_notification(host_name, notif) -> dict
Main entry point. Dispatches to owner + managers.
from hbd.server.notify import send_notification, Notification
send_notification(
"webserver01",
Notification(
title="[CRITICAL] webserver01",
body="cpu_percent = 95.2 (threshold: > 90.0)",
level="CRITICAL",
url="https://hbd.example.com/plugins#webserver01",
),
)
Returns {channel_name: bool} for each channel dispatched.
setup(cfg, loop=None)
Called once at startup from main.py. Pass the running asyncio event loop so Matrix sends work correctly.
Troubleshooting
No notifications sent:
- Check that users are configured (
users:section in yaml) - Check that the host has an
ownerormanagersset - Check that users have
notification_channelslisted - Check that the channel names in user config match keys under
notification_channels: - If a user can't select a channel, check whether it has an
ownerother than that user
min_level filtering too aggressive:
- Default is
WARNING— both WARNING and CRITICAL are sent - Set
min_level: WARNINGexplicitly if you were expecting warnings but set CRITICAL
Matrix sends time out:
- Verify the access token is valid and the bot is in the room
matrix-niomust be installed:pip install matrix-nio
voip.ms SMS fails:
- Enable the API in your voip.ms account (Account → API)
- Verify the DID is SMS-capable in your voip.ms account
Signal not found:
- Specify full
cli_path - Run
signal-cli -u +NUMBER receiveto sync trust store
Email authentication failed:
- Use app-specific passwords for Gmail/Fastmail
- Verify port: 587 for STARTTLS, 465 for SSL
Pushover 400 errors:
- Double-check
token(app) anduser(user key) — they are different values