heartbeat

Public Access

Author	SHA1	Message	Date
andreas	ca908ee967	fix: remove unused imports from oauth module and tests	2026-05-08 13:26:51 -04:00
andreas	73c697b6c5	feat: add oauth module skeleton and is_enabled() Add hbd/server/oauth.py with OAuthError, _gitea_cfg(), and is_enabled() to detect when all three required Gitea OAuth2 config keys are present. Add "oauth": {} default to SERVER_DEFAULTS in config.py. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 13:24:27 -04:00
andreas	b81a0d2a6c	plugins: persist owner chip in glance strip across JS updates Store owner in data-owner attribute; updateHostHeader always prepends it so it survives innerHTML replacement. Render it immediately on page load before JS fetches plugin data. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 09:57:58 -04:00
andreas	1a19088cfe	udp: resolve host owner from config, default_owner, or os_info on each PLG Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 09:50:42 -04:00
andreas	172f6e950f	plugins: show host owner in glance strip for admin users Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 09:12:02 -04:00
andreas	b3aa7b585f	udp/config: fall back to default_owner when os_info has no owner; log debug - When os_info arrives with no owner field, apply default_owner from server config - Stop applying default_owner unconditionally in get_host_access (now deferred to os_info handling) - os_info plugin logs debug message when injecting owner from client config Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 08:49:42 -04:00
andreas	88a3c09b51	hbc/server: request InfoPlugin refresh when host has no plugin data; update docs - Server sets request_update=1 in ACK when host.plugin_data is empty - hbc: AsyncConnection.request_info_event; handle_ack sets it on request_update - hbc: _info_plugin_refresh_loop clears InfoPlugin caches and resends on demand - hbc_mini: same via _request_info event and _info_refresh_loop - docs/USERS.md: document client-declared owner config key - docs/PLUGIN_DEVELOPMENT.md: document server-initiated InfoPlugin refresh Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 07:37:41 -04:00
andreas	0504402a8a	hbc/hbc_mini: add owner config; include in os_info; server applies to host - owner: optional top-level config key in ~/.hbc.yaml / ~/.hbc.json - Propagated into plugin configs at load time so os_info can include it - os_info PLG data carries owner field when set - udp: sets host.owner from os_info if not already configured server-side - live.html: format event log timestamps as YYYY-MM-DD HH:MM:SS (24-hour) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 07:25:47 -04:00
andreas	ca58c18802	eventlog: store structured dicts; filter by user; clock: fix minute hand step - eventlog() now stores {ts, host, level, service, message} dicts instead of strings - WebSocket sends/broadcasts filter event log messages by the user's managed hosts - live.html renders structured log entries with level-coloured spans - Swiss railway clock minute hand now holds until second hand reaches 12, then steps Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 07:00:17 -04:00
andreas	1ddc4b8132	threshold/alerts: strip _status_code suffix from displayed metric names Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-08 06:19:16 -04:00
andreas	5e1720ed32	notify: use plain URL in Mattermost plugin metrics link Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-07 10:43:18 -04:00
andreas	7ab17e26e2	hbc/hbc_mini: log name and version at startup; ui: bump alert-metric font size Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-07 10:15:03 -04:00
andreas	28f5fa951c	ui: show metric name inline with hostname in alerts and notifications Alerts page: move metric name into the header row alongside hostname. Notifications: include metric name in title (hostname metric) and strip the metric prefix from the body so it contains only value/detail. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-07 06:26:27 -04:00
andreas	ca8ba84e65	fix: silence aiohttp.access log and strip plugin prefix in alerts UI - main: disable aiohttp.access propagation unless --debug is active - alerts.html: strip plugin-name prefix from metric_path display (nagios_runner.check_disk_root_status_code → check_disk_root_status_code) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-06 07:39:55 -04:00
andreas	1e4263b793	fix: threshold and logging improvements - threshold: fix crash when display is None (_format_display now falls back to default format string instead of calling None.format()) - threshold: shorten notification messages by stripping plugin-name prefix from metric_path (cpu_percent instead of cpu_monitor.cpu_percent) - main: demote aiohttp.access log records from INFO to DEBUG - udp: replace debug print with proper logger.info for new host sign-on Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-06 07:06:56 -04:00
andreas	1824f637b4	fix: always show THRESHOLD_DEFAULTS in Settings threshold config Seed threshold_configs["default"] from THRESHOLD_DEFAULTS at the start of _parse_config() so the Settings page displays built-in defaults regardless of whether the server config uses the multi-config format, the legacy thresholds: format, or has no threshold config at all. _parse_multi_config() overwrites the seed with the fully-merged effective defaults when a threshold_configs section is present. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-05 13:02:28 -04:00
andreas	a534c06b26	feat: nagios operator for direct exit-code severity mapping Add ComparisonOperator.NAGIOS ("nagios") that maps Nagios exit codes directly to alert levels (0=OK 1=WARNING 2=CRITICAL 3=UNKNOWN) without requiring numeric warning/critical thresholds. Hysteresis is bypassed for discrete codes. Display template defaults to "{check_name}: {output}". _format_display() handles None threshold_value gracefully. Add nagios_runner.status_code as a built-in default threshold config so nagios checks alert out of the box. Also: fix alerts.html scrolling (override html,body), make hostname a link to /plugins#<hostname>, remove overall_status/overall_status_code/plugin_count from nagios_runner and hbc_mini, replace with computed worst-status in plugins.html via nagiosWorstStatus() helper. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-05 12:26:56 -04:00
andreas	ae447ac4a6	feat: nagios_runner improvements and alerts page fixes - nagios_runner: remove overall_status/overall_status_code/plugin_count fields; each command still reports its own <name>_status and <name>_status_code - threshold: expose {output} and {status} aliases in display templates for nagios_runner generic matches (mapped from <check_name>_output/status) - alerts.html: fix scrolling by overriding html,body height/overflow (style.css sets both); make hostname a link to /plugins/<hostname> Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-05 11:05:45 -04:00
andreas	b1985d0eb2	feat: generic threshold matching for nagios_runner with {check_name} display support _find_threshold() now returns the stripped prefix ("check_name") alongside the ThresholdConfig, enabling a single generic entry (e.g. nagios_runner.status_code) to cover all per-command metrics (check_disk_root_status_code, check_load_status_code, …). The prefix is threaded through to _format_display() as {check_name}, with {metric_name} also available in display templates. purge_stale_alerts() updated to use generic matching so it does not incorrectly drop alerts on generic-matched metrics. README updated with Display Format Templates and Generic Threshold Matching sections. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-05 10:48:17 -04:00
andreas	de778f680f	fix: reduce default hysteresis 10%→2%; show recovery threshold in alerts UI The 10% default hysteresis created an unreasonably wide recovery band: a 95% threshold would only clear once the value dropped below 85.5%, causing alerts to linger long after the metric was well below the trigger level. Change default hysteresis to 2% across all threshold parsers (plugin metrics, partitions, RTT). For a 95% threshold, recovery is now at 93.1% instead of 85.5%. Add AlertState.hysteresis field (set on every check, cleared on OK) and expose recovery_threshold in to_dict() so the Alerts dashboard can display "recovers < 93.1" alongside the trigger threshold, making the hysteresis band visible to the user. Pickle backward-compatible via __setstate__. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-04 14:47:50 -04:00
andreas	c93dbdc0f4	fix: settings thresholds show correct per-config metrics; misc hbc fixes Settings page: pass threshold_checker to http.start so the Threshold Configurations section has data. Use threshold_checker's already-parsed ThresholdConfig objects instead of re-parsing the raw nested YAML. Named (non-default) configs now display only their explicit overrides via threshold_raw_configs, not the full merged set with defaults. hbc/hbc_mini: send boot and shutdown messages on first connection only to avoid duplicate packets when multiple servers are configured. Replace print("Daemonizing...") with logging.info so output goes to syslog in daemon mode. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-04 09:12:39 -04:00
andreas	3a546a1e5c	feat: fetch-based Update/Delete buttons with toast notification on Host Overview Replace href navigation with fetch() so the server response is captured and displayed in a slide-up toast at the bottom of the page. Delete also removes the host card from the DOM on success without a page reload. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-04 08:16:54 -04:00
andreas	3301dbfe34	feat: owner Update/Delete buttons on Host Overview; purge stale alerts on reload Host Overview (plugins.html): show Update and Delete buttons in the host-right zone when the logged-in user is the host owner (or admin / unauthenticated mode). Buttons link to /u?h=<host> and /d?h=<host> with stopPropagation so they don't toggle the accordion; Delete prompts for confirmation first. ThresholdChecker.purge_stale_alerts(): removes alert states whose metric_path has no matching threshold in the current config. Called after startup pickle restore and after every SIGHUP config reload so alerts orphaned by upgrades or config changes do not persist indefinitely. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-04 08:03:46 -04:00
andreas	d00d903e7d	fix: make Alerts page scrollable Override the global style.css body height/overflow that locks all pages to the viewport height (a remnant of the old drawer-menu layout). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-04 13:33:08 +02:00
Andreas Wrede	a99b6b54c7	feat: add alert pie chart to nav bar Show a colour-coded pie chart (red=critical, yellow=warning, green=ok) to the left of the clock in the nav bar. Backed by a new GET /api/0/alert_summary endpoint that counts hosts per alert level for the current user's visible hosts. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-05-03 13:45:15 -04:00
Andreas Wrede	a76d0fc840	feat: generic ping_monitor thresholds; round RTT to nearest ms - threshold.py: add _find_threshold() with suffix fallback so thresholds like ping_monitor.rtt_avg match ping_monitor.8_8_8_8_rtt_avg etc.; each pinged host keeps its own alert state - hbdclass.py: format RTT as integer ms (round()) - live.html: JS RTT display rounded to nearest ms (Math.round) Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-03 06:08:11 -04:00
Andreas Wrede	ae60844a8a	feat: link hostnames in Live Dashboard to Host Overview Hostnames in the live dashboard table are now links to /plugins#hostname, which expands and scrolls to that host's card in the Host Overview page. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 14:37:08 -04:00
Andreas Wrede	49fa310361	feat: add Threshold Configurations section to settings page Reads threshold_configs (or legacy thresholds) from config and renders per-named-config tables showing metric path, operator, warning/critical values, hysteresis, and count. Disabled entries are dimmed. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 14:30:31 -04:00
Andreas Wrede	28e2180f7b	fix: suppress notifications on alert de-escalation (e.g. CRITICAL→WARNING) Only notify on worsening transitions (OK→WARNING, OK→CRITICAL, WARNING→CRITICAL) and recovery (any→OK). De-escalation within alert states no longer sends a duplicate notification since the metric never recovered. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 14:27:18 -04:00
Andreas Wrede	ce0590f015	fix: suppress recover messages for down durations under 4 seconds Transient blips caused by hbc client restarts no longer generate eventlog entries or notifications. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 14:18:58 -04:00
Andreas Wrede	72fc82b91f	feat: add ZFS pool renderer to Host Overview Add renderZfsTables() to plugins.html with health/capacity/frag/dedup table and cumulative I/O table; colour-code health and capacity thresholds; add zfs_monitor to plugin_order and summary/render dispatch. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 13:21:28 -04:00
Andreas Wrede	691f62aa69	feat: host-level watch flag suppresses notifications; filter dashboard/overview by owner/manager; add ZFS monitor plugin - watch: true (default) per host; watch: false suppresses all notifications for that host in udp.py and threshold.py - Live Dashboard and Host Overview now show only hosts where the logged-in user is owner or manager (admins see all); WebSocket broadcasts filtered per-connection by the same rule - Add hbd/client/plugins/zfs_monitor.py: collects per-pool health, capacity, fragmentation, dedup ratio, and cumulative I/O ops/bandwidth via zpool(8) Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 12:42:35 -04:00
Andreas Wrede	cffc9805f9	fix: mask api_password and access_token in settings page; add List to threshold imports Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 11:51:55 -04:00
Andreas Wrede	917d6a401b	feat: composable threshold_config list for per-host threshold layering threshold_config in the hosts section now accepts a list of named configs applied left-to-right on top of the defaults, so focused override profiles can be mixed without duplication. Single-string and legacy host_threshold_mapping forms are unchanged. - Add threshold_raw_configs to store per-config overrides separately - Normalise threshold_config to list on parse (string or list) - get_thresholds_for_host folds the list over the default base - Update README and docs/THRESHOLD_ALERTING.md with examples Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-02 10:35:23 -04:00
Andreas Wrede	c4f09e9ced	version 5.1.8 Release / release (push) Successful in 5s Details - fix: matrix/sms_voipms notifications blocked the event loop on timeout; make send_notification async, dispatch all channel drivers as non-blocking tasks (asyncio.to_thread for sync drivers, asyncio.wait_for for async); update all call sites to fire-and-forget via create_task - feat: add /about page with version, runtime, uptime counter, and repo link - fix: hbc_mini plugin data format now matches full hbc client so Host Overview displays memory, disk, and network metrics correctly Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-05-01 05:33:27 -04:00
Andreas Wrede	64710fd4cd	tweak h1 margins	2026-05-01 04:51:11 -04:00
Andreas Wrede	1f5e7465a3	fix nav bar position	2026-05-01 04:32:04 -04:00
Andreas Wrede	b6dcce4f35	simplify eventlog usage, fix arguments	2026-04-30 15:38:46 -04:00
Andreas Wrede	c5ce41762e	feat: update hbc via hb_install.sh instead of code patching Server now sends a bare UPD command; client runs hb_install.sh to reinstall from the package registry, then restarts. hb_install.sh also copies itself alongside hbc on client installs. Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-04-30 13:55:15 -04:00
Andreas Wrede	ddf7067d13	feat: redesign Plugin Metrics page as Host Overview Replace pill-tab plugin view with an accordion layout that shows key metrics (CPU%, MEM%, top disk%, net delta, nagios status) at a glance in each host card header. Plugin sections expand as structured tables. - Rename page to "Host Overview" (URL /plugins unchanged) - Three-wave parallel data loading: glance plugins on host expand, on-demand fetch for filesystem_info and extras - Per-plugin table renderers with inline percent bars and threshold colour coding - Add escHtml() for XSS-safe rendering of all field values - Remove stale planning docs (REFACTORING.md, hbd/Plan.md) Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>	2026-04-30 08:12:07 -04:00
andreas	990c658e65	Apply grace period to all threshold alerts before logging/notifying Threshold alerts (plugin metrics, RTT) were firing immediately on the first breach. Now every state transition to WARNING/CRITICAL starts a grace-period timer (grace_seconds from the 'grace' config key). The notification is deferred until the next heartbeat after grace_seconds have elapsed. If the metric recovers within the grace window, both the alert and the recovery are suppressed — no spurious pages for transient spikes. Two helper methods added to ThresholdChecker: - _apply_grace: handles the state-change path (defer or suppress) - _check_pending_or_renotify: handles the stable-alert path (fire deferred notification once grace expires, or fall through to reminders) The overdue case is unchanged — on_overdue already fires only after interval+grace seconds of silence, which is equivalent behaviour. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-24 12:00:40 +02:00
andreas	b78d6ac0fe	Fix RECOVER routing: use consistent level name and route via alerted channel threshold.py was emitting level="RECOVERED" for metric recoveries, which failed the is_recover check in send_notification (which only matched "RECOVER"), bypassing _alerted_channels routing and the min_level bypass added in the previous commit. Changed to "RECOVER" so all recovery paths are consistent. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-24 11:29:04 +02:00
andreas	afd5060f59	Fix early reminder notifications and lost recovery notifications - AlertState.update() now resets last_notification when the alert level changes, so a WARNING→CRITICAL escalation restarts the reminder interval rather than inheriting a nearly-expired timer. - _dispatch_to_channel() bypasses min_level for RECOVER, so recovery notifications are delivered even after a server restart when _alerted_channels is empty and the fallback dispatch path is used. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>	2026-04-22 18:11:22 +02:00
Andreas Wrede	5c382d2b8d	One more nit	2026-04-13 09:31:35 -04:00
Andreas Wrede	35bba451f5	Various formating nits	2026-04-13 09:27:51 -04:00
Andreas Wrede	80edfba0c0	fix inconsistencies in page layout, add swiss clock	2026-04-13 08:45:50 -04:00
Andreas Wrede	6bc8de192e	fix non-alerting of overdue hosts	2026-04-12 18:44:36 -04:00
Andreas Wrede	d0c8c186f4	Fix typo	2026-04-12 13:04:17 -04:00
Andreas Wrede	19f7c8312e	Mkae columns sortabel agian, check hbc version, provide modile html pages	2026-04-12 12:53:00 -04:00
Andreas Wrede	24b0e362fb	provide cli function stop, restart and reload for hbd Thought for 1s	2026-04-12 12:06:07 -04:00

1 2 3

131 Commits