Skip to main content

Design: Outbound Todo Notify Hooks

Context​

A channel doorbell (ADR-0013) needs a live process that has loaded Switchboard as a channel. claude -p, cron sweeps, CI jobs and start-on-demand supervisors have no such process. Today they poll, or the operator makes the upstream sender deliver twice, which moves trust configuration outside Switchboard and bypasses routing rules.

ADR-0029 (accepted) chose per-endpoint notify hooks managed over MCP. SPEC-0024 is the requirement set. Nothing outbound exists for todos today:

  • internal/push/ssrf.go is a DNS-rebinding-aware validator with no callers.
  • push_notification_configs (migration 0012) is A2A task-scoped, sits behind SWITCHBOARD_A2A, and is unused.

The consumer this unblocks first is Harness's event triggers (Harness ADR-0021 / SPEC-0014): a [webhook.*] listener that starts a one-shot when Switchboard says work is waiting. Together with Cairn's annotation events, this is the "on-demand one-shot" milestone.

Goals / Non-Goals​

Goals​

  • A headless consumer learns that work is waiting within seconds, with no polling and no second upstream webhook.
  • No instance-wide surface. Every hook has an owner, a scope and a verb that granted it.
  • Reuse what exists: the doorbell hook, the sender gate, the SSRF validator and the secret envelope.
  • Ingest never waits on a tenant's receiver.

Non-Goals​

  • Guaranteed delivery. The queue is the guarantee, and the hook is a hint.
  • Carrying the work. The body is a pointer, and the consumer claims.
  • Coalescing bursts (not in v1; see Open Questions).
  • A test_notify_hook verb. ADR-0029 fixed the verb set at four. A receiver is tested by a real todo, and list_notify_hooks shows the result.
  • Replacing A2A push (SPEC-0019). That path should adopt this worker later, not the other way round.

Decisions​

A table of its own, keyed to the endpoint​

Choice: a new notify_hooks table with endpoint_id as a cascading foreign key. The table has no owner columns of its own.

Rationale: every endpoint-attached resource inherits the endpoint's owner scope. Teams (ADR-0038 / SPEC-0033) keep that rule, and let a team own an endpoint. A hook that carries its own owner_human_id would be a second ownership fact that could disagree with the endpoint's. The cascade means revoking and deleting an endpoint takes its hooks with it.

Alternatives considered:

  • Reuse push_notification_configs: task-scoped, behind a flag that is off by default, configured by the remote caller. Rejected in ADR-0029 (option C).

Validation runs twice, and the dial is pinned​

Choice: push.Validator runs at create and before every attempt. The attempt's http.Transport uses a DialContext that:

  1. resolves the host once through the validator's injected resolver;
  2. validates every returned address, and drops the ones that fail;
  3. dials a surviving address directly;
  4. sets TLS ServerName to the URL host.

CheckRedirect returns http.ErrUseLastResponse.

Rationale: validating and then letting net/http resolve again leaves a time-of-check/time-of-use window that a zero-TTL rebinding record can use. Pinning closes it. The validator stays the single source of truth for "allowed", as ADR-0021 intended.

Alternatives considered:

  • Validate only at create: defeated by rebinding. Rejected by ADR-0029.
  • Egress proxy: the right answer for a large deployment, but an operational dependency the self-hosted default cannot assume. It MAY be layered on later with HTTPS_PROXY plus the same validator.

Private ranges are an operator bound, off by default​

Choice: SWITCHBOARD_NOTIFY_HOOK_ALLOW_CIDRS (comma-separated CIDRs) exempts listed ranges from the private-address rejection. Loopback and link-local (including cloud metadata addresses) are exempted only by an entry lying wholly inside them, such as 127.0.0.1/32 for a Harness listener on the same host; a broad entry like 0.0.0.0/0 never opens them. The metadata services that sit in CGNAT or ULA space (100.100.100.200, fd00:ec2::254) are held to the same bar: only an entry for exactly that address opens one. Switchboard's own listen address and port stay rejected even when listed, compared as address plus port. Today WithOwnListenAddrs in internal/push/ssrf.go records IPs only and leaves loopback binds to the loopback rule, so the story extends it to ports.

Rationale: a self-hoster running Switchboard and Harness on one LAN needs the "server it can reach" path to work without a public hop. On a multi-tenant instance this opens those ranges to every tenant, so the choice belongs to the operator (an instance role that may bound tenant behaviour), and the guide says so loudly.

Standard Webhooks, with dual signatures on rotation​

Choice: webhook-id / webhook-timestamp / webhook-signature: v1,<sig> over id.timestamp.body. The secret is whsec_ plus base64 of 32 random bytes. Rotation keeps the previous secret for 24 hours and signs with both.

Rationale: this is the scheme ADR-0029 chose. Switchboard's own generic receiver already honours Webhook-Id inbound, and off-the-shelf verifier libraries exist in most languages. Dual signing is part of the Standard Webhooks spec and makes rotation zero-downtime.

The secret encoding is an interop trap. The inbound minter, mintWebhookSecret in internal/mcp/webhooks.go, returns whsec_ followed by hex. Hex is a valid base64 alphabet, so a Standard Webhooks verifier decodes a hex secret without complaint, derives a different key, and rejects every delivery. Neither side reports anything useful. Notify hooks therefore get their own minter that uses padded standard base64. A test verifies a real notification with a reference Standard Webhooks implementation, not with Switchboard's own signer.

Cross-repo consequence: Harness ADR-0021 deferred timestamped schemes, so Harness [webhook.*] cannot yet verify this header set. The receiving end is stump.wtf/harness#466 ("standard-webhooks verification — receive Switchboard notify hooks"), under the SPEC-0014 amendment stump.wtf/harness#427 and epic stump.wtf/harness#451. Its tests encode four properties that this spec must honour:

  • the base64 secret;
  • dual signatures during rotation;
  • a webhook-id that is stable across retries;
  • a fresh webhook-timestamp on each attempt, within a 5-minute tolerance.

It also filters on the body's top-level type. The F-X2 webhook path is blocked until Harness verifies these signed hooks. Switchboard does not add a second, body-only signature to paper over the gap: that would drop timestamp replay protection for every receiver, to save one receiver a story.

Fire from the existing doorbell hook, plus the requeue paths​

Choice: store.SetTodoDoorbellHook gains a second subscriber, the hook dispatcher, next to mcp.PublishTodoReady. The two re-queue paths call the same fan-out after their statement returns rows:

  • the lease reaper (ReapExpired, claimed → pending);
  • the retry scheduler (RequeueDueRetries, failed with a due next_retry_at → pending).

Neither re-queue statement applies the sender gate today. They move rows and emit the todo_ready wakeup, and the gate lives only in the doorbell read queries (RingUnclaimed, RingOnAttach, PendingDoorbellTodos). So the re-queue fan-out MUST re-apply the SPEC-0011 predicate to the rows it returns before any hook is considered. That predicate is: the event exists, and verified OR trust_mode = 'token'. The statement becomes a CTE: WITH moved AS (UPDATE … RETURNING …) SELECT moved.*, (e.verified OR e.trust_mode = 'token') AS push_eligible FROM moved LEFT JOIN events e ON e.id = moved.event_id. Only push_eligible rows reach the dispatcher. A test re-queues an unverified todo and asserts that no hook fires.

Rationale: one trigger, one gate. A second notification path with a looser gate is the one attackers would use (ADR-0029 driver "One sender gate"). The requeue fire is what lets a start-on-demand supervisor replace a worker that died holding a lease. It is also the hook that Switchboard ADR-0039 (attempt history, F-X3) builds its relay loop on.

Alternatives considered:

  • Fire on heartbeat re-rings as well: the sweep exists to reach a live session that missed a push. Firing it at a dispatcher would start a new worker every 5, 20 and 60 minutes for a todo that a live worker may be holding off on deliberately. Rejected for v1.

A per-instance bounded queue, with no persistence​

Choice: two buffered channels of 1024 per instance, drained by one worker pool of 8 goroutines. The first holds ready todos waiting to be matched; each carries only the ids, queue, attempt, reason and the sender text already cut to its body limit, never the payload, so a full queue pins kilobytes rather than the todos' multi-megabyte payloads. Matching turns each todo into one delivery job per matching hook on the second channel, so one hook's timeouts and retries never hold back a sibling hook's first attempt. The endpoint, its scope, the hook and its secrets are re-read before each attempt, so a revoke, a scope shrink, a delete and a disable all stop the next attempt. Retries run inside the worker with time.Timer, not by re-enqueueing.

Rationale: the ADR-0013 contract is that a restart loses hints and nothing else. A persistent outbox would turn a hint into a delivery guarantee we have promised no one. It would also add a table that grows during an outage, exactly when it hurts.

Health lives on the hook row​

Choice: last_attempt_at, last_status, last_error, consecutive_failures, enabled, disabled_reason and disabled_at are columns. Each is updated with a single UPDATE … WHERE id = $1 after each delivery (not after each attempt). The disable decision is made in the same statement:

UPDATE notify_hooks
SET consecutive_failures = consecutive_failures + 1,
last_attempt_at = now(), last_status = $2, last_error = $3,
enabled = CASE WHEN consecutive_failures + 1 >= $4 THEN false ELSE enabled END,
disabled_reason = CASE WHEN consecutive_failures + 1 >= $4
THEN 'consecutive_failures' ELSE disabled_reason END,
disabled_at = CASE WHEN consecutive_failures + 1 >= $4 THEN now() ELSE disabled_at END
WHERE id = $1 AND enabled
RETURNING enabled;

Rationale: two instances failing the same hook concurrently cannot both miss the threshold, and the row itself is the audit record.

Presence, and the resolution of ADR-0029's open question​

Choice: respect presence by default. Add ignore_presence per hook. On out → in, send one todos.backlog to each held hook that has matching pending work. That notification is decided by the SPEC-0022 transition claim.

Rationale: out is the agent's or its human's statement that no turns should be spent on this endpoint right now. A shift that says "weekdays only" is a statement about cost and attention, and it applies to a one-shot started by a hook as much as to a live session. The dispatcher case, where something else decides whether to start a worker, is real but is the minority, so it is the opt-out. Without todos.backlog, a held hook would never hear about work that arrived while the endpoint was out, because hooks only fire on transitions. With it, the return is surfaced once, with counts only, exactly as the digest doorbell does for sessions.

This work is gated on SPEC-0022 landing. Until then every endpoint is in and the flag is inert. The presence story is therefore filed separately and blocked, and the rest of the feature ships without it.

The digest question is carried, with a measurement plan​

Choice: no coalescing in v1. Each ready todo is one notification, and each notification is dedupable by webhook-id. A per-hook limit of 120 notifications per minute bounds amplification.

Rationale: coalescing adds a timer, a flush rule and a second body shape to every consumer. Nobody has measured a burst problem yet. What would show one:

  • a rising switchboard_notify_hook_notifications_total{outcome="dropped"} rate;
  • receivers reporting duplicate worker starts for one queue within seconds.

If either appears, the design is a todos.ready batch type behind a per-hook coalesce_ms, defaulting to off.

Architecture​

Schema​

A new migration, taking the next free number when it is written. Other planned specs also claim migrations, so the number is not reserved here.

CREATE TABLE notify_hooks (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
endpoint_id uuid NOT NULL REFERENCES endpoints(id) ON DELETE CASCADE,
url text NOT NULL CHECK (length(url) <= 2048),
secret text NOT NULL, -- internal/cred envelope, always; no key = no hook (fail closed)
prev_secret text, -- dual-sign grace after rotation
prev_secret_expires_at timestamptz,
queues text[] NOT NULL DEFAULT '{}',
ignore_presence boolean NOT NULL DEFAULT false,
enabled boolean NOT NULL DEFAULT true,
disabled_reason text CHECK (disabled_reason IN ('consecutive_failures', 'operator')),
disabled_at timestamptz,
consecutive_failures integer NOT NULL DEFAULT 0,
last_attempt_at timestamptz,
last_status integer,
last_error text,
created_at timestamptz NOT NULL DEFAULT now(),
rotated_at timestamptz
);
CREATE INDEX idx_notify_hooks_endpoint ON notify_hooks (endpoint_id);

Unlike the inbound webhook secret, a hook secret has no plaintext fallback: the store refuses to create or rotate a hook without SWITCHBOARD_SECRET_ENCRYPTION_KEY, and refuses to sign with a stored value that is not envelope ciphertext. The index is not partial, because Postgres does not index a foreign key's referencing column and the list, the ceiling count and the endpoint cascade all filter on endpoint_id alone.

The migration is additive. Rolling it back means dropping the table, which loses only hook registrations.

MCP shapes​

// create_notify_hook
{"url": "https://dispatch.example.com/sb", "queues": ["reviews"], "ignore_presence": false}
// → result
{"hook_id": "8c1e…", "url": "https://dispatch.example.com/sb", "queues": ["reviews"],
"ignore_presence": false, "enabled": true, "signing_secret": "whsec_…"}

// list_notify_hooks → result
{"hooks": [{"hook_id": "8c1e…", "url": "https://dispatch.example.com/sb", "queues": ["reviews"],
"ignore_presence": false, "enabled": true, "disabled_reason": null,
"consecutive_failures": 0, "last_status": 202, "last_error": null,
"last_attempt_at": "…", "created_at": "…", "rotated_at": null}],
"ceiling": {"max": 5, "used": 1}}

// rotate_notify_hook {"hook_id": "8c1e…"} → {"hook_id": "…", "signing_secret": "whsec_…",
// "previous_secret_valid_until": "…", "enabled": true}
// delete_notify_hook {"hook_id": "8c1e…"} → {"deleted": true}

Errors use the stable error shape from SPEC-0006 REQ "Structured Output and Stable Error Shape": invalid_argument, forbidden, not_found and ceiling_exceeded.

Configuration​

VariableDefaultMeaning
SWITCHBOARD_NOTIFY_HOOK_MAX5Per-endpoint hook ceiling (operator bound)
SWITCHBOARD_NOTIFY_HOOK_ALLOW_CIDRSemptyCIDRs exempt from the SSRF private-range rule; loopback and link-local only via entries wholly inside them; every tenant can reach them
SWITCHBOARD_PUSH_ALLOW_HTTPunsetExisting opt-in; also permits http:// hook URLs

Fixed values (constants, not configuration): attempt timeout 5s, 3 attempts, disable after 10 consecutive failed deliveries, 24h rotation grace, 120 notifications per minute per hook, a queue of 1024 per instance, and 8 workers.

Receiver guide (docs)​

A new guide section, "Wake a consumer with a notify hook", MUST cover:

  • verifying the signature, with a 20-line Go and Python verifier and a pointer to the Standard Webhooks libraries;
  • rejecting a webhook-timestamp more than 5 minutes old;
  • deduplicating on webhook-id;
  • answering 2xx fast and doing the work asynchronously;
  • claiming by todo_id, and treating summary as data;
  • ignore_presence;
  • the end-to-end F-X2 recipe with a Harness [webhook.*] trigger.

Risks / Trade-offs​

  • Switchboard now dials tenant-chosen hosts. → The validator runs twice, the dial is pinned, HTTPS is the default, redirects are refused, private ranges are closed by default, and each hook is rate limited. These bound the risk. They do not remove it.
  • A hook can silently stop working. → Per-hook health, auto-disable with a reason, the endpoint card, and the counters. The queue still holds the todo.
  • A consumer treats the hook as the delivery. → The body carries too little to work from, and the docs say to claim.
  • Double work when a session and a hook both wake a consumer. → The lease allows one claim. The second consumer finds nothing and exits, which costs one turn.
  • Harness cannot verify the signature yet. → A cross-repo story, tracked in the epic as a blocker for the F-X2 acceptance test.

Migration Plan​

  1. Ship the migration and the store with the verbs unregistered (dark).
  2. Ship the verbs and the dispatcher together. Existing endpoints gain nothing until their human grants the verbs.
  3. Ship the web UI rows and the receiver guide.
  4. Ship the presence integration after SPEC-0022 is implemented.

Rollback at any step: remove the verbs from registration (a hook with no dispatcher never fires), then drop the table.

Open Questions​

  • Presence (carried from ADR-0029). Resolved (design review 2026-09-22): hooks follow clock-in and clock-out by default, and a per-hook ignore_presence flag is the opt-out (REQ-9). A held hook gets one counts-only todos.backlog on clock-in.
  • Digest / coalescing (carried from ADR-0029). Resolved (design review 2026-09-22): no coalescing in v1. The measurement plan above decides whether it is ever designed.
  • Can the operator disable notify hooks outright with SWITCHBOARD_NOTIFY_HOOK_MAX=0? Resolved (design review 2026-09-22): yes, as proposed. A ceiling of 0 refuses every create and also stops dispatch for existing hooks, and a startup log line reports it (REQ-1).