Design: Outbound Todo Notify Hooks
Context
A channel doorbell (ADR-0013) needs a live
process that has loaded Switchboard as a channel. claude -p, cron sweeps, CI jobs and
start-on-demand supervisors have no such process. Today they poll, or the operator makes the
upstream sender deliver twice, which moves trust configuration outside Switchboard and bypasses
routing rules.
ADR-0029 (accepted) chose per-endpoint notify hooks managed over MCP. SPEC-0024 is the requirement set. Nothing outbound exists for todos today:
internal/push/ssrf.gois a DNS-rebinding-aware validator with no callers.push_notification_configs(migration 0012) is A2A task-scoped, sits behindSWITCHBOARD_A2A, and is unused.
The consumer this unblocks first is Harness's event triggers (Harness ADR-0021 / SPEC-0014):
a [webhook.*] listener that starts a one-shot when Switchboard says work is waiting. Together
with Cairn's annotation events, this is the "on-demand one-shot" milestone.
Goals / Non-Goals
Goals
- A headless consumer learns that work is waiting within seconds, with no polling and no second upstream webhook.
- No instance-wide surface. Every hook has an owner, a scope and a verb that granted it.
- Reuse what exists: the doorbell hook, the sender gate, the SSRF validator and the secret envelope.
- Ingest never waits on a tenant's receiver.
Non-Goals
- Guaranteed delivery. The queue is the guarantee, and the hook is a hint.
- Carrying the work. The body is a pointer, and the consumer claims.
- Coalescing bursts (not in v1; see Open Questions).
- A
test_notify_hookverb. ADR-0029 fixed the verb set at four. A receiver is tested by a real todo, andlist_notify_hooksshows the result. - Replacing A2A push (SPEC-0019). That path should adopt this worker later, not the other way round.
Decisions
A table of its own, keyed to the endpoint
Choice: a new notify_hooks table with endpoint_id as a cascading foreign key. The table
has no owner columns of its own.
Rationale: every endpoint-attached resource inherits the endpoint's owner scope. Teams
(ADR-0038 / SPEC-0033) keep that rule, and let a team own an
endpoint. A hook that carries its own owner_human_id would be a second ownership fact that could
disagree with the endpoint's. The cascade means revoking and deleting an endpoint takes its hooks
with it.
Alternatives considered:
- Reuse
push_notification_configs: task-scoped, behind a flag that is off by default, configured by the remote caller. Rejected in ADR-0029 (option C).
Validation runs twice, and the dial is pinned
Choice: push.Validator runs at create and before every attempt. The attempt's
http.Transport uses a DialContext that:
- resolves the host once through the validator's injected resolver;
- validates every returned address, and drops the ones that fail;
- dials a surviving address directly;
- sets TLS
ServerNameto the URL host.
CheckRedirect returns http.ErrUseLastResponse.
Rationale: validating and then letting net/http resolve again leaves a
time-of-check/time-of-use window that a zero-TTL rebinding record can use. Pinning closes it. The
validator stays the single source of truth for "allowed", as ADR-0021 intended.
Alternatives considered:
- Validate only at create: defeated by rebinding. Rejected by ADR-0029.
- Egress proxy: the right answer for a large deployment, but an operational dependency the
self-hosted default cannot assume. It MAY be layered on later with
HTTPS_PROXYplus the same validator.
Private ranges are an operator bound, off by default
Choice: SWITCHBOARD_NOTIFY_HOOK_ALLOW_CIDRS (comma-separated CIDRs) exempts listed ranges
from the private-address rejection. Loopback and link-local (including cloud metadata addresses) are
exempted only by an entry lying wholly inside them, such as 127.0.0.1/32 for a Harness listener on
the same host; a broad entry like 0.0.0.0/0 never opens them. The metadata services that sit in
CGNAT or ULA space (100.100.100.200, fd00:ec2::254) are held to the same bar: only an entry for
exactly that address opens one. Switchboard's own listen address and
port stay rejected even when listed, compared as address plus port. Today WithOwnListenAddrs in
internal/push/ssrf.go records IPs only and leaves loopback binds to the loopback rule, so the story
extends it to ports.
Rationale: a self-hoster running Switchboard and Harness on one LAN needs the "server it can reach" path to work without a public hop. On a multi-tenant instance this opens those ranges to every tenant, so the choice belongs to the operator (an instance role that may bound tenant behaviour), and the guide says so loudly.
Standard Webhooks, with dual signatures on rotation
Choice: webhook-id / webhook-timestamp / webhook-signature: v1,<sig> over
id.timestamp.body. The secret is whsec_ plus base64 of 32 random bytes. Rotation keeps the
previous secret for 24 hours and signs with both.
Rationale: this is the scheme ADR-0029 chose. Switchboard's own generic receiver already
honours Webhook-Id inbound, and off-the-shelf verifier libraries exist in most languages. Dual
signing is part of the Standard Webhooks spec and makes rotation zero-downtime.
The secret encoding is an interop trap. The inbound minter, mintWebhookSecret in
internal/mcp/webhooks.go, returns whsec_ followed by hex. Hex is a valid base64 alphabet, so a
Standard Webhooks verifier decodes a hex secret without complaint, derives a different key, and
rejects every delivery. Neither side reports anything useful. Notify hooks therefore get their own
minter that uses padded standard base64. A test verifies a real notification with a reference
Standard Webhooks implementation, not with Switchboard's own signer.
Cross-repo consequence: Harness ADR-0021 deferred timestamped schemes, so Harness [webhook.*]
cannot yet verify this header set. The receiving end is
stump.wtf/harness#466 ("standard-webhooks verification — receive Switchboard notify hooks"),
under the SPEC-0014 amendment stump.wtf/harness#427 and epic stump.wtf/harness#451. Its tests
encode four properties that this spec must honour:
- the base64 secret;
- dual signatures during rotation;
- a
webhook-idthat is stable across retries; - a fresh
webhook-timestampon each attempt, within a 5-minute tolerance.
It also filters on the body's top-level type. The F-X2 webhook path is blocked until Harness verifies these signed hooks. Switchboard does not add a
second, body-only signature to paper over the gap: that would drop timestamp replay protection for
every receiver, to save one receiver a story.
Fire from the existing doorbell hook, plus the requeue paths
Choice: store.SetTodoDoorbellHook gains a second subscriber, the hook dispatcher, next to
mcp.PublishTodoReady. The two re-queue paths call the same fan-out after their statement
returns rows:
- the lease reaper (
ReapExpired,claimed→pending); - the retry scheduler (
RequeueDueRetries,failedwith a duenext_retry_at→pending).
Neither re-queue statement applies the sender gate today. They move rows and emit the
todo_ready wakeup, and the gate lives only in the doorbell read queries (RingUnclaimed,
RingOnAttach, PendingDoorbellTodos). So the re-queue fan-out MUST re-apply the SPEC-0011
predicate to the rows it returns before any hook is considered. That predicate is: the event
exists, and verified OR trust_mode = 'token'. The statement becomes a CTE:
WITH moved AS (UPDATE … RETURNING …) SELECT moved.*, (e.verified OR e.trust_mode = 'token') AS push_eligible FROM moved LEFT JOIN events e ON e.id = moved.event_id. Only push_eligible rows
reach the dispatcher. A test re-queues an unverified todo and asserts that no hook fires.
Rationale: one trigger, one gate. A second notification path with a looser gate is the one attackers would use (ADR-0029 driver "One sender gate"). The requeue fire is what lets a start-on-demand supervisor replace a worker that died holding a lease. It is also the hook that Switchboard ADR-0039 (attempt history, F-X3) builds its relay loop on.
Alternatives considered:
- Fire on heartbeat re-rings as well: the sweep exists to reach a live session that missed a push. Firing it at a dispatcher would start a new worker every 5, 20 and 60 minutes for a todo that a live worker may be holding off on deliberately. Rejected for v1.
A per-instance bounded queue, with no persistence
Choice: two buffered channels of 1024 per instance, drained by one worker pool of 8
goroutines. The first holds ready todos waiting to be matched; each carries only the ids, queue,
attempt, reason and the sender text already cut to its body limit, never the payload, so a full
queue pins kilobytes rather than the todos' multi-megabyte payloads. Matching turns each todo into
one delivery job per matching hook on the second channel, so one hook's timeouts and retries never
hold back a sibling hook's first attempt. The endpoint, its scope, the hook and its secrets are
re-read before each attempt, so a revoke, a scope shrink, a delete and a disable all stop the next
attempt. Retries run inside the worker with time.Timer, not by re-enqueueing.
Rationale: the ADR-0013 contract is that a restart loses hints and nothing else. A persistent outbox would turn a hint into a delivery guarantee we have promised no one. It would also add a table that grows during an outage, exactly when it hurts.
Health lives on the hook row
Choice: last_attempt_at, last_status, last_error, consecutive_failures, enabled,
disabled_reason and disabled_at are columns. Each is updated with a single UPDATE … WHERE id = $1 after each delivery (not after each attempt). The disable decision is made in the same
statement:
UPDATE notify_hooks
SET consecutive_failures = consecutive_failures + 1,
last_attempt_at = now(), last_status = $2, last_error = $3,
enabled = CASE WHEN consecutive_failures + 1 >= $4 THEN false ELSE enabled END,
disabled_reason = CASE WHEN consecutive_failures + 1 >= $4
THEN 'consecutive_failures' ELSE disabled_reason END,
disabled_at = CASE WHEN consecutive_failures + 1 >= $4 THEN now() ELSE disabled_at END
WHERE id = $1 AND enabled
RETURNING enabled;
Rationale: two instances failing the same hook concurrently cannot both miss the threshold, and the row itself is the audit record.
Presence, and the resolution of ADR-0029's open question
Choice: respect presence by default. Add ignore_presence per hook. On out → in, send one
todos.backlog to each held hook that has matching pending work. That notification is decided by
the SPEC-0022 transition claim.
Rationale: out is the agent's or its human's statement that no turns should be spent on this
endpoint right now. A shift that says "weekdays only" is a statement about cost and attention, and
it applies to a one-shot started by a hook as much as to a live session. The dispatcher case,
where something else decides whether to start a worker, is real but is the minority, so it is the
opt-out. Without todos.backlog, a held hook would never hear about work that arrived while the
endpoint was out, because hooks only fire on transitions. With it, the return is surfaced once,
with counts only, exactly as the digest doorbell does for sessions.
This work is gated on SPEC-0022 landing. Until then every endpoint is in and the flag is inert.
The presence story is therefore filed separately and blocked, and the rest of the feature ships
without it.
The digest question is carried, with a measurement plan
Choice: no coalescing in v1. Each ready todo is one notification, and each notification is
dedupable by webhook-id. A per-hook limit of 120 notifications per minute bounds amplification.
Rationale: coalescing adds a timer, a flush rule and a second body shape to every consumer. Nobody has measured a burst problem yet. What would show one:
- a rising
switchboard_notify_hook_notifications_total{outcome="dropped"}rate; - receivers reporting duplicate worker starts for one queue within seconds.
If either appears, the design is a todos.ready batch type behind a per-hook coalesce_ms,
defaulting to off.
Architecture
Schema
A new migration, taking the next free number when it is written. Other planned specs also claim migrations, so the number is not reserved here.
CREATE TABLE notify_hooks (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
endpoint_id uuid NOT NULL REFERENCES endpoints(id) ON DELETE CASCADE,
url text NOT NULL CHECK (length(url) <= 2048),
secret text NOT NULL, -- internal/cred envelope, always; no key = no hook (fail closed)
prev_secret text, -- dual-sign grace after rotation
prev_secret_expires_at timestamptz,
queues text[] NOT NULL DEFAULT '{}',
ignore_presence boolean NOT NULL DEFAULT false,
enabled boolean NOT NULL DEFAULT true,
disabled_reason text CHECK (disabled_reason IN ('consecutive_failures', 'operator')),
disabled_at timestamptz,
consecutive_failures integer NOT NULL DEFAULT 0,
last_attempt_at timestamptz,
last_status integer,
last_error text,
created_at timestamptz NOT NULL DEFAULT now(),
rotated_at timestamptz
);
CREATE INDEX idx_notify_hooks_endpoint ON notify_hooks (endpoint_id);
Unlike the inbound webhook secret, a hook secret has no plaintext fallback: the store refuses to
create or rotate a hook without SWITCHBOARD_SECRET_ENCRYPTION_KEY, and refuses to sign with a
stored value that is not envelope ciphertext. The index is not partial, because Postgres does not
index a foreign key's referencing column and the list, the ceiling count and the endpoint cascade
all filter on endpoint_id alone.
The migration is additive. Rolling it back means dropping the table, which loses only hook registrations.
MCP shapes
// create_notify_hook
{"url": "https://dispatch.example.com/sb", "queues": ["reviews"], "ignore_presence": false}
// → result
{"hook_id": "8c1e…", "url": "https://dispatch.example.com/sb", "queues": ["reviews"],
"ignore_presence": false, "enabled": true, "signing_secret": "whsec_…"}
// list_notify_hooks → result
{"hooks": [{"hook_id": "8c1e…", "url": "https://dispatch.example.com/sb", "queues": ["reviews"],
"ignore_presence": false, "enabled": true, "disabled_reason": null,
"consecutive_failures": 0, "last_status": 202, "last_error": null,
"last_attempt_at": "…", "created_at": "…", "rotated_at": null}],
"ceiling": {"max": 5, "used": 1}}
// rotate_notify_hook {"hook_id": "8c1e…"} → {"hook_id": "…", "signing_secret": "whsec_…",
// "previous_secret_valid_until": "…", "enabled": true}
// delete_notify_hook {"hook_id": "8c1e…"} → {"deleted": true}
Errors use the stable error shape from SPEC-0006 REQ "Structured Output and Stable Error Shape":
invalid_argument, forbidden, not_found and ceiling_exceeded.
Configuration
| Variable | Default | Meaning |
|---|---|---|
SWITCHBOARD_NOTIFY_HOOK_MAX | 5 | Per-endpoint hook ceiling (operator bound) |
SWITCHBOARD_NOTIFY_HOOK_ALLOW_CIDRS | empty | CIDRs exempt from the SSRF private-range rule; loopback and link-local only via entries wholly inside them; every tenant can reach them |
SWITCHBOARD_PUSH_ALLOW_HTTP | unset | Existing opt-in; also permits http:// hook URLs |
Fixed values (constants, not configuration): attempt timeout 5s, 3 attempts, disable after 10 consecutive failed deliveries, 24h rotation grace, 120 notifications per minute per hook, a queue of 1024 per instance, and 8 workers.
Receiver guide (docs)
A new guide section, "Wake a consumer with a notify hook", MUST cover:
- verifying the signature, with a 20-line Go and Python verifier and a pointer to the Standard Webhooks libraries;
- rejecting a
webhook-timestampmore than 5 minutes old; - deduplicating on
webhook-id; - answering
2xxfast and doing the work asynchronously; - claiming by
todo_id, and treatingsummaryas data; ignore_presence;- the end-to-end F-X2 recipe with a Harness
[webhook.*]trigger.
Risks / Trade-offs
- Switchboard now dials tenant-chosen hosts. → The validator runs twice, the dial is pinned, HTTPS is the default, redirects are refused, private ranges are closed by default, and each hook is rate limited. These bound the risk. They do not remove it.
- A hook can silently stop working. → Per-hook health, auto-disable with a reason, the endpoint card, and the counters. The queue still holds the todo.
- A consumer treats the hook as the delivery. → The body carries too little to work from, and the docs say to claim.
- Double work when a session and a hook both wake a consumer. → The lease allows one claim. The second consumer finds nothing and exits, which costs one turn.
- Harness cannot verify the signature yet. → A cross-repo story, tracked in the epic as a blocker for the F-X2 acceptance test.
Migration Plan
- Ship the migration and the store with the verbs unregistered (dark).
- Ship the verbs and the dispatcher together. Existing endpoints gain nothing until their human grants the verbs.
- Ship the web UI rows and the receiver guide.
- Ship the presence integration after SPEC-0022 is implemented.
Rollback at any step: remove the verbs from registration (a hook with no dispatcher never fires), then drop the table.
Open Questions
- Presence (carried from ADR-0029). Resolved (design review 2026-09-22): hooks follow clock-in and clock-out by default, and a
per-hook
ignore_presenceflag is the opt-out (REQ-9). A held hook gets one counts-onlytodos.backlogon clock-in. - Digest / coalescing (carried from ADR-0029). Resolved (design review 2026-09-22): no coalescing in v1. The measurement plan above decides whether it is ever designed.
- Can the operator disable notify hooks outright with
SWITCHBOARD_NOTIFY_HOOK_MAX=0? Resolved (design review 2026-09-22): yes, as proposed. A ceiling of 0 refuses every create and also stops dispatch for existing hooks, and a startup log line reports it (REQ-1).