Skip to main content

Upgrading

This page documents the changes that need action when you move between Switchboard releases. It is written for people running their own Switchboard; if you are on the published image, check which release you are actually running first.

To find your version, run switchboard version, or read it from /healthz. Newer builds also report it over MCP as serverInfo.version.

Upgrading to v0.5.0​

Read this first​

This release narrows what friend endpoints and sibling endpoints may do, and one migration cannot be fully reversed. It also makes routing rules fail closed, with no switch: see Routing rules fail closed below. After you upgrade:

  • a webhook's routes and rules can be managed only from the endpoint that owns the webhook. The rule and route verbs called from any other endpoint, including another endpoint of the same person, answer not_found;
  • every endpoint minted by approving a friend request is narrowed in place to create_for and the drain verbs (list_todos, get_todo, claim, claim_next, complete, fail, release, heartbeat), and its friendship's recorded grant is narrowed with it;
  • a friend request (over A2A, or from the Friends page) that names only verbs a friend can never be granted is refused with a 400 instead of being stored.

Migration 0025_friend_edges_own_authority rewrites endpoints.scope_verbs and friend_edges.granted_verbs in place. Its index change can be undone, but the verbs it removes cannot be put back from the database. Back up first.

The migration's own narrowing list predates the release verb (#507) and does not carry it, so a friend endpoint that holds release loses it on upgrade. A follow-up migration, 0028_friend_release_verb, repairs that on the same upgrade: every approved edge that requested release and no longer holds it gets it back, and its endpoint is re-broadened with it. (A grant is always a subset of the request, so requesting release is the only way an edge could ever have held it.) One limit: an approver who deliberately withheld release while granting the other requested verbs is indistinguishable from a stripped grant and gets it back too — revoke the edge and re-approve with a narrower scope if that matters. The list below is the correct one, and the verification query says what to expect.

What breaks, and who is affected​

You are affected if any of these is true:

  • an agent edits a webhook's rules or routes (list_webhook_rules, set_webhook_rules, add_webhook_rule, update_webhook_rule, move_webhook_rule, remove_webhook_rule, test_webhook_rules, list_webhook_routes, add_webhook_route, remove_webhook_route) from an endpoint other than the one that created the webhook; or
  • you approved a friendship that granted webhook, rule, route or event-history verbs. Those verbs are removed from the friend's endpoint, and a friend agent that relied on them now gets forbidden; or
  • a client sends friend requests that ask only for such verbs.

You are not affected if every webhook is managed from the endpoint that created it and no friendship was granted anything beyond create_for and the drain verbs.

Why this changed​

A friend endpoint is vended on the approver's agent, so every verb on it acted with the approver's authority. A friend granted set_webhook_rules or test_webhook_rules could rewrite the approver's rules or read their payloads. The same check also let any endpoint of a person configure all of that person's webhooks, so one leaked credential reached every webhook they own. Both now follow the endpoint that owns the webhook (SPEC-0033 F3 and F19).

Before you upgrade: back up​

pg_dump --format=custom --file=switchboard-pre-friend-authority.dump "$SWITCHBOARD_DATABASE_URL"

Keep that file somewhere other than the database host until you have confirmed the upgrade works.

Steps​

  1. Manage each webhook from its own endpoint. list_webhooks on an endpoint lists the webhooks it owns. Point any agent that edits a webhook's rules or routes at that endpoint's credential.

  2. Review routes that deliver to friend endpoints. Before this release, a friend endpoint holding add_webhook_route could route your webhook's deliveries to itself. The migration removes the verb but leaves any route it made, and such a route looks exactly like one you added. List them:

    SELECT r.webhook_id, r.target_endpoint_id, r.granted_at, f.from_persona AS friend
    FROM webhook_routes r
    JOIN friend_edges f ON f.endpoint_id = r.target_endpoint_id
    ORDER BY r.granted_at;

    For each row you did not add yourself, call remove_webhook_route from the webhook's own endpoint, and check that webhook's rules with list_webhook_rules.

Verify​

  • No friend endpoint carries anything beyond create_for and the drain verbs. This should return no rows:

    SELECT e.id, e.scope_verbs
    FROM endpoints e
    JOIN friend_edges f ON f.endpoint_id = e.id
    WHERE f.state = 'approved'
    AND NOT e.scope_verbs <@ ARRAY['create_for','list_todos','get_todo','claim','claim_next','complete','fail','release','heartbeat'];
  • From a webhook's own endpoint, list_webhook_rules returns its rules; from any other endpoint it returns not_found.

If you need to go back​

Restore the dump with pg_restore --clean --if-exists --dbname "$SWITCHBOARD_DATABASE_URL" switchboard-pre-friend-authority.dump, then run the previous image. Anything written after the upgrade is lost.

Upgrading to v0.4.0​

Replay targets are owned by the endpoint​

What breaks. The instance settings replay_default_target and replay_allowed_targets are removed. Migration 0026_owned_replay_targets deletes both rows from settings; after the upgrade nothing reads them and nothing warns about them at startup. Their targets were exempt from the SSRF checks for every tenant's replay, which is why they are gone rather than kept for compatibility.

replay_webhook_event now:

  • replays to the target_url the call names, or else to the calling endpoint's first owned replay target;
  • fails with the new error code replay_target_required when it has neither;
  • sends every target, owned or not, through the shared SSRF guard, both when the call is made and again when it connects: the target must be https and every address its host resolves to must be public. Loopback, private (RFC 1918 and IPv6 unique-local), link-local and cloud-metadata, and shared (100.64.0.0/10) addresses are refused, and so is plain http.

Who is affected. Anyone who set either setting, or who replays to a consumer on localhost, a private network, or plain http. An agent that called replay_webhook_event without target_url now gets replay_target_required until its endpoint owns a target.

What to do.

  1. Back up first. The migration deletes the two rows and cannot put them back:

    pg_dump --format=custom --file=switchboard-pre-replay-targets.dump "$SWITCHBOARD_DATABASE_URL"
  2. For each endpoint that should replay by default, vend a replacement that owns its targets. Replay targets are part of an endpoint's scope, which is fixed at vend time:

    curl -sS -X POST "$SWITCHBOARD_URL/api/v1/endpoints" \
    -H "Authorization: Bearer $OPERATOR_TOKEN" -H 'Content-Type: application/json' \
    -d '{"name":"my-agent","replay_targets":["https://consumer.example.com/hook"]}'

    A target the guard refuses fails the vend with 400 and mints nothing. Otherwise, pass target_url on each replay call.

  3. Move any replay consumer that lived on localhost or a private network to a public https URL, or stop replaying to it. There is no allowlist for internal hosts.

Verify.

  • GET /api/v1/endpoints lists replay_targets for each of your endpoints.

  • replay_webhook_event with no target_url on an endpoint without targets returns replay_target_required.

  • Nothing is left in settings. This should print 0:

    psql "$SWITCHBOARD_DATABASE_URL" -Atc \
    "SELECT count(*) FROM settings WHERE key IN ('replay_default_target','replay_allowed_targets')"

Routing rules fail closed​

:::warning Behaviour change, no switch A routing rule that errors, times out, runs out of memory or budget, or no longer compiles used to count as "no match": the delivery fell through to later rules or the default. It now stops evaluation, and the delivery is recorded as faulted and held in the owner's quarantine instead of reaching any queue. A webhook with rules on an instance whose rule sandbox cannot run answers 503 instead of routing by default. There is no setting to turn this off. :::

Who is affected: any webhook whose rules fault on some of its traffic today. Before this release those faults were recorded on the delivery's routing trace while it routed anyway. Before upgrading, find them with this read-only query against your database (the last 7 days; widen the interval if your traffic is sparse):

SELECT webhook_id, count(*) AS faulted_deliveries, max(received_at) AS last_seen
FROM events
WHERE received_at > now() - interval '7 days'
AND jsonb_array_length(COALESCE(routing_trace->'faults', '[]'::jsonb)) > 0
GROUP BY webhook_id ORDER BY faulted_deliveries DESC;

For each webhook it lists, dry-run a recent delivery with test_webhook_rules {"event_id": …}, fix the rule (usually a // default or a type guard), and save it. After the upgrade, list_webhook_events {"disposition": "faulted"} lists deliveries that faulted, and a rule save that would fault on the webhook's 50 most recent deliveries is refused.

The upgrade also adds events.disposition (migration 0022), a metadata-only column add that backfills drops from their stored traces.

Upgrading past v0.3.0 (unreleased): quarantine is a reserved queue name​

:::warning The migration refuses to run if the name is in use Deliveries that are held for review (an untrusted actor, a faulting rule, or a rule's {"quarantine": true}) now wait on a reserved queue called quarantine on the webhook owner's endpoint, where no agent can list, claim or be rung for them. Migration 0028_quarantine aborts, naming the counts, if any todo, endpoint scope or webhook-queue ceiling, or webhook target already uses a queue with that name, and Switchboard will not start until it applies. :::

Who is affected: only a deployment that already routes work to a queue literally named quarantine. Check before upgrading with this read-only query:

SELECT (SELECT count(*) FROM todos WHERE queue = 'quarantine') AS todos,
(SELECT count(*) FROM endpoints
WHERE 'quarantine' = ANY(scope_queues) OR 'quarantine' = ANY(webhook_queues)) AS endpoints,
(SELECT count(*) FROM endpoint_webhooks WHERE target_queue = 'quarantine') AS webhooks;

If any count is non-zero, move that work to a differently named queue first: drain or re-route the todos, re-vend the endpoints without quarantine in their scope or ceiling, and point the webhooks at the new queue. From this release on, quarantine is refused as a webhook target, a vend scope or ceiling, a friend request queue and a rule action's queue.

Upgrading to v0.3.0​

Read this first​

v0.3.0 is a breaking change for any deployment that receives webhooks. The instance-wide receivers configured through environment variables have been removed. After you upgrade:

  • the environment variables listed below are ignored, with no warning and no message at startup;
  • deliveries to /webhooks/github, /webhooks/gitea, /webhooks/stripe, /webhooks/slack and /webhooks/generic/* will fail, because those routes no longer exist;
  • the migration that drops the old provider rows cannot be reversed.

Back up your database before you upgrade. The backup is the only way back.

What breaks, and who is affected​

You are affected if any of these is true:

  • you set SWITCHBOARD_GITHUB_SECRET, SWITCHBOARD_GITEA_SECRET, SWITCHBOARD_STRIPE_SECRET or SWITCHBOARD_SLACK_SECRET; or
  • you set SWITCHBOARD_LEGACY_RECEIVER_ENDPOINT_ID; or
  • a sender delivers to POST /webhooks/<provider> rather than to a per-webhook URL that contains a token.

You are not affected if every sender already delivers to a webhook you created with create_webhook or the vend wizard, and you never set those variables. In that case the upgrade needs no changes from you; skip to "Verify".

Why this changed​

The old receivers were instance-wide: one secret and one URL shared by every sender, and every delivery landed in a registry the operator owned. That meant Switchboard could not name the endpoint that owned the todos a delivery minted, so the todos belonged to nobody, the operator board could not attribute them, and the provider registry held secrets with nothing able to rotate or remove them.

A webhook now belongs to an endpoint, carries its own signing secret, and mints todos owned by that endpoint's owner. The trade is that each sender gets its own URL and secret, and nothing is shared. This is the same reason the "receiver not configured" 503 and the empty providers page existed; both are gone with the old path.

Before you upgrade: back up​

Migration 0021_drop_adapters deletes the old provider rows. It is not reversible, and no downgrade can restore them. Take a dump first:

pg_dump --format=custom --file=switchboard-pre-v0.3.0.dump "$SWITCHBOARD_DATABASE_URL"

Keep that file somewhere other than the database host until you have confirmed the upgrade works.

The variables that are now ignored​

Each is ignored silently. The replacement is the same in every case: a webhook created through create_webhook (or the vend wizard), which returns the ingest URL and signing secret that sender should use instead.

Retired variableReplacement
SWITCHBOARD_GITHUB_SECRETA webhook created with create_webhook for the GitHub sender; paste its ingest_url and signing_secret into the GitHub webhook settings.
SWITCHBOARD_GITEA_SECRETSame, for the Gitea sender.
SWITCHBOARD_STRIPE_SECRETSame, for the Stripe sender.
SWITCHBOARD_SLACK_SECRETSame, for the Slack sender.
SWITCHBOARD_LEGACY_RECEIVER_ENDPOINT_IDNot replaced. It named the endpoint that claimed every legacy delivery; ownership now comes from the webhook itself, so create_webhook takes as_endpoint instead.

Remove the variables from your environment once the senders have moved. Leaving them set is harmless, but it hides the fact that the migration is unfinished.

Move each sender to its own webhook​

Do this per sender. Work through them one at a time and confirm each before starting the next, so a failure is attributable.

Before — one shared secret, one shared URL for every sender:

SWITCHBOARD_GITHUB_SECRET=<one secret, shared>
# GitHub delivers to: https://switchboard.example.com/webhooks/github

After — one webhook per sender, each with its own token in the URL and its own secret:

# No SWITCHBOARD_*_SECRET variables. The secret lives with the webhook, not the process.

For each sender:

  1. Have the agent or operator that owns the target endpoint create the webhook. The ingest URL embeds a per-webhook token, so it differs for every sender:

    create_webhook(name="github-push", as_endpoint=<your endpoint>)
    → ingest_url: https://switchboard.example.com/webhooks/w/<token>
    → signing_secret: <secret, shown once>
  2. Paste both values into the sender's webhook configuration, replacing the old /webhooks/<provider> URL and the shared secret. Set the same content type the sender used before (GitHub and Gitea send application/json).

  3. Send one real delivery from that sender, and confirm it arrives as a todo on the queue you expect.

Verify​

  • Send a test delivery from every sender that was migrated, and confirm each one appears as a todo (list_todos, or the operator board).

  • Confirm nothing still reads the retired variables. This should print nothing:

    env | cut -d= -f1 | grep -E '^SWITCHBOARD_((GITHUB|GITEA|STRIPE|SLACK)_SECRET|LEGACY_RECEIVER_ENDPOINT_ID)$'
  • Confirm the build reports the new version (switchboard version, or /healthz), so you know the upgrade actually took rather than failing back to the old binary.

If you need to go back​

Restore the dump:

pg_restore --clean --if-exists --dbname "$SWITCHBOARD_DATABASE_URL" switchboard-pre-v0.3.0.dump

Then run the previous image. Anything ingested after the upgrade is lost, because the rows that referenced the dropped provider rows cannot be reconstructed. This is why the backup comes first.

About the published image​

ghcr.io/stump-wtf/switchboard:latest moves only when a v* tag is pushed. Before v0.3.0 was tagged it resolved to the v0.2.0 build, so a deployment pulling :latest could have been running a release two weeks behind main without saying so. Tag latest by its digest rather than trusting the name, and pin a version tag (ghcr.io/stump-wtf/switchboard:0.3.0) if you want an upgrade to be something you choose rather than something that happens.

Upgrading to v0.2.0​

v0.2.0 is the first release you can run from a published image. There is no upgrade path from anything earlier, because there was no earlier release; fresh deployments follow Self-hosting.