Skip to content

dist: skip stale cert when rebuilding scheduler HTTP client - #2707

Open
timn-nexthop wants to merge 4 commits into
mozilla:mainfrom
nexthop-ai:upstream/cert-fix
Open

timn-nexthop wants to merge 4 commits into
mozilla:mainfrom
nexthop-ai:upstream/cert-fix

Conversation

@timn-nexthop

@timn-nexthop timn-nexthop commented May 12, 2026 •

Copy link
Copy Markdown

When a build server registers with a new certificate, the scheduler rebuilds its outbound reqwest client with the new cert plus every other cert it knows about. The loop over certs.values() ran before certs.insert(...) overwrote the map entry, so it still contained the stale cert for the same server_id — meaning both the old and new self-signed certs for that server were installed as trust anchors in the rebuilt client.

Each build server's cert is self-signed, so old and new share a Subject DN. TLS validators index trust anchors by Subject; with two anchors having the same name, path building can pick the stale one, fail signature verification against its public key, and reject the handshake. The result is that cert rotation on a build server deterministically (depending on TLS library) breaks the scheduler's ability to talk to it until the scheduler restarts.

Skip the entry matching server_id when iterating existing certs so only the up-to-date cert for that server ends up as a trust anchor.

We've been running with this change since mid December and it's fixed this problem for us.

When a build server registers with a new certificate, the scheduler
rebuilds its outbound reqwest client with the new cert plus every
other cert it knows about. The loop over `certs.values()` ran before
`certs.insert(...)` overwrote the map entry, so it still contained
the stale cert for the same `server_id` — meaning both the old and
new self-signed certs for that server were installed as trust anchors
in the rebuilt client.

Each build server's cert is self-signed, so old and new share a
Subject DN. TLS validators index trust anchors by Subject; with two
anchors having the same name, path building can pick the stale one,
fail signature verification against its public key, and reject the
handshake. The result is that cert rotation on a build server
deterministically breaks the scheduler's ability to talk to it until
the scheduler restarts.

Skip the entry matching `server_id` when iterating existing certs so
only the up-to-date cert for that server ends up as a trust anchor.
Comment thread src/dist/http.rs Outdated
client_builder = client_builder.add_root_certificate(
reqwest::Certificate::from_pem(cert_pem).expect("previously valid cert"),
reqwest::Certificate::from_pem(existing_cert_pem)
.expect("previously valid cert"),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

since you're touching it, could you please drop the expect() and propagate the error?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done.

Comment thread src/dist/http.rs Outdated
);
for (_, cert_pem) in certs.values() {
// Add all OTHER existing certificates (skip the one we're updating)
for (sid, (_, existing_cert_pem)) in certs.iter() {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

could be simpler: insert into certs first, then loop over all values, no skip needed?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if it's simpler, but done.

Comment thread src/dist/http.rs Outdated
@@ -742,14 +742,19 @@ mod server {
server_id.addr()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please add a test for the rotation case; maybe_update_certs could be moved out of the closure to make it testable.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Also done.

timn-nexthop and others added 3 commits September 29, 2026 16:06
When rebuilding the scheduler's HTTP client, the certificates already
stored for other servers were parsed with `expect()`, so a bad stored
cert would panic the request handler. Return the error instead so only
that heartbeat fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Insert the server's new certificate into the map first, then build the
client from the map's values. The new cert replaces the stale entry for
that server, so it no longer needs to be added separately and the old
one skipped.

If building the client fails, restore the previous map entry. Otherwise
the map would hold a cert the client doesn't trust, and the next
heartbeat with the same digest would skip the rebuild.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move `maybe_update_certs` out of `Scheduler::start` so it can be tested
directly, and add a test that rotates a build server's certificate and
checks which certs the scheduler's client trusts afterwards, including
when an update fails.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@timn-nexthop

Copy link
Copy Markdown
Author

Note that I couldn't reproduce this bug with the test on my server today. I think it's still running using openssl on the box, so maybe something changed there.

I still think this change does make the code more correct, even if (today) it's not causing trouble for me.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants