Repository navigation
Document B35 and D16 runner failure modes - #9053
Conversation
Co-authored-by: lpcox <15877973+lpcox@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
B35’s status is outdated and the required portable-agent updates are missing.
Review effort: Balanced
Findings: 1
Open (2)
What changed in this PR
Adds runner-doctor guidance for B35 cache-memory permission failures and D16 VMM confinement false positives.
Changes:
- Adds catalog entries and diagnostic mappings.
- Adds regression assertions.
- Refreshes the compiled workflow lock.
| File | Description |
|---|---|
scripts/ci/self-hosted-runner-doctor-workflow.test.ts |
Tests new guidance. |
.github/workflows/shared/self-hosted-failure-modes.md |
Adds B35 and D16. |
.github/workflows/self-hosted-runner-doctor.md |
Adds diagnostic mappings. |
.github/workflows/self-hosted-runner-doctor.lock.yml |
Refreshes generated metadata. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| it('documents B35 and D16 in the shared catalog and diagnostic playbook', () => { | ||
| const source = fs.readFileSync(sourcePath, 'utf-8'); | ||
| const shared = fs.readFileSync(sharedPath, 'utf-8'); |
| | B32 | A repeated/persistent-runner workflow intermittently blocks allowlisted `api.github.com`/`github.com` traffic with `403` or DNS `SERVFAIL`, recurring across otherwise-healthy runs | Squid's default `negative_dns_ttl` is 1 minute, so one transient upstream `SERVFAIL` is negatively cached and replayed for up to 60 seconds even after DNS recovers | **Fixed in AWF (PR github/gh-aw-firewall#8171, merged 2026-09-05):** `generateDnsSection()` emits `negative_dns_ttl 1 seconds`, `dns_retransmit_interval 1 seconds`, and `dns_timeout 10 seconds`. Upgrade AWF to include github/gh-aw-firewall#8171. | Inspect generated `squid.conf` for `negative_dns_ttl 1 seconds`, `dns_retransmit_interval 1 seconds`, `dns_timeout 10 seconds`, and `dns_nameservers`; correlate Squid `TCP_DENIED`/SERVFAIL bursts with concurrent startup or resolver load on unpatched AWF | github/gh-aw-firewall#8168, github/gh-aw-firewall#8171 | | ||
| | B33 | `[DEBUG] Could not check Squid logs: EACCES ... access.log` during diagnostics, or `[DEBUG] Could not preserve squid logs: chmod ... Operation not permitted` during artifact preservation, although logs are intact | Squid writes logs as UID 13; the previous shutdown-time repair only changed mode bits, ran after mid-run diagnostics, and never transferred ownership to the runner before artifact preservation | **Fixed in AWF (PR github/gh-aw-firewall#8251, merged 2026-09-07):** reusable `fixSquidLogPermissions()` now `chown`s to the runner UID/GID via `docker exec -e` and `chmod`s; `runAgentCommand()` repairs permissions before `checkSquidLogs()`, and preserved-log `chmod` is reported separately after rename. **Additional fix (PR github/gh-aw-firewall#8624, merged 2026-09-16):** topology startup can recreate Squid mid-run (for example, for new `extra_hosts`); startup preflight repairs ownership of the explicit top-level `access.log`, `audit.jsonl`, and `cache.log` before Squid drops to the `proxy` user. The scoped, best-effort repair avoids both stale unwritable logs and a new startup-abort path. Upgrade AWF to include github/gh-aw-firewall#8624. | Trigger a blocked-domain or upstream-error diagnostic and confirm `access.log` is readable; after the run, `ls -la <preserved-squid-logs-dir>` should show runner-UID ownership and no `Could not check Squid logs`/`Could not preserve squid logs` messages; recreate Squid through topology startup and confirm `docker compose up -d` succeeds with pre-existing log files | github/gh-aw-firewall#8249, github/gh-aw-firewall#8251, github/gh-aw-firewall#8615, github/gh-aw-firewall#8624 | | ||
| | B34 | `host.docker.internal` or `(host.docker.internal/redacted)` appears in `network.allowDomains`, but an Ollama or other host-side service remains unreachable from inside AWF even when its host port is allowlisted | Host-gateway trigger detection did not recognize the canonical `host.docker.internal` keyword or its redacted audit form, so AWF omitted the host-gateway mapping needed to route the request to the runner host | **Fixed in AWF (PR github/gh-aw-firewall#8172, merged 2026-09-05):** host-gateway detection recognizes `host.docker.internal` and `(host.docker.internal/redacted)` forms and emits the required host mapping. Upgrade AWF to include github/gh-aw-firewall#8172. | Inspect generated `docker-compose.yml` for the `host.docker.internal:host-gateway` mapping and verify `getent hosts host.docker.internal` plus a request to the allowlisted host service from inside the agent | github/gh-aw-firewall#8172 | | ||
| | B35 | After a workflow using `tools.cache-memory` completes inside AWF, the host-side `validateMemoryStep` (gh-aw ≥ v0.89.21) fails with `EACCES: permission denied` writing `/tmp/gh-aw/memory-validation/cache-default.ok` | AWF remaps the agent UID/GID before running the command, while `/tmp` is bind-mounted read-write into the chroot. Files and directories created under `/tmp/gh-aw` can therefore have ownership and modes that prevent the original host runner identity from writing there after the container exits | **UNVERIFIED — do not recommend the proposal unless current status is checked:** PR github/gh-aw-firewall#9029 tracks a proposed change; consult [Provisional remediation tracking](#provisional-remediation-tracking) for historical context, then verify the live PR before recommending its proposal. | After a `cache-memory`-enabled AWF run, inspect `ls -la /tmp/gh-aw/memory-validation` (or another host-side-touched `/tmp/gh-aw` path); on affected builds, differing owner and no group-write bit reproduce the mode | github/gh-aw#63472, github/gh-aw-firewall#9028, github/gh-aw-firewall#9029 | |
|
@copilot address review feedback |
I addressed the B35 review comment in ceec789 and a follow-up commit, but the portable agent comment is still open.
|
|
✅ Copilot review passed with no inline comments. @copilot Add the |
✅ Coverage Check PassedOverall Coverage
📁 Per-file Coverage Changes (1 files)
Coverage comparison generated by |


Adds runner-doctor guidance for cache-memory permission failures after AWF exits and intermittent NVX/Cloud Hypervisor confinement false positives caused by VMM thread churn.