ci(s3tables): stop Lakekeeper flaking on Docker Hub pull timeouts (#9920)

* ci(s3tables): drop docker pre-pull from Lakekeeper job

The lakekeeper repro is pure Go against the local weed binary; the job
kept failing on Docker Hub timeouts pulling python:3 and localstack
images the test never runs. Also drop the stale python-in-docker
comments left from the old harness.

* ci(s3tables): serve python:3 from GHA cache in the STS job

Retried pulls still die when both mirror.gcr.io and registry-1.docker.io
are unreachable from the runner. Cache the saved image tarball under a
weekly key: an exact hit skips the registry entirely, a miss pulls fresh
and refreshes the cache, and a stale tarball from a previous week is the
fallback when Docker Hub is down.

* ci(spark): pre-pull the spark tag the test actually runs

The workflow warmed apache/spark:3.5.8 with retries while the
testcontainers setup runs apache/spark:3.5.1, so the real image was
pulled at test time with no retry at all.
This commit is contained in:
Chris Lu
2026-06-10 13:26:30 -07:00
committed by GitHub
parent 594fc667d5
commit caadd6ca79
3 changed files with 27 additions and 17 deletions
+1 -1
View File
@@ -45,7 +45,7 @@ jobs:
- name: Pre-pull Spark image
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull apache/spark:3.5.8
pull apache/spark:3.5.1
- name: Run S3 Spark integration tests
working-directory: test/s3/spark