Compare commits

..
27 Commits
Author SHA1 Message Date
Chris Lu 0b5d1c0c64 format: restore the hls-ts sniff and assert adapter capabilities
Sniff identifies TS assets for the coming policy-driven alignment even
though nothing reaches it through repack today. Compile-time assertions
now state each adapter's capabilities explicitly.
2026-08-10 20:03:58 -07:00
Chris Lu 56468c83e4 format: make Sniff an optional capability
Sniff sat in the mandatory adapter interface but has exactly one
caller, the repack gate, which only Indexer adapters can reach - the
hls-ts implementation was dead code. Move it to a Sniffer capability
discovered by assertion like the others: parquet keeps it, hls-ts
drops it, and repack skips the gate when an adapter cannot sniff.
2026-08-10 19:57:24 -07:00
Chris Lu bfe4b810bf filer: reject file sizes beyond int64 in the format paths
The stored size is uint64; converting a larger value wrapped negative,
sizing the repack sniff buffer with make([]byte, -1) - a handler panic -
and passing Validate a negative size, which means skip the size check.
Refuse repack and answer views stale instead.
2026-08-10 19:28:43 -07:00
Chris Lu 618febdba5 filer: inline content disqualifies format views and repack
The identity digests covered only chunks, while the view's extent path
would serve from inline Content when present - a gRPC update could set
Content with the chunks and size unchanged and segments were served
from bytes the layout never described. Format entries are never written
with inline content, so treat it as disqualifying: views answer stale,
repack rejects it up front, the extent path no longer reads it, and the
source identity digests it so it cannot appear mid-repack unnoticed.
2026-08-10 19:23:09 -07:00
Chris Lu ed9d1eec64 filer: repack conflicts on every input to its output
The commit-time check compared only the chunk list, so a concurrent
change that kept the chunks - clearing the TTL, moving the expiry
anchor, hard-linking, going remote - passed verification, and the swap
paired that fresh metadata with chunks uploaded under the old inputs: a
permanent entry pointing at chunks that still expire. Digest everything
the repack consumed - chunk fingerprint, file size, TTL, both time
anchors, the S3-expiry flag, hard-link and remote state - and answer
409 when any of it moved.
2026-08-10 19:02:58 -07:00
Chris Lu f07aabb39f filer: repack verifies the entry against the store at commit time
The entry lock is filer-local, so a writer on another filer could
commit between repack's read and its swap, and repack then restored the
old bytes over an acknowledged update. Re-read the entry and revalidate
the chunk fingerprint and WORM right before the swap, answering 409 on
any change, and build the new entry from the fresh read so concurrent
metadata-only updates are carried forward. This shrinks the unguarded
window from the whole repack to the commit itself; closing it entirely
needs owner routing.
2026-08-10 18:45:49 -07:00
Chris Lu 17fe96e620 filer: fingerprint every read-relevant chunk field
The layout binding hashed only offset and file id, so a mutation that
kept both - a truncate shrinking chunk.Size, then a sparse extend back
to the original length - passed both the size and fingerprint checks
and served a stale view. Digest size, modification timestamp, cipher
key, compression, manifest status, and SSE type as well.
2026-08-10 18:42:53 -07:00
Chris Lu 35e9f84334 filer: an unrepresentable remaining TTL means no volume TTL
The fallback capped the remainder at MaxInt32 seconds, which the volume
TTL grid encodes as 68 years - about 35 days shorter than the entry's
lifetime. For the narrow band nothing can round up within int32, store
the chunks without a volume TTL instead: they outlive the entry rather
than predecease it.
2026-08-10 18:42:18 -07:00
Chris Lu 10b64686ba filer: repack chunk TTLs round up and follow the S3 expiry anchor
SecondsToTTL truncates to the volume TTL grid, so 3599 remaining
seconds became 59m and anything under a minute became no TTL at all -
permanent chunks under an expiring entry. Round the remaining lifetime
up to the smallest representable value instead, and anchor it the way
FindEntry expires entries: S3-expiring entries age from Mtime, others
from Crtime, so a recently overwritten S3 object is no longer treated
as nearly expired.

Also bind each layout to a digest of the chunk list it described.
Offset writes and appends keep Extended while changing the chunks, so
a same-size partial write used to leave the old playlist and extents
being served over new bytes; the views now detect the mismatch and
answer 404 until the file is re-ingested or repacked.
2026-08-10 18:25:03 -07:00
Chris Lu 045c834dcf filer: revalidate WORM under the commit lock
WORM was checked before the entry lock was acquired, so a concurrent
writer could enable it while an ingest, repack, or plain HTTP overwrite
waited, and the commit then replaced a protected entry. Repack now
checks under its lock, and ingest and saveMetaData recheck at commit
time.
2026-08-10 18:23:34 -07:00
Chris Lu 318e1c64d6 filer: let repack handle S3-versioned entries
Every object version owns its chunk list, so rewriting one version's
chunks cannot affect a sibling. Drop the guard.
2026-08-10 17:32:37 -07:00
Chris Lu 2c84bb1161 filer: serialize HTTP entry commits on the entry lock
gRPC writers, renames, and repack already took the per-path entry lock,
but plain HTTP overwrites committed without it: an overwrite landing
between repack's read and its update was silently replaced, orphaning
its chunks. Take the lock around the saveMetaData and format-ingest
commits, so repack's exclusive hold spans every writer.
2026-08-10 17:27:03 -07:00
Chris Lu 61348b147b filer: derive a view-specific validator for format views
Views validated conditional requests against the media entry's ETag,
so re-ingesting identical bytes with a different sidecar changed the
playlist and segment boundaries while clients kept getting 304s. Fold
the encoded layout and the request's view parameters into the ETag the
view serves and checks.
2026-08-10 17:23:46 -07:00
Chris Lu a7fec8004e filer: repack refuses versioned entries and keeps the remaining TTL
S3 object versions may share one chunk list, so deleting the old chunks
after a repack could corrupt sibling versions; reject those entries
until chunk ownership is tracked.

New chunks also carried the full original TTL, restarting needle
expiry at repack time while entry expiry stayed anchored to creation: a
nearly expired entry left chunks stored for almost a full extra span.
Assign the remaining lifetime instead, and reject entries already past
it.
2026-08-10 17:23:23 -07:00
Chris Lu b65bcd4afa format: fix the hls-ts media-sequence decode bound
Ingest admits mediaSequence up to MaxInt64-(count-1), but the payload
decoder rejected anything above MaxInt64-count, so a boundary playlist
ingested successfully and then failed every view. Mirror the ingest
bound, covered by a round-trip test at the boundary.
2026-08-10 17:22:54 -07:00
Chris Lu 79297b549e filer: tighten the format HTTP surface
- namespace the query parameters as format.ingest, format.repack and
  format.view, following the mv.from/cp.from dotted convention, so the
  general endpoints cannot collide with pass-through client parameters;
  requests naming both ingest and repack are rejected
- state Accept-Ranges: none on view responses, which always answer with
  whole documents or whole extents
- derive the small-content permission from the boundary source instead
  of a second positional bool that a call site could silently swap
- validate the hls-ts layout before returning it, making the formattest
  invariant enforced rather than emergent
2026-08-10 00:59:38 -07:00
Chris Lu 4c7e5afbfe format: bound the encoded adapter name symmetrically
DecodeLayout rejected names over 256 bytes while Encode accepted them,
so an oversized name encoded fine and then failed every decode. Enforce
the bound in Validate, shared by both directions.
2026-08-10 00:59:38 -07:00
Chris Lu 9cb7dc7204 filer: repack keeps the entry TTL and notifies subscribers
New chunks were assigned with the TTL the request query implied while
the entry kept its own, so repacking a permanent file with ?ttl= made
its chunks expire under permanent metadata. Force the entry TTL onto
the storage option instead.

Filer.UpdateEntry only writes the store, so metadata subscribers never
heard about the new chunk ids while the old ones were queued for
deletion. Emit the update event the way the gRPC UpdateEntry path does.
2026-08-10 00:31:16 -07:00
Chris Lu 77a2b1b378 format: compute interior chunk cuts lazily
The cutter materialized every interior cut up front, so a tiny sidecar
declaring one enormous extent could allocate gigabytes of cut offsets
before any media byte arrived. Keep only the extent start offsets,
bounded by the extent count, and derive each cut arithmetically.
2026-08-10 00:31:16 -07:00
Chris Lu 7387866fd6 filer: bound format chunk sizes when no maxMB is configured
Extent chunks are buffered in memory, so an absent limit must not mean
unlimited. Also close the repack chunk reader to release its private
reader cache, and drop the arithmetic capacity hint on the extended-map
allocation.
2026-08-09 23:56:21 -07:00
Chris Lu 4bf126944f format: assert the chunk count in the align-clamp cutter test 2026-08-09 23:56:21 -07:00
Chris Lu 1737211ffa filer: keep every written chunk in the manifestization failure return
A merge failure part-way through returned only the flat data chunks,
dropping the manifests already written: cleanup paths could not delete
those needles, and AppendToEntry, which keeps the returned list after
logging the error, lost the wrapped chunks. Return the manifests plus
the not-yet-wrapped remainder instead - a complete representation of
every byte, safe to delete or to keep.
2026-08-09 23:56:21 -07:00
Chris Lu 4ee57f214c filer: wire format adapters into ingest, serving, and repack
Three hooks, all on the entry's real path so JWT scopes, WORM, and
read-only rules apply unchanged:

- POST /path?format=<name> ingests a multipart index sidecar plus media
  and cuts storage chunks on the extents the sidecar declares
- GET /path?view=<name> serves adapter views; rendered documents and
  extent streams both ride the normal prefetch path with entry ETag,
  preconditions, and HEAD support
- POST /path?repack=<name> derives the layout from the stored bytes and
  rewrites the chunks cut on extent boundaries, swapping the entry under
  the entry lock and queueing the old chunks for deletion

The layout is advisory: a stale one 404s its views while plain reads
stay untouched. Repack refuses hard-linked, remote, and SSE entries.
2026-08-09 14:36:18 -07:00
Chris Lu 211bf4d2fe format: add parquet adapter cutting extents at row-group starts
The footer already names every row-group byte range, so the adapter
only reads metadata: one extent per row group, the leading magic riding
with the first, and a trailing extent for the page indexes and footer.
Engines that fetch row groups by offset then read exactly the covering
chunks. Parquet needs no view; alignment alone delivers the benefit.
2026-08-09 14:31:39 -07:00
Chris Lu deb8b9bef1 format: add hls-ts adapter for single-file MPEG-TS VOD assets
The ingest sidecar is an FFmpeg-style EXT-X-BYTERANGE media playlist;
its segments become the extents and 188 becomes the align quantum, so
every storage chunk holds whole TS packets of one segment. The view
renders a playback playlist with plain numbered segment URLs for
clients that do not speak byte-range HLS, and maps ?seq=N to the
segment's extent. Playlist state the generated playlist cannot
reproduce (EXT-X-KEY, MAP, DISCONTINUITY, GAP, I-frame-only) is
rejected at ingest.
2026-08-09 14:30:56 -07:00
Chris Lu 53fe128511 format: add adapter registry mapping file structure to chunk extents
A format adapter reduces one container format to three things the core
understands: extent sizes, an alignment quantum, and an opaque payload.
Capabilities beyond identity (Indexer, SidecarIndexer, Viewer) are
discovered by type assertion. The layout persists in one compact
extended attribute keyed by extent sizes rather than chunk ids, so it
survives chunk manifest folding, and the Cutter turns it into upload
chunk boundaries clamped by maxMB and the align quantum. The formattest
kit holds every adapter to no-panic parsing of truncated input.
2026-08-09 14:29:27 -07:00
Chris Lu 6863412f4e filer: let the upload loop cut chunks at caller-chosen boundaries
The chunking loop always cut at a fixed size. Accept a ChunkBoundaries
source instead, with the fixed size as the default implementation, so a
caller can align storage chunks to structure inside the file. Inline
small-content storage is disabled in that mode because it would drop the
first boundary.
2026-08-09 14:28:06 -07:00
988 changed files with 15895 additions and 93764 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.9
uses: github/codeql-action/init@v4.37.4
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.9
uses: github/codeql-action/autobuild@v4.37.4
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.9
uses: github/codeql-action/analyze@v4.37.4
+8 -39
View File
@@ -6,7 +6,6 @@ on:
paths:
- 'weed/**'
- 'seaweed-volume/**'
- 'seaweed-worker/**'
- 'docker/**'
- 'go.mod'
- 'go.sum'
@@ -17,7 +16,7 @@ permissions:
jobs:
# ── Pre-build the Rust binaries natively ────────────────────────────
# ── Pre-build Rust volume server binaries natively ──────────────────
build-rust-binaries:
runs-on: ubuntu-22.04
strategy:
@@ -32,6 +31,9 @@ jobs:
- name: Checkout
uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -56,26 +58,10 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
seaweed-worker/target/${{ matrix.target }}/release
key: rust-docker-dev-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-worker/Cargo.lock') }}
key: rust-docker-dev-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-docker-dev-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build normal variant
env:
SEAWEEDFS_COMMIT: ${{ github.sha }}
@@ -84,19 +70,11 @@ jobs:
cargo build --release --target ${{ matrix.target }} --no-default-features
cp target/${{ matrix.target }}/release/weed-volume ../weed-volume-normal-${{ matrix.arch }}
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
cp target/${{ matrix.target }}/release/weed-worker ../weed-worker-${{ matrix.arch }}
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: rust-bins-${{ matrix.arch }}
path: |
weed-volume-normal-${{ matrix.arch }}
weed-worker-${{ matrix.arch }}
name: rust-volume-${{ matrix.arch }}
path: weed-volume-normal-${{ matrix.arch }}
build-dev-containers:
needs: [build-rust-binaries]
@@ -109,7 +87,7 @@ jobs:
- name: Download pre-built Rust binaries
uses: actions/download-artifact@v8
with:
pattern: rust-bins-*
pattern: rust-volume-*
merge-multiple: true
path: ./rust-bins
@@ -123,16 +101,7 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Docker meta
id: docker_meta
+9 -47
View File
@@ -59,7 +59,7 @@ jobs:
echo "publish=true" >> "$GITHUB_OUTPUT"
fi
# ── Pre-build the Rust binaries natively ────────────────────────────
# ── Pre-build Rust volume server binaries natively ──────────────────
build-rust-binaries:
runs-on: ubuntu-22.04
strategy:
@@ -76,6 +76,9 @@ jobs:
with:
ref: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.source_ref || github.ref }}
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -100,26 +103,10 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
seaweed-worker/target/${{ matrix.target }}/release
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-worker/Cargo.lock') }}
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-docker-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build large-disk variant
env:
SEAWEEDFS_COMMIT: ${{ github.sha }}
@@ -136,20 +123,13 @@ jobs:
cargo build --release --target ${{ matrix.target }} --no-default-features
cp target/${{ matrix.target }}/release/weed-volume ../weed-volume-normal-${{ matrix.arch }}
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
cp target/${{ matrix.target }}/release/weed-worker ../weed-worker-${{ matrix.arch }}
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: rust-bins-${{ matrix.arch }}
name: rust-volume-${{ matrix.arch }}
path: |
weed-volume-large-disk-${{ matrix.arch }}
weed-volume-normal-${{ matrix.arch }}
weed-worker-${{ matrix.arch }}
build:
needs: [setup, build-rust-binaries]
@@ -197,7 +177,7 @@ jobs:
- name: Download pre-built Rust binaries
uses: actions/download-artifact@v8
with:
pattern: rust-bins-*
pattern: rust-volume-*
merge-multiple: true
path: ./rust-bins
@@ -211,16 +191,7 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Docker meta
id: docker_meta
@@ -318,7 +289,7 @@ jobs:
if: needs.setup.outputs.publish != 'true'
uses: actions/download-artifact@v8
with:
pattern: rust-bins-*
pattern: rust-volume-*
merge-multiple: true
path: ./rust-bins
- name: Place Rust binaries in Docker context for local scan
@@ -336,16 +307,7 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Create BuildKit config for local scan build
if: needs.setup.outputs.publish != 'true'
run: |
@@ -405,7 +367,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.37.4
if: always()
with:
sarif_file: trivy-results.sarif
+10 -40
View File
@@ -42,10 +42,9 @@ concurrency:
jobs:
# ── Pre-build the Rust binaries natively ────────────────────────────
# The volume server and the Rust maintenance worker, cross-compiled for
# amd64 and arm64 without QEMU, turning a 5-hour emulated cargo build into
# ~15 minutes of native compilation.
# ── Pre-build Rust volume server binaries natively ──────────────────
# Cross-compiles for amd64 and arm64 without QEMU, turning a 5-hour
# emulated cargo build into ~15 minutes of native compilation.
build-rust-binaries:
runs-on: ubuntu-22.04
strategy:
@@ -60,6 +59,9 @@ jobs:
- name: Checkout
uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -84,26 +86,10 @@ jobs:
~/.cargo/registry
~/.cargo/git
seaweed-volume/target
seaweed-worker/target/${{ matrix.target }}/release
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock', 'seaweed-worker/Cargo.lock') }}
key: rust-docker-${{ matrix.target }}-${{ hashFiles('seaweed-volume/Cargo.lock') }}
restore-keys: |
rust-docker-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build large-disk variant
env:
SEAWEEDFS_COMMIT: ${{ github.sha }}
@@ -120,20 +106,13 @@ jobs:
cargo build --release --target ${{ matrix.target }} --no-default-features
cp target/${{ matrix.target }}/release/weed-volume ../weed-volume-normal-${{ matrix.arch }}
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
cp target/${{ matrix.target }}/release/weed-worker ../weed-worker-${{ matrix.arch }}
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: rust-bins-${{ matrix.arch }}
name: rust-volume-${{ matrix.arch }}
path: |
weed-volume-large-disk-${{ matrix.arch }}
weed-volume-normal-${{ matrix.arch }}
weed-worker-${{ matrix.arch }}
# One job per (variant, platform) on a native runner, pushed by digest;
# the merge job stitches the digests into one multi-arch tag.
@@ -179,7 +158,7 @@ jobs:
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
uses: actions/download-artifact@v8
with:
pattern: rust-bins-*
pattern: rust-volume-*
merge-multiple: true
path: ./rust-bins
@@ -194,16 +173,7 @@ jobs:
echo "Placed pre-built Rust binary for ${arch}"
fi
done
mkdir -p docker/weed-worker-prebuilt
for arch in amd64 arm64; do
src="./rust-bins/weed-worker-${arch}"
if [ -f "$src" ]; then
cp "$src" "docker/weed-worker-prebuilt/weed-worker-${arch}"
echo "Placed pre-built Rust worker for ${arch}"
fi
done
ls -la docker/weed-volume-prebuilt/
ls -la docker/weed-worker-prebuilt/
- name: Free Disk Space
if: github.event_name != 'workflow_dispatch' || github.event.inputs.variant == 'all' || github.event.inputs.variant == matrix.variant
@@ -430,7 +400,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.9
uses: github/codeql-action/upload-sarif@v4.37.4
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
+6 -3
View File
@@ -61,10 +61,13 @@ jobs:
- name: Install dependencies
run: |
# Use faster mirrors and install with timeout
sudo rm -f /etc/apt/sources.list.d/azure-cli.list /etc/apt/sources.list.d/microsoft-prod.list
# Same helper the e2e image installs through: the runner's own list is
# azure-only too, and an outage there fails this step outright.
sudo ./apt-install fuse
echo "deb http://azure.archive.ubuntu.com/ubuntu/ $(lsb_release -cs) main restricted universe multiverse" | sudo tee /etc/apt/sources.list
echo "deb http://azure.archive.ubuntu.com/ubuntu/ $(lsb_release -cs)-updates main restricted universe multiverse" | sudo tee -a /etc/apt/sources.list
sudo apt-get update --fix-missing
sudo DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends fuse
# Verify FUSE installation
echo "FUSE version: $(fusermount --version 2>&1 || echo 'fusermount not found')"
+2 -5
View File
@@ -30,7 +30,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code into the Go module directory
@@ -43,10 +43,7 @@ jobs:
- name: Run EC Integration Tests
working-directory: test/erasure_coding
run: |
# The suite now includes the interruption matrix and runs close to Go's
# default 10m binary timeout on slower runners; bound it by the job's
# 30m budget instead.
go test -v -timeout 25m
go test -v
- name: Collect server logs on failure
if: failure()
+3 -4
View File
@@ -40,11 +40,10 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Configure FUSE
- name: Install FUSE dependencies
run: |
# Nothing to install: fuse3 ships fusermount3 and is pre-installed,
# and go-fuse is pure Go, so the libfuse headers were never linked
# against.
sudo apt-get update
sudo apt-get install -y libfuse3-dev
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
-76
View File
@@ -1,76 +0,0 @@
name: "FUSE Volume Server Failover Tests"
on:
pull_request:
paths:
- 'weed/command/mount*.go'
- 'weed/mount/**'
- 'weed/filer/**'
- 'weed/wdclient/**'
- 'weed/operation/upload_content.go'
- 'test/fuse_failover/**'
- '.github/workflows/fuse-failover.yml'
- '.github/actions/fix-fusermount-setuid/**'
push:
branches: [master]
paths:
- 'weed/command/mount*.go'
- 'weed/mount/**'
- 'weed/filer/**'
- 'weed/wdclient/**'
- 'weed/operation/upload_content.go'
- 'test/fuse_failover/**'
- '.github/workflows/fuse-failover.yml'
- '.github/actions/fix-fusermount-setuid/**'
concurrency:
group: ${{ github.head_ref || github.ref }}/fuse-failover
cancel-in-progress: true
permissions:
contents: read
jobs:
fuse-failover:
name: FUSE Volume Server Failover
runs-on: ubuntu-22.04
timeout-minutes: 40
steps:
- name: Check out code
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Configure FUSE
run: |
# Nothing to install: fuse3 ships fusermount3 and is pre-installed,
# and go-fuse is pure Go, so the libfuse headers were never linked
# against.
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
- name: Repair the fusermount3 setuid bit
uses: ./.github/actions/fix-fusermount-setuid
- name: Build SeaweedFS
run: go build -o weed/weed -buildvcs=false ./weed
- name: Run failover integration tests
timeout-minutes: 35
env:
WEED_BINARY: ${{ github.workspace }}/weed/weed
run: go test -v -count=1 -timeout=30m ./test/fuse_failover/...
- name: Upload logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: fuse-failover-test-logs
path: /tmp/seaweedfs-fuse-failover-logs/
retention-days: 3
+5 -3
View File
@@ -38,10 +38,12 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Configure FUSE
- name: Install FUSE and dependencies
run: |
# Nothing to install: fuse3 ships fusermount3 and is pre-installed, and
# go-fuse is pure Go, so the libfuse headers were never linked against.
sudo apt-get update
# fuse3 is pre-installed on ubuntu-22.04 runners and conflicts
# with the legacy fuse package, so only install the dev headers.
sudo apt-get install -y libfuse3-dev
# Allow non-root FUSE mounts with allow_other
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
+3 -4
View File
@@ -46,11 +46,10 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Configure FUSE
- name: Install FUSE dependencies
run: |
# Nothing to install: fuse3 ships fusermount3 and is pre-installed,
# and go-fuse is pure Go, so the libfuse headers were never linked
# against.
sudo apt-get update
sudo apt-get install -y libfuse3-dev
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
-12
View File
@@ -111,16 +111,6 @@ jobs:
test:
name: Test
runs-on: ubuntu-latest
services:
redis:
image: redis:8
ports:
- 6379:6379
options: >-
--health-cmd "redis-cli ping"
--health-interval 10s
--health-timeout 5s
--health-retries 5
steps:
- name: Check out code into the Go module directory
uses: actions/checkout@v7
@@ -129,8 +119,6 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Test
env:
RUN_REDIS_TESTS: "1"
run: cd weed; go test -tags "elastic gocdk sqlite ydb tarantool tikv rclone" -v ./...
test-32bit:
+4 -108
View File
@@ -3,10 +3,10 @@ name: "helm: lint and test charts"
on:
push:
branches: [ master ]
paths: ['k8s/**', '.github/workflows/helm_ci.yml']
paths: ['k8s/**']
pull_request:
branches: [ master ]
paths: ['k8s/**', '.github/workflows/helm_ci.yml']
paths: ['k8s/**']
permissions:
contents: read
@@ -58,52 +58,10 @@ jobs:
grep -q "kind: Deployment" /tmp/s3.yaml && grep -q "seaweedfs-s3" /tmp/s3.yaml
echo "S3 deployment renders correctly"
echo "=== Testing S3 credentials from an existing secret ==="
credential_args=(
--set s3.credentials.admin.existingSecret=minio-root
--set s3.credentials.admin.accessKeyKey=root-user
--set s3.credentials.admin.secretKeyKey=root-password
--set s3.credentials.read.existingSecret=minio-root
)
for workload in \
"s3.enabled=true,s3.enableAuth=true" \
"filer.s3.enabled=true,filer.s3.enableAuth=true" \
"allInOne.enabled=true,allInOne.s3.enabled=true,allInOne.s3.enableAuth=true"
do
helm template test $CHART_DIR --set "$workload" "${credential_args[@]}" > /tmp/s3-existing-credentials.yaml
# The identities file names the variables, and each is bound to the key it was pointed at.
grep -q 'accessKey":"${SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID}' /tmp/s3-existing-credentials.yaml
grep -q 'secretKey":"${SEAWEEDFS_S3_ADMIN_SECRET_ACCESS_KEY}' /tmp/s3-existing-credentials.yaml
grep -q 'accessKey":"${SEAWEEDFS_S3_READ_ACCESS_KEY_ID}' /tmp/s3-existing-credentials.yaml
for pair in \
"SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID root-user" \
"SEAWEEDFS_S3_ADMIN_SECRET_ACCESS_KEY root-password" \
"SEAWEEDFS_S3_READ_ACCESS_KEY_ID read_access_key_id" \
"SEAWEEDFS_S3_READ_SECRET_ACCESS_KEY read_secret_access_key"
do
set -- $pair
grep -A 4 -- "- name: $1\$" /tmp/s3-existing-credentials.yaml | grep -q "name: \"minio-root\""
grep -A 4 -- "- name: $1\$" /tmp/s3-existing-credentials.yaml | grep -q "key: \"$2\""
done
# The keys stay in the user's secret rather than being copied into the chart's.
! grep -qE "^ (admin|read)_(access_key_id|secret_access_key):" /tmp/s3-existing-credentials.yaml
echo "S3 credentials reference the existing secret for $workload"
done
echo "=== Testing with all-in-one mode ==="
helm template test $CHART_DIR --set allInOne.enabled=true > /tmp/allinone.yaml
grep -q "seaweedfs-all-in-one" /tmp/allinone.yaml
echo "All-in-one deployment renders correctly"
echo "=== Testing the all-in-one s3 secret is created by every flag that mounts it ==="
for auth in allInOne.s3.enableAuth s3.enableAuth filer.s3.enableAuth; do
helm template test $CHART_DIR \
--set allInOne.enabled=true --set allInOne.s3.enabled=true --set "$auth=true" \
> /tmp/allinone-s3-auth.yaml
grep -q "secretName: test-seaweedfs-s3-secret" /tmp/allinone-s3-auth.yaml
grep -q "name: test-seaweedfs-s3-secret" /tmp/allinone-s3-auth.yaml
echo "All-in-one s3 secret renders for $auth"
done
echo "=== Testing with security enabled ==="
helm template test $CHART_DIR --set global.seaweedfs.enableSecurity=true > /tmp/security.yaml
@@ -482,41 +440,6 @@ jobs:
--set filer.s3.enableAuth=true > /tmp/filer-s3.yaml
echo "Filer S3 gateway renders correctly"
echo "=== Testing the mysql filer store gates its secret and env ==="
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MONGODB_ENABLED=true \
--set-string filer.extraEnvironmentVars.WEED_LEVELDB2_ENABLED=false > /tmp/filer-mongodb.yaml
! grep -q "db-secret" /tmp/filer-mongodb.yaml
! grep -q "WEED_MYSQL" /tmp/filer-mongodb.yaml
grep -q "name: WEED_MONGODB_ENABLED" /tmp/filer-mongodb.yaml
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MYSQL_ENABLED=true > /tmp/filer-mysql.yaml
grep -q "name: test-seaweedfs-db-secret" /tmp/filer-mysql.yaml
grep -q "name: WEED_MYSQL_USERNAME" /tmp/filer-mysql.yaml
grep -q "name: WEED_MYSQL_PASSWORD" /tmp/filer-mysql.yaml
grep -q "name: WEED_MYSQL_HOSTNAME" /tmp/filer-mysql.yaml
# Secret-backed keys follow the same rule as the plain ones.
helm template test $CHART_DIR \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.name=db \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.key=password > /tmp/filer-mysql-off-secret.yaml
! grep -q "WEED_MYSQL" /tmp/filer-mysql-off-secret.yaml
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MYSQL_ENABLED=true \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.name=db \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.key=password > /tmp/filer-mysql-on-secret.yaml
grep -A 4 -- "- name: WEED_MYSQL_PASSWORD$" /tmp/filer-mysql-on-secret.yaml | grep -q "name: db"
# A flag the chart cannot read counts as selected, not as off.
helm template test $CHART_DIR \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_ENABLED.secretKeyRef.name=store \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_ENABLED.secretKeyRef.key=enabled > /tmp/filer-mysql-secret.yaml
grep -q "name: test-seaweedfs-db-secret" /tmp/filer-mysql-secret.yaml
grep -q "name: WEED_MYSQL_HOSTNAME" /tmp/filer-mysql-secret.yaml
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MYSQL2_HOSTNAME=other > /tmp/filer-mysql2.yaml
grep -q "name: WEED_MYSQL2_HOSTNAME" /tmp/filer-mysql2.yaml
! grep -q "WEED_MYSQL_" /tmp/filer-mysql2.yaml
echo "The mysql secret and env follow the selected filer store"
echo "=== Testing SFTP enabled ==="
helm template test $CHART_DIR --set sftp.enabled=true > /tmp/sftp.yaml
grep -q "seaweedfs-sftp" /tmp/sftp.yaml
@@ -1549,37 +1472,11 @@ jobs:
echo "All template rendering tests passed!"
- name: Resolve an image tag that is published
run: |
set -e
# A release bumps appVersion on master ~40 minutes before the container
# build publishes that tag, and this workflow runs on the bump commit.
# Install the last released image for the length of that window instead
# of failing on ImagePullBackOff.
IMAGE=$(helm template test k8s/charts/seaweedfs \
-s templates/master/master-statefulset.yaml | awk '$1 == "image:" {print $2; exit}')
REPO=${IMAGE%:*}
TAG=${IMAGE##*:}
# Anything but a published tag installs latest, which between releases
# is the same digest as the chart's own appVersion, so a registry blip
# costs nothing while failing the job on one would cost a red build.
STATUS=$(curl -sSL --connect-timeout 5 --max-time 15 -o /dev/null \
-w '%{http_code}' "https://hub.docker.com/v2/repositories/$REPO/tags/$TAG" || true)
if [ "$STATUS" = 200 ]; then
echo "installing $IMAGE"
else
echo "$IMAGE is unavailable (HTTP ${STATUS:-none}), installing $REPO:latest"
TAG=latest
fi
echo "IMAGE_TAG=$TAG" >> $GITHUB_ENV
- name: Create kind cluster
uses: helm/kind-action@v1.14.0
- name: Run chart-testing (install)
run: |
ct install --target-branch ${{ github.event.repository.default_branch }} --all --chart-dirs k8s/charts \
--helm-extra-set-args "--set=image.tag=$IMAGE_TAG"
run: ct install --target-branch ${{ github.event.repository.default_branch }} --all --chart-dirs k8s/charts
- name: Verify SFTP host key secret lifecycle
run: |
@@ -1587,7 +1484,7 @@ jobs:
CHART_DIR="k8s/charts/seaweedfs"
NS="sftp-hostkey"
SECRET="hk-seaweedfs-sftp-ssh-secret"
SFTP_ARGS="--set image.tag=$IMAGE_TAG --set sftp.enabled=true --set master.enabled=false --set volume.enabled=false --set filer.enabled=false"
SFTP_ARGS="--set sftp.enabled=true --set master.enabled=false --set volume.enabled=false --set filer.enabled=false"
kubectl create namespace "$NS"
echo "=== install generates a host key, upgrade keeps it ==="
@@ -1664,7 +1561,6 @@ jobs:
# release if the hook Job does not finish, so a clean install is the
# assertion.
helm install np $CHART_DIR -n "$NS" --wait --timeout 8m \
--set image.tag=$IMAGE_TAG \
--set s3.enabled=true \
--set s3.createBuckets[0].name=testbucket \
--set networkPolicy.enabled=true \
+1 -1
View File
@@ -34,7 +34,7 @@ jobs:
id: go
- name: Set up Java
uses: actions/setup-java@v6
uses: actions/setup-java@v5
with:
java-version: ${{ matrix.java }}
distribution: 'temurin'
+1 -1
View File
@@ -73,7 +73,7 @@ jobs:
echo "version=${VERSION}" >> "$GITHUB_OUTPUT"
- name: Set up JDK 17
uses: actions/setup-java@v6
uses: actions/setup-java@v5
with:
java-version: '17'
distribution: 'temurin'
+1 -1
View File
@@ -26,7 +26,7 @@ jobs:
uses: actions/checkout@v7
- name: Set up Java
uses: actions/setup-java@v6
uses: actions/setup-java@v5
with:
java-version: ${{ matrix.java }}
distribution: 'temurin'
+1 -1
View File
@@ -42,7 +42,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
cache-dependency-path: |
**/go.sum
+14 -14
View File
@@ -43,7 +43,7 @@ jobs:
matrix:
container-id: [unit-tests-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 1.0 --memory 1g --hostname kafka-unit-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -53,7 +53,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code
@@ -87,7 +87,7 @@ jobs:
matrix:
container-id: [integration-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 2.0 --memory 2g --ulimit nofile=1024:1024 --hostname kafka-integration-${{ matrix.container-id }}
env:
GOMAXPROCS: 2
@@ -98,7 +98,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code
@@ -134,7 +134,7 @@ jobs:
matrix:
container-id: [e2e-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 2.0 --memory 2g --hostname kafka-e2e-${{ matrix.container-id }}
env:
GOMAXPROCS: 2
@@ -148,7 +148,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
cache-dependency-path: |
**/go.sum
@@ -313,7 +313,7 @@ jobs:
matrix:
container-id: [consumer-group-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 1.0 --memory 2g --ulimit nofile=512:512 --hostname kafka-consumer-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -327,7 +327,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
cache-dependency-path: |
**/go.sum
@@ -475,7 +475,7 @@ jobs:
matrix:
container-id: [client-compat-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 1.0 --memory 1.5g --shm-size 256m --hostname kafka-client-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -489,7 +489,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
cache-dependency-path: |
**/go.sum
@@ -633,7 +633,7 @@ jobs:
matrix:
container-id: [smq-integration-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 1.0 --memory 2g --hostname kafka-smq-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -647,7 +647,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
cache-dependency-path: |
**/go.sum
@@ -794,7 +794,7 @@ jobs:
matrix:
container-id: [protocol-1]
container:
image: golang:1.26-alpine
image: golang:1.24-alpine
options: --cpus 1.0 --memory 1g --tmpfs /tmp:exec --hostname kafka-protocol-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -805,7 +805,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code
-203
View File
@@ -1,203 +0,0 @@
name: "mount: benchmark"
# Manual benchmark: native WinFsp mount vs rclone+WebDAV on the same Windows
# runner, with a Linux FUSE mount of the same build as a reference. Numbers
# from shared runners are noisy; this is for finding factor-of-N gaps, not
# regressions of a few percent.
on:
workflow_dispatch:
push:
branches: [ 'winfsp-bench**' ]
concurrency:
group: ${{ github.workflow }}-${{ github.head_ref || github.run_id }}
cancel-in-progress: true
permissions:
contents: read
jobs:
bench-windows:
name: Windows native vs rclone
runs-on: windows-latest
timeout-minutes: 60
env:
CGO_ENABLED: 0
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Install WinFsp and rclone
run: choco install winfsp rclone -y --no-progress
- name: Build weed.exe
run: go build -o weed.exe ./weed
- name: Benchmark both mounts
shell: pwsh
run: |
$ErrorActionPreference = 'Stop'
function Test-Port($port) {
$client = New-Object System.Net.Sockets.TcpClient
try { $client.Connect('127.0.0.1', $port); return $client.Connected }
catch { return $false }
finally { $client.Dispose() }
}
function Wait-Drive($drive, $what) {
$deadline = (Get-Date).AddMinutes(2)
while ((Get-Date) -lt $deadline) {
if (Test-Path "${drive}\") { Write-Host "$what is mounted on $drive"; return }
Start-Sleep -Seconds 2
}
throw "$what never appeared on $drive"
}
function Invoke-Bench($dir, $label, $out) {
Write-Host "::group::bench $label"
& go run ./test/mount_bench -dir $dir -label $label -filer 127.0.0.1:8888 -out $out
$code = $LASTEXITCODE
Write-Host "::endgroup::"
if ($code -ne 0) { throw "bench $label failed with exit $code" }
}
New-Item -ItemType Directory -Force -Path C:\seaweed-data | Out-Null
Start-Process -FilePath .\weed.exe `
-ArgumentList '-logtostderr','mini','-dir=C:\seaweed-data','-ip=127.0.0.1' `
-RedirectStandardOutput C:\seaweed-mini.log -RedirectStandardError C:\seaweed-mini.err.log
$deadline = (Get-Date).AddMinutes(3)
while ((Get-Date) -lt $deadline) {
if ((Test-Port 8888) -and (Test-Port 18888) -and (Test-Port 7333)) { break }
Start-Sleep -Seconds 3
}
if (-not ((Test-Port 8888) -and (Test-Port 18888) -and (Test-Port 7333))) {
Get-Content C:\seaweed-mini.log, C:\seaweed-mini.err.log -ErrorAction SilentlyContinue
throw "mini cluster never came up"
}
Write-Host "filer on 8888/18888, webdav on 7333"
# --- native WinFsp mount ---
Start-Process -FilePath .\weed.exe `
-ArgumentList '-logtostderr','mount','-filer=127.0.0.1:8888','-dir=S:' `
-RedirectStandardOutput C:\seaweed-mount.log -RedirectStandardError C:\seaweed-mount.err.log
Wait-Drive 'S:' 'weed mount'
Invoke-Bench 'S:\bench-native' 'winfsp-native' 'C:\results-native.json'
Get-CimInstance Win32_Process -Filter "Name = 'weed.exe'" |
Where-Object { $_.CommandLine -like '*mount*' } |
ForEach-Object { Stop-Process -Id $_.ProcessId -Force }
$deadline = (Get-Date).AddMinutes(1)
while ((Get-Date) -lt $deadline -and (Test-Path S:\)) { Start-Sleep -Seconds 2 }
# --- rclone + WebDAV on the same WinFsp ---
# rclone serves listings from a directory cache it fills lazily, so the
# big-listing files have to exist before it mounts or it never sees them.
& go run ./test/mount_bench -filer 127.0.0.1:8888 -seed bench-rclone/biglist
if ($LASTEXITCODE -ne 0) { throw "seeding failed" }
$env:RCLONE_CONFIG_SEAWEED_TYPE = 'webdav'
$env:RCLONE_CONFIG_SEAWEED_URL = 'http://127.0.0.1:7333'
$env:RCLONE_CONFIG_SEAWEED_VENDOR = 'other'
Start-Process -FilePath rclone `
-ArgumentList 'mount','seaweed:','T:','--vfs-cache-mode=writes','-v','--log-file=C:\rclone.log'
Wait-Drive 'T:' 'rclone mount'
Invoke-Bench 'T:\bench-rclone' 'rclone-webdav' 'C:\results-rclone.json'
Stop-Process -Name rclone -Force -ErrorAction SilentlyContinue
# --- comparison ---
$table = & go run ./test/mount_bench -compare C:\results-native.json,C:\results-rclone.json
$table | Write-Host
"## Windows: native WinFsp vs rclone+WebDAV" | Out-File -Append $env:GITHUB_STEP_SUMMARY
$table | Out-File -Append $env:GITHUB_STEP_SUMMARY
- name: Logs
if: always()
shell: pwsh
run: |
foreach ($f in 'C:\seaweed-mount.log','C:\seaweed-mount.err.log','C:\rclone.log','C:\seaweed-mini.log','C:\seaweed-mini.err.log') {
if (Test-Path $f) { Write-Host "===== $f"; Get-Content $f -Tail 100 }
}
- name: Results
if: always()
uses: actions/upload-artifact@v7
with:
name: results-windows
path: C:\results-*.json
if-no-files-found: ignore
bench-linux:
name: Linux FUSE reference
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Repair the fusermount3 setuid bit
uses: ./.github/actions/fix-fusermount-setuid
- name: Allow non-root FUSE mounts with allow_other
run: |
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
- name: Build weed
run: go build -o /tmp/weed ./weed
- name: Benchmark FUSE mount
run: |
set -e
mkdir -p /tmp/seaweed-data
/tmp/weed -logtostderr mini -dir=/tmp/seaweed-data -ip=127.0.0.1 > /tmp/mini.log 2>&1 &
for i in $(seq 1 60); do
if nc -z 127.0.0.1 8888 && nc -z 127.0.0.1 18888; then break; fi
sleep 3
done
nc -z 127.0.0.1 8888 || { cat /tmp/mini.log; echo "filer never came up"; exit 1; }
mkdir -p "$HOME/mnt"
/tmp/weed -logtostderr mount -filer=127.0.0.1:8888 -dir="$HOME/mnt" > /tmp/mount.log 2>&1 &
for i in $(seq 1 60); do
if mountpoint -q "$HOME/mnt"; then break; fi
sleep 2
done
mountpoint -q "$HOME/mnt" || { cat /tmp/mount.log; echo "mount never appeared"; exit 1; }
go run ./test/mount_bench -dir "$HOME/mnt/bench-linux" -label linux-fuse -filer 127.0.0.1:8888 -out /tmp/results-linux.json
{
echo "## Linux FUSE reference"
go run ./test/mount_bench -compare /tmp/results-linux.json
} >> "$GITHUB_STEP_SUMMARY"
- name: Logs
if: always()
run: |
tail -n 100 /tmp/mount.log /tmp/mini.log 2>/dev/null || true
- name: Results
if: always()
uses: actions/upload-artifact@v7
with:
name: results-linux
path: /tmp/results-linux.json
if-no-files-found: ignore
+8
View File
@@ -175,6 +175,10 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
if: matrix.impl == 'rust'
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
if: matrix.impl == 'rust'
uses: dtolnay/rust-toolchain@stable
@@ -327,6 +331,10 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
if: matrix.impl == 'rust'
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
if: matrix.impl == 'rust'
uses: dtolnay/rust-toolchain@stable
+1 -1
View File
@@ -41,7 +41,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code
+12 -133
View File
@@ -1,17 +1,4 @@
name: "release: bump version and cut the release"
# One entry point for a SeaweedFS release:
# 1. bump MAJOR/MINOR in constants.go and the Helm Chart.yaml, commit to master
# 2. push the <appVersion> tag, which fans out to the workflows that trigger on
# `push: tags` (binaries_release*, container_release_unified, helm_manual_release)
# 3. create the GitHub release, with GitHub's generated notes
# 4. dispatch "Prepare release" in seaweedfs-csi-driver and seaweedfs-operator,
# which pick up the new master through `go get -u`, and wait for both
#
# Events raised by the default GITHUB_TOKEN do not start other workflows, and it
# cannot reach the other two repositories at all. Add a repo secret RELEASE_PAT
# with `contents: write` here and `actions: write` on the csi-driver and operator
# repos. Without it the tag is still pushed, but nothing downstream of it runs.
name: "release: bump version"
on:
workflow_dispatch:
@@ -27,37 +14,15 @@ on:
description: "Explicit MAJOR.MINOR to set, e.g. 4.36 (overrides 'bump')"
type: string
required: false
downstream:
description: "Also release the CSI driver and the operator"
type: boolean
default: true
dry_run:
description: "Show the version bump, but change nothing"
type: boolean
default: false
permissions:
contents: write
jobs:
release:
bump-version:
runs-on: ubuntu-latest
permissions:
contents: write
outputs:
app_version: ${{ steps.compute.outputs.app_version }}
sha: ${{ steps.tag.outputs.sha }}
steps:
- uses: actions/checkout@v7
with:
ref: master
fetch-depth: 0
token: ${{ secrets.RELEASE_PAT || secrets.GITHUB_TOKEN }}
- name: Check the release token
env:
HAS_PAT: ${{ secrets.RELEASE_PAT != '' }}
run: |
if [ "$HAS_PAT" != "true" ]; then
echo "::warning::RELEASE_PAT is not set. The tag will be pushed with GITHUB_TOKEN, so the binary, container and helm workflows will not start on their own."
fi
- name: Compute new version
id: compute
@@ -134,103 +99,17 @@ jobs:
sed -i -E "s/^version:.*/version: ${CHART_VERSION}/" "$CHART"
cat "$CHART"
- name: Commit, and push the tag
id: tag
- name: Commit and push
env:
TAG: ${{ steps.compute.outputs.app_version }}
DRY_RUN: ${{ inputs.dry_run }}
APP_VERSION: ${{ steps.compute.outputs.app_version }}
run: |
set -euo pipefail
if git ls-remote --exit-code --tags origin "refs/tags/${TAG}" >/dev/null 2>&1; then
echo "::error::Tag ${TAG} already exists."
exit 1
fi
if [ "$DRY_RUN" = "true" ]; then
git --no-pager diff --stat
echo "sha=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
if git diff --quiet; then
echo "No version change to commit."
exit 0
fi
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
if git diff --quiet; then
echo "::warning::Version files are already at ${TAG}; tagging the current HEAD."
else
git add weed/util/version/constants.go k8s/charts/seaweedfs/Chart.yaml
git commit -m "${TAG}"
git push
fi
# Push the tag with git so the `push: tags` triggers fire. Creating the
# tag through the release API alone would only emit a `create` event.
git tag "$TAG"
git push origin "$TAG"
echo "sha=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
- name: Create the release
if: ${{ !inputs.dry_run }}
env:
GH_TOKEN: ${{ secrets.RELEASE_PAT || secrets.GITHUB_TOKEN }}
TAG: ${{ steps.compute.outputs.app_version }}
run: gh release create "$TAG" --title "$TAG" --generate-notes --verify-tag
downstream:
needs: release
if: ${{ inputs.downstream && !inputs.dry_run }}
runs-on: ubuntu-latest
permissions: {}
strategy:
fail-fast: false
matrix:
include:
- repo: seaweedfs/seaweedfs-csi-driver
workflow: prepare_release.yaml
- repo: seaweedfs/seaweedfs-operator
workflow: prepare_release.yml
steps:
- name: Release ${{ matrix.repo }}
env:
GH_TOKEN: ${{ secrets.RELEASE_PAT }}
REPO: ${{ matrix.repo }}
WORKFLOW: ${{ matrix.workflow }}
SHA: ${{ needs.release.outputs.sha }}
MODULE: github.com/seaweedfs/seaweedfs
run: |
set -euo pipefail
if [ -z "${GH_TOKEN}" ]; then
echo "::error::RELEASE_PAT with actions:write on ${REPO} is required to release it"
exit 1
fi
# The dispatched workflow pins seaweedfs with `go get -u ...@latest`, so
# wait until the proxy serves the release commit as the tip. Asking for
# the commit by name is what makes the proxy fetch it.
for _ in $(seq 30); do
curl -sf "https://proxy.golang.org/${MODULE}/@v/${SHA}.info" >/dev/null || true
TIP=$(curl -sf "https://proxy.golang.org/${MODULE}/@latest" | jq -r '.Origin.Hash // ""' || true)
[ "$TIP" = "$SHA" ] && break
sleep 10
done
if [ "$TIP" != "$SHA" ]; then
echo "::error::the module proxy still serves ${TIP} as the tip, so ${REPO} would pin a pre-release commit. Run ${WORKFLOW} there once it catches up."
exit 1
fi
# Wait on the release the dispatched workflow publishes, not on the run
# that publishes it: a dispatch cannot be told apart from a concurrent
# one through the API, and the release is what we are here for.
released() { gh api "repos/${REPO}/releases?per_page=30" --jq '[.[].tag_name]'; }
BEFORE=$(released)
gh workflow run -R "$REPO" "$WORKFLOW" --ref master -f bump=patch -f update_seaweedfs=true
for _ in $(seq 80); do
sleep 15
NEW=$(released | jq -c --argjson before "$BEFORE" '. - $before')
[ "$(jq length <<<"$NEW")" -gt 0 ] && break
done
if [ "$(jq length <<<"$NEW")" -eq 0 ]; then
echo "::error::${REPO} published no release within 20 minutes; see https://github.com/${REPO}/actions/workflows/${WORKFLOW}"
exit 1
fi
echo "${REPO} released $(jq -r 'join(", ")' <<<"$NEW")"
git add weed/util/version/constants.go k8s/charts/seaweedfs/Chart.yaml
git commit -m "${APP_VERSION}"
git push
@@ -36,6 +36,9 @@ jobs:
- name: Checkout code
uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -76,6 +79,9 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -158,6 +164,9 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
-82
View File
@@ -1,82 +0,0 @@
name: "Rust Plugin Worker Tests"
on:
pull_request:
branches: [ master ]
paths:
- 'seaweed-worker/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
push:
branches: [ master, main ]
paths:
- 'seaweed-worker/**'
- 'weed/pb/plugin.proto'
- '.github/workflows/rust-worker-tests.yml'
concurrency:
group: ${{ github.workflow }}-${{ github.head_ref || github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
rust-worker-build:
name: Rust Plugin Worker Build and Unit Tests
runs-on: ubuntu-22.04
timeout-minutes: 45
steps:
- name: Checkout code
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
# cargo tracks its own inputs but not the runner's C toolchain, so a cached
# target/ can carry C objects built against a different glibc than we link against.
- name: Fingerprint build toolchain
id: toolchain
run: echo "fingerprint=$(getconf GNU_LIBC_VERSION | tr ' ' '-')-rustc-$(rustc -V | awk '{print $2}')" >> "$GITHUB_OUTPUT"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-worker/target/release
key: rust-worker-${{ steps.toolchain.outputs.fingerprint }}-${{ hashFiles('seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-worker-${{ steps.toolchain.outputs.fingerprint }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
# The release profile is what ships, and it is where the release and the
# container builds would otherwise discover a break for the first time.
- name: Build the plugin workers
run: cd seaweed-worker && cargo build --release
# The tests that need a live gateway skip themselves without one, the way
# the Go integration tests skip without Docker; the lifecycle suite in
# test/s3tables/lifecycle is what runs them against a real cluster.
# Release, so this reuses the build above rather than compiling lance,
# arrow and datafusion a second time in another profile.
- name: Run unit tests
run: cd seaweed-worker && cargo test --release --workspace
+6
View File
@@ -41,6 +41,9 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -107,6 +110,9 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: brew install protobuf
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
+10 -97
View File
@@ -1,4 +1,4 @@
name: "rust: build versioned binaries"
name: "rust: build versioned volume server binaries"
on:
push:
@@ -28,6 +28,9 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -113,102 +116,6 @@ jobs:
weed-volume_${{ matrix.asset_suffix }}.tar.gz
weed-volume_${{ matrix.asset_suffix }}.tar.gz.md5
# The Rust maintenance worker: Linux only, because it runs beside the cluster
# it maintains rather than on a laptop, and its dependency tree (lance, arrow,
# datafusion) makes every extra target an expensive build.
build-rust-worker-linux:
permissions:
contents: write
runs-on: ubuntu-22.04
strategy:
matrix:
include:
- target: x86_64-unknown-linux-gnu
asset_suffix: linux_amd64
- target: aarch64-unknown-linux-gnu
asset_suffix: linux_arm64
cross: true
steps:
- uses: actions/checkout@v7
with:
# The upload step is handed a token explicitly; a cargo build script
# should not find another one sitting in the checkout's git config.
persist-credentials: false
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
targets: ${{ matrix.target }}
- name: Install cross-compilation tools
if: matrix.cross
run: |
sudo dpkg --add-architecture arm64
sudo sed -i 's/^deb /deb [arch=amd64] /' /etc/apt/sources.list
echo "deb [arch=arm64] http://ports.ubuntu.com/ jammy main restricted universe multiverse" | sudo tee /etc/apt/sources.list.d/arm64.list
echo "deb [arch=arm64] http://ports.ubuntu.com/ jammy-updates main restricted universe multiverse" | sudo tee -a /etc/apt/sources.list.d/arm64.list
sudo apt-get update
sudo apt-get install -y gcc-aarch64-linux-gnu
echo "CARGO_TARGET_AARCH64_UNKNOWN_LINUX_GNU_LINKER=aarch64-linux-gnu-gcc" >> "$GITHUB_ENV"
- name: Cache cargo registry and target
uses: actions/cache@v6
with:
path: |
~/.cargo/registry
~/.cargo/git
seaweed-worker/target/${{ matrix.target }}/release
key: rust-worker-release-${{ matrix.target }}-${{ hashFiles('seaweed-worker/Cargo.lock') }}
restore-keys: |
rust-worker-release-${{ matrix.target }}-
# lance's build scripts compile their own protos and look for a protoc.
# Point them at the one protoc-bin-vendored ships, which seaweed-worker's
# own build already uses, so no job depends on a system package and every
# build sees the same version.
- name: Use the vendored protoc
run: |
cd seaweed-worker
cargo fetch
# The version from the lock, not whatever else a restored cache holds.
version=$(awk '/^name = "protoc-bin-vendored-linux-x86_64"$/{found=1; next} found && /^version = /{gsub(/"/,"",$3); print $3; exit}' Cargo.lock)
test -n "$version" || { echo "protoc-bin-vendored-linux-x86_64 is not in Cargo.lock" >&2; exit 1; }
protoc=$(find ~/.cargo/registry/src -path "*protoc-bin-vendored-linux-x86_64-$version/bin/protoc" | head -1)
test -x "$protoc" || { echo "no vendored protoc $version in the registry" >&2; exit 1; }
echo "PROTOC=$protoc" >> "$GITHUB_ENV"
- name: Build the Rust maintenance worker
run: |
cd seaweed-worker
cargo build --release -p weed-lance-worker --target ${{ matrix.target }}
- name: Package binary
run: |
cp seaweed-worker/target/${{ matrix.target }}/release/weed-worker weed-worker
tar czf weed-worker_${{ matrix.asset_suffix }}.tar.gz weed-worker
rm weed-worker
md5sum weed-worker_${{ matrix.asset_suffix }}.tar.gz > weed-worker_${{ matrix.asset_suffix }}.tar.gz.md5
- name: Upload release assets
if: startsWith(github.ref, 'refs/tags/')
uses: softprops/action-gh-release@v3
with:
files: |
weed-worker_${{ matrix.asset_suffix }}.tar.gz
weed-worker_${{ matrix.asset_suffix }}.tar.gz.md5
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- name: Upload artifacts
if: ${{ !startsWith(github.ref, 'refs/tags/') }}
uses: actions/upload-artifact@v7
with:
name: rust-worker-${{ matrix.asset_suffix }}
path: |
weed-worker_${{ matrix.asset_suffix }}.tar.gz
weed-worker_${{ matrix.asset_suffix }}.tar.gz.md5
build-rust-volume-darwin:
permissions:
contents: write
@@ -224,6 +131,9 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: brew install protobuf
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -301,6 +211,9 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: choco install protoc -y
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -64,14 +64,6 @@ jobs:
echo "=== Running S3 Empty Directory Marker Tests ==="
go test -v -timeout=180s -run TestS3ListObjectsEmptyDirectoryMarkers ./...
- name: Run S3 Prefix Object Tests
timeout-minutes: 15
working-directory: test/s3/normal
run: |
set -x
echo "=== Running S3 Prefix Object Tests ==="
go test -v -timeout=180s -run TestS3PrefixObjectKeys ./...
- name: Run IAM Integration Tests
timeout-minutes: 15
working-directory: test/s3/normal
+2 -2
View File
@@ -41,7 +41,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
- name: Set up Python ${{ matrix.python-version }}
@@ -148,7 +148,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
cache: true
- name: Run Go unit tests
+3 -303
View File
@@ -6,7 +6,6 @@ on:
- 'weed/s3api/**'
- 'weed/filer/**'
- 'weed/server/**'
- 'weed/worker/tasks/iceberg/**'
- 'test/s3tables/**'
- 'go.mod'
- 'go.sum'
@@ -355,19 +354,9 @@ jobs:
retention-days: 3
clickhouse-iceberg-catalog-tests:
name: ClickHouse Iceberg Catalog Integration Tests (${{ matrix.tag }})
name: ClickHouse Iceberg Catalog Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
include:
# Pinned baseline, and latest so new ClickHouse releases are
# exercised without a code change.
- clickhouse-image: clickhouse/clickhouse-server:25.8
tag: "25.8"
- clickhouse-image: clickhouse/clickhouse-server:latest
tag: latest
steps:
- name: Check out code
@@ -387,7 +376,7 @@ jobs:
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull ${{ matrix.clickhouse-image }}
pull clickhouse/clickhouse-server:25.8
pull python:3.11-slim
- name: Run go mod tidy
@@ -400,8 +389,6 @@ jobs:
- name: Run ClickHouse Iceberg Catalog Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_clickhouse
env:
CLICKHOUSE_IMAGE: ${{ matrix.clickhouse-image }}
run: |
set -x
set -o pipefail
@@ -435,7 +422,7 @@ jobs:
if: failure()
uses: actions/upload-artifact@v7
with:
name: clickhouse-iceberg-catalog-test-logs-${{ matrix.tag }}
name: clickhouse-iceberg-catalog-test-logs
path: test/s3tables/catalog_clickhouse/test-output.log
retention-days: 3
@@ -875,293 +862,6 @@ jobs:
path: test/s3tables/unity_catalog/test-output.log
retention-days: 3
lancedb-namespace-tests:
name: LanceDB Namespace Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 30
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull python:3.11-slim
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run LanceDB Namespace Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_lancedb
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting LanceDB Namespace Tests ==="
go test -v -timeout 20m . 2>&1 | tee test-output.log || {
echo "LanceDB namespace integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_lancedb
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: lancedb-namespace-test-logs
path: test/s3tables/catalog_lancedb/test-output.log
retention-days: 3
table-lifecycle-tests:
name: Table Lifecycle Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 40
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull python:3.11-slim
pull duckdb/duckdb:latest
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run Table Lifecycle Integration Tests
timeout-minutes: 35
working-directory: test/s3tables/lifecycle
env:
# The Rust worker's own tests cover its handlers; a cold build of the
# lance crate costs more here than the layer it would be checking.
WEED_LANCE_MAINTENANCE: library
run: |
set -x
set -o pipefail
go test -v -timeout 30m . 2>&1 | tee test-output.log || {
echo "Table lifecycle integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/lifecycle
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: table-lifecycle-test-logs
path: test/s3tables/lifecycle/test-output.log
retention-days: 3
duckdb-lance-tests:
name: DuckDB Lance Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 30
steps:
- name: Check out code
uses: actions/checkout@v7
with:
# The job uploads a test log on failure; nothing here needs to push,
# so do not leave a token in the checkout for it to pick up.
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull duckdb/duckdb:latest
pull python:3.11-slim
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run DuckDB Lance Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_duckdb_lance
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting DuckDB Lance Tests ==="
go test -v -timeout 20m . 2>&1 | tee test-output.log || {
echo "DuckDB Lance integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_duckdb_lance
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker|duckdb)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: duckdb-lance-test-logs
path: test/s3tables/catalog_duckdb_lance/test-output.log
retention-days: 3
spark-lance-namespace-tests:
name: Spark Lance Namespace Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 40
steps:
- name: Check out code
uses: actions/checkout@v7
with:
# The job uploads a test log on failure; nothing here needs to push,
# so do not leave a token in the checkout for it to pick up.
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull apache/spark:3.5.1
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run Spark Lance Namespace Integration Tests
timeout-minutes: 35
working-directory: test/s3tables/catalog_spark_lance
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting Spark Lance Namespace Tests ==="
go test -v -timeout 30m . 2>&1 | tee test-output.log || {
echo "Spark Lance namespace integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_spark_lance
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker|spark)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: spark-lance-namespace-test-logs
path: test/s3tables/catalog_spark_lance/test-output.log
retention-days: 3
s3-tables-build-verification:
name: S3 Tables Build Verification
runs-on: ubuntu-22.04
@@ -34,7 +34,7 @@ jobs:
uses: actions/checkout@v7
- name: Set up JDK 11
uses: actions/setup-java@v6
uses: actions/setup-java@v5
with:
java-version: '11'
distribution: 'temurin'
@@ -36,7 +36,7 @@ jobs:
- uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
- name: Build SeaweedFS
run: |
+1 -1
View File
@@ -30,7 +30,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code into the Go module directory
@@ -30,7 +30,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.26
go-version: ^1.25
id: go
- name: Check out code into the Go module directory
+12 -15
View File
@@ -61,14 +61,14 @@ Table of Contents
* [Features](#features)
* [Additional Features](#additional-features)
* [Filer Features](#filer-features)
* [Example: Using Seaweed Blob Store](#example-using-seaweed-blob-store)
* [Example: Using Seaweed Object Store](#example-using-seaweed-object-store)
* [Architecture](#object-store-architecture)
* [Compared to Other File Systems](#compared-to-other-file-systems)
* [Compared to HDFS](#compared-to-hdfs)
* [Compared to GlusterFS, Ceph](#compared-to-glusterfs-ceph)
* [Compared to GlusterFS](#compared-to-glusterfs)
* [Compared to Ceph](#compared-to-ceph)
* [Compared to MinIO, RustFS](#compared-to-minio-rustfs)
* [Compared to Minio](#compared-to-minio)
* [Dev Plan](#dev-plan)
* [Installation Guide](#installation-guide)
* [Disk Related Topics](#disk-related-topics)
@@ -90,7 +90,7 @@ S3_BUCKET=my-bucket \
./weed mini -dir=/data
```
That's it — the S3 endpoint is at http://localhost:8333, `my-bucket` already exists, and `admin`/`secret` are valid credentials. `S3_BUCKET` accepts a comma-separated list (e.g. `raw,processed`); use `S3_TABLE_BUCKET` for S3 Tables buckets, each `name` or `name:FORMAT` where the format is `ICEBERG` (the default) or `LANCE`. Drop any of the env vars to skip that piece (no AWS keys → S3 runs in unauthenticated "Allow All" mode for development).
That's it — the S3 endpoint is at http://localhost:8333, `my-bucket` already exists, and `admin`/`secret` are valid credentials. `S3_BUCKET` accepts a comma-separated list (e.g. `raw,processed`); use `S3_TABLE_BUCKET` for S3 Tables (Iceberg) buckets. Drop any of the env vars to skip that piece (no AWS keys → S3 runs in unauthenticated "Allow All" mode for development).
The same command starts everything else too:
- **S3 Endpoint**: http://localhost:8333
@@ -466,8 +466,7 @@ The architectures are mostly the same. SeaweedFS aims to store and read files fa
| GlusterFS | hashing | | FUSE, NFS | | |
| Ceph | hashing + rules | | FUSE | Yes | |
| MooseFS | in memory | | FUSE | | No |
| MinIO | separate meta file per drive for each file | | | Yes | No |
| RustFS | separate meta file per drive for each file | | | Yes | No |
| MinIO | separate meta file for each file | | | Yes | No |
[Back to TOC](#table-of-contents)
@@ -509,26 +508,24 @@ SeaweedFS Filer uses off-the-shelf stores, such as MySql, Postgres, Sqlite, Mong
[Back to TOC](#table-of-contents)
### Compared to MinIO, RustFS ###
### Compared to MinIO ###
Please note, as Apr 25, 2026 MinIO ceased development. It's strongly discouraged to use that unmaintained software with multiple security bugs. RustFS is a MinIO reimplementation in Rust, Apache 2.0 licensed and still developed, keeping MinIO's storage model down to a byte-compatible on-disk format. So the points below apply to both.
Please note, as Apr 25, 2026 MinIO ceased development. It's strongly discouraged to use that unmaintained software with multiple security bugs.
MinIO followed AWS S3 closely and was ideal for testing for S3 API. It had good UI, policies, versionings, etc. SeaweedFS is trying to catch up here.
The metadata are in simple files. Each file write incurs extra writes to the corresponding meta file, on every drive of the erasure set. Changing only tags or retention rewrites that meta file on all of them, so the write amplification does not shrink with object size.
MinIO metadata were in simple files. Each file write will incur extra writes to corresponding meta file.
There is no optimization for lots of small files. The files are simply stored as is to local disks.
MinIO did not have optimization for lots of small files. The files were simply stored as is to local disks.
Plus the extra meta file and shards for erasure coding, it only amplifies the LOSF problem.
Multiple disk IO are needed to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.
MinIO had multiple disk IO to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.
Erasure coding is full-time. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.
MinIO had full-time erasure coding. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.
No POSIX-like API support.
MinIO did not have POSIX-like API support.
There are specific requirements on storage layout, which makes it hard to scale out and to maintain. An erasure set must be 2 to 16 drives and must divide the drive list symmetrically, and capacity grows or shrinks a whole pool at a time. In SeaweedFS, just start one volume server pointing to the master. That's all.
[Back to TOC](#table-of-contents)
MinIO had specific requirements on storage layout. It is not flexible to adjust capacity. In SeaweedFS, just start one volume server pointing to the master. That's all.
## Dev Plan ##
-696
View File
@@ -1,696 +0,0 @@
# Lance Catalog for SeaweedFS
A second catalog surface next to the Iceberg REST catalog, speaking the Lance Namespace
REST spec, over the same table buckets and the same filer.
## Why
Gravitino 1.1 added a Lance REST service and 1.3 ships it as a standalone server; Lakekeeper
added Lance in the same window by a completely different route. That is the useful signal:
two unrelated catalogs decided independently that Lance had to be first-class, not a niche.
The client side is already there — `lance-spark` (`LanceNamespaceSparkCatalog`
with `impl=rest`), `lance-ray`, and the generated Python/Java/Rust clients all talk the same
OpenAPI. Implementing the spec means those engines work against SeaweedFS with no
SeaweedFS-specific code on the client.
The second reason is that Gravitino's own documentation names the gap it cannot close:
DuckDB, pandas and DataFusion "do not support Lance REST natively yet" and have to fetch a
location from the catalog and then open the dataset directly. Gravitino cannot help there,
because it does not own the storage. SeaweedFS does. That is the whole design opportunity
below.
## Prior art: three families
Upstream lists twelve catalog implementations, and they fall into three shapes. Knowing
which one we are building matters more than any individual API decision.
**1. Storage-native, no service.** The Lance Directory Catalog. V1 is a directory listing
where every `<name>.lance/` child of a prefix is a table; V2 adds a `__manifest` table —
itself a Lance table — holding `object_id`/`object_type`/`location` rows, with nested
namespaces, hash-prefixed table directories, and optional managed versioning. No server, no
credentials, no governance. This is the floor every other implementation has to beat.
**2. Protocol-native server.** Someone implements the Lance Namespace REST OpenAPI and
clients connect with `impl=rest`. Gravitino is the only one of the twelve that does this,
and it is what this design proposes.
**3. Client-side adapters onto an existing catalog.** Nine of the twelve. The Lance client
translates namespace operations into whatever the backing catalog already speaks: Apache
Polaris, Unity Catalog, AWS Glue, Hive Metastore v2 and v3, Google BigLake, Dataproc,
Microsoft OneLake — and Apache Iceberg REST. Two flavors:
- Catalogs with a real non-Iceberg table concept mark the format directly. Polaris uses its
Generic Table API with `format = lance`; Unity uses an `EXTERNAL` table with
`table_type=lance` in properties and the path in `storage_location`; Glue uses
`EXTERNAL_TABLE` plus `table_type=lance` in `Parameters`, path in
`StorageDescriptor.Location`.
- Catalogs with no such concept fake one. The Iceberg REST adapter registers **a regular
Iceberg table with a dummy schema — a single nullable string column named `dummy`** —
carrying the property `table_type=lance`, and treats the Iceberg table location as the
Lance dataset root.
Every adapter in family 3 lands in the same place: `DeclareTable`/`ListTables`/
`DescribeTable`/`DeregisterTable` only, `DropNamespace` in RESTRICT mode only,
`load_detailed_metadata=false` only, and `managed_versioning=false`. They are a name-to-
location map and nothing more.
Lakekeeper is the instructive outlier. It has the same generic-table concept Polaris has,
but no upstream adapter exists for it — there is no `lance-namespace` reference anywhere in
its repository and no page for it in the supported-catalogs list. So Polaris's generic tables
are reachable from a stock Lance client and Lakekeeper's are not, despite being the same
idea. Shipping the concept is not the same as shipping the integration.
## Gravitino and Lakekeeper: the two opposite bets
Both shipped Lance support in the same window and did not build the same thing.
**Gravitino implements the protocol.** Its `lance/` module serves the Lance Namespace REST
spec on its own port (`:9101/lance`), so stock `lance-spark` and `lance-ray` connect with
`impl=rest` and no vendor-specific client. The cost is governance: storage credentials are
static properties on the catalog (`lance.storage.access_key_id`, `secret_access_key`,
`endpoint`, `region`, `allow_http`), optionally overridden per table, handed to the engine
as-is. No STS, no expiry, no per-table scoping.
**Lakekeeper refuses the protocol and governs the object instead.** There is no
`lance-namespace` anywhere in the repository; Lance arrived in 0.13.0 (2026-06-30, issue
#1673 `Generic Table API with Lance`) as one `format` string on a Lakekeeper-native Generic
Table API:
```
POST/GET/DELETE /lakekeeper/v1/{prefix}/namespaces/{ns}/generic-tables[/{table}]
GET /lakekeeper/v1/{prefix}/namespaces/{ns}/generic-tables/{table}/credentials
POST /lakekeeper/v1/{prefix}/generic-tables/rename
```
`format` is opaque, `schema` and `statistics` are stored but never validated, and the
catalog writes no format-specific metadata — engines go straight to the location. In
exchange Lance tables get everything Iceberg tables get: STS-vended prefix-scoped
credentials, OpenFGA per-action permissions (16 actions), soft-delete with undrop, a
protection flag, rename, pagination, and name uniqueness across Iceberg tables, views and
generic tables in one namespace. The price is that no stock Lance client can talk to it —
you need `pylakekeeper`, which exists mainly to translate vended credentials into
`lance_storage_options`.
So: protocol fidelity and weak governance, or strong governance and client lock-in. Both
documented their limit honestly, and it is the same limit. Lakekeeper's capability table
says it outright — "Commit coordination: the catalog does not arbitrate writes — engines
write directly." Gravitino does not claim it either. Neither of them coordinates a Lance
commit, which is exactly the thing a store can do and a control plane cannot.
We do not have to choose. Serve the Lance protocol natively the way Gravitino does, over
the `s3tables` entries that already carry ARNs, policies, tags and maintenance config, and
the governance comes from the layer underneath rather than from a proprietary API on top.
That is only available to us because we are the store, which is also what makes the third
option — arbitrating the commit — available.
## We are probably already a Lance catalog, and that is a problem
The Iceberg REST adapter does not care whose Iceberg catalog it is talking to. It needs
`/v1/config?warehouse=`, `/v1/{prefix}/namespaces`, `/v1/{prefix}/namespaces/{ns}/tables`
and unit-separator (`\x1F`) multi-level namespaces. We serve all of those, and
`parseNamespace` in `weed/s3api/iceberg/utils.go:22` already splits on `\x1F`. So a stock
Lance client pointed at our Iceberg catalog on :8181 with the Iceberg impl should already
create, list, describe and deregister Lance tables today, with no SeaweedFS change at all.
That is worth testing before writing a line of the design above, for two reasons. It is a
free baseline — and possibly a free announcement. And it is a data-loss hazard.
A Lance table registered this way is an Iceberg table whose metadata references no data
files, sitting on top of a Lance dataset that uses `data/` for its fragments — the same
subdirectory name Iceberg uses. The maintenance worker's orphan cleaner walks exactly
`<table>/metadata` and `<table>/data`, and deletes every file not referenced by a snapshot
and older than `orphan_older_than_hours`
(`weed/worker/tasks/iceberg/operations.go:331`, default 72). Against an adapter-registered
Lance table, every fragment is unreferenced by construction. Run maintenance and the
dataset is deleted.
Maintenance is disabled by default (`handler.go:334`), so this is a latent hazard rather
than a live one: it needs an operator to enable Iceberg maintenance on a bucket that also
holds adapter-registered Lance tables. But it costs nothing to close — detection should
skip any table carrying a non-Iceberg format marker (`table_type` property, or
`Format != "ICEBERG"` once the format field is honest), and that guard is worth landing on
its own regardless of whether the rest of this design ever gets built. It is the same
"catalog-only, no maintenance" marker the generic-format question needs.
## Where we differ from Gravitino
Gravitino is a metadata service in front of somebody else's object store:
```
Spark / Ray Spark / Ray / pandas / duckdb
| |
Lance REST Lance REST (direct S3)
| | |
Gravitino SeaweedFS S3 gateway ----+
| |
S3 keys handed out SeaweedFS filer + volumes
|
somebody else's S3
```
It resolves a name to a location plus `lance.storage.*` credentials, and steps out of the
way. Everything a Lance table actually is — `_versions/`, `data/`, `_indices/` — is opaque
to it.
We are the store. Three things follow that Gravitino cannot do:
1. The catalog and a plain directory listing can be made to agree, so a client with no
catalog at all still sees the right tables.
2. `_versions/` is a filer directory listing, not an object-store `LIST`. Version history
is cheap and can back the admin UI.
3. We can offer a genuinely atomic commit reservation. Lance's commit protocol needs
put-if-not-exists; our S3 layer does not currently provide one (see
[Commit safety](#commit-safety)). The filer does.
## Placement
The Iceberg catalog is a thin HTTP shell over `s3tables.Manager`; the storage work lives in
`weed/s3api/s3tables`. Table buckets live under `TablesPath = s3_constants.DefaultBucketsPath`,
i.e. the same filer tree the S3 gateway serves, so `s3://bucket/ns/table/` is simultaneously
a catalog entry and an S3 prefix. Catalog entries are filer directories carrying `s3tables.*`
extended attributes. `Table.Format` already exists and is hard-checked against `"ICEBERG"`
in `weed/s3api/s3tables/handler_table.go:48`.
So:
```
weed/s3api/lance/ new: HTTP surface, id codec, error model
weed/s3api/s3tables/ extended: Format "LANCE", lance state xattr, version entries
weed/command/s3.go new: -port.lance (default 9101), startLanceServer
```
`Format: "LANCE"` on the table entry is the whole storage-model change for phase 1.
Everything else — namespaces, ARNs, policies, tags, ownership — is shared verbatim.
```
s3tables.Manager (filer)
|
+------------------------+------------------------+
| |
weed/s3api/iceberg weed/s3api/lance
Iceberg REST :8181 Lance REST :9101
| |
Iceberg tables Lance datasets
\ /
+-------------------- s3 :8333 -----------------+
|
SeaweedFS volumes
```
## Identifier mapping
Lance identifiers are `["ns", ..., "table"]`, encoded in the URL as a single string joined
by a delimiter that defaults to `$`. The delimiter alone means the root namespace, so
`/v1/namespace/$/list` lists the root's children.
Iceberg had to invent a warehouse selector because its identifier is flat and every table
bucket is a separate catalog. Lance does not need that — its identifier is already
hierarchical, and Gravitino uses exactly three levels (`["lance_catalog", "sales", "orders"]`).
That maps onto us without inventing anything:
```
$ root -> list of table buckets
$analytics level 1 -> a table bucket
$analytics$sales level 2 -> a namespace in that bucket
$analytics$sales$orders table
```
`spark.sql.catalog.lance.parent = analytics` then makes `sales.orders` resolve, which is the
same shape Gravitino's Spark example uses.
Levels 2..N join into one `s3tables` namespace with `.`, matching what the Iceberg catalog
already does with `flattenNamespacePath`. The flattened form is only the directory name —
`namespaceMetadata.Namespace []string` in the xattr keeps the authoritative parts, so the
mapping stays invertible even though `.` is a legal character inside a namespace part.
Reject `$` in any name part with `InvalidInput`; our charsets already exclude it, so no
escaping scheme is needed.
Root-level `ListNamespaces` returning table buckets means an unauthenticated or
broadly-scoped caller can enumerate buckets. Filter it through the same
`s3tables/permissions.go` check `ListTableBuckets` uses, not a separate path.
`CreateNamespace` on a one-part identifier creates a table bucket, and it does so only if
the caller is permitted to — the namespace never creates a bucket as a side effect of
creating something inside it. A table bucket is a tenant resource with its own policy, ARN
and lifecycle, and conjuring one because a client said `CREATE SCHEMA` is a privilege
escalation dressed as a convenience. Lakekeeper draws the same line explicitly: its client
creates tables, not warehouses.
## Storage layout
Lay tables out as:
```
s3://<table-bucket>/<flattened-namespace>/<table>/
data/
_versions/
_indices/
```
**Built without the `.lance` suffix this design originally proposed.** The suffix would have
made every namespace prefix a valid Lance Directory Catalog V1 root, since V1 recognises a
table by exactly that naming. It does not survive contact with the storage layer: the
catalog entry *is* the dataset directory, `validateTableName` excludes `.` from the charset,
and a suffixed entry name would leak into ARNs, policy documents and the S3 Tables API,
where the same table would answer to two different names. Making `GetTablePath` format-aware
instead spreads an "unless it is Lance" branch through code that has no business knowing —
the exact cross-cutting cost this design rejects family 3 for.
So one name, one directory. What survives is direct access by URI, which is the larger half
of the story and needs no naming convention at all:
```python
# with the catalog
spark.sql("SELECT * FROM lance.sales.orders")
# without it, same bytes
lance.dataset("s3://analytics/sales/orders")
```
DuckDB, pandas and DataFusion still reach the data with no catalog running, which is the gap
Gravitino's documentation admits to. What they no longer get for free is *enumeration* — a
directory-catalog client pointed at the namespace prefix will not list these as tables. If
that turns out to matter, the cheapest fix is a repair-style tool that materialises `.lance`
aliases, not a rename of the catalog entry.
Note also that the directory catalog's own V2 mode puts child-namespace tables in
`<hash>_<ns$table>` directories at the root and creates no physical subdirectories for
namespaces, so full directory-catalog fidelity was never on offer anyway. We are a
server-backed catalog; the human-readable prefix layout is worth more than partial V1
lookalike behaviour.
## The table bucket was not a neutral container
This design assumed a table bucket is a place to put a table's files. It is
not: `validateTableBucketObjectPath` runs on every S3 write into one and
validated the path against Iceberg's layout, so a Lance client got 403 on
`data/*.lance`, on `_versions/`, and on `_transactions/` — a directory Lance
writes that neither the spec documentation nor this design anticipated. Nothing
about the catalog worked end to end until that changed.
The layout guard now admits the union of what the supported formats write, and
treats any underscore-prefixed top-level directory as belonging to the format,
checking only that the path stays inside the table. Enumerating Lance's
internal directories by name is exactly the mistake that missed
`_transactions`. Iceberg writes none of them, so it loses nothing.
Found by pointing the real Python client at a running gateway, not by reading
the spec. Worth remembering for the next format: the premise to check first is
whether the storage layer will accept its files at all.
## Table lifecycle
Lance has three table states, and the spec pins them to marker files:
| State | Marker | Created by | Visible in ListTables |
| --- | --- | --- | --- |
| declared | `.lance-reserved` | `DeclareTable` | yes, when `include_declared=true` |
| created | `_versions/` present | client writes, or `CreateTable` | yes |
| deregistered | `.lance-deregistered` | `DeregisterTable` | no; data preserved |
Record the state in an xattr (`s3tables.lanceState`) on the catalog entry *and* write the
marker file into the table directory. The xattr is what the catalog reads; the marker is
what keeps a directory-catalog client honest. Dual-write is the price of the interop claim
above, and it is one extra filer write on three rarely-called operations.
`DeclareTable` is the operation `lance-spark` actually calls on `CREATE TABLE` (it replaced
the legacy `create-empty`), so it is not optional in practice even though the spec marks
only a subset as required.
`DeregisterTable` preserving data is the same shape as our Iceberg rename, where the catalog
entry moves and the data stays put — reuse `TableDataDirFromMetadataLocation`'s idea rather
than re-deriving the data path from the catalog name.
## Commit safety
This is the part I got wrong, and the correction removed a feature rather than adding one.
Lance commits a version by writing `_versions/{v}.manifest` with put-if-not-exists: exactly
one writer is supposed to win, and the loser rebases. In lance 10 that path is not optional
and needs nothing bolted on — `commit_handler_from_url` hands every `s3://` dataset a
`ConditionalPutCommitHandler`, which calls `put_opts` with `PutMode::Create`, which
object_store's S3 backend sends as `If-None-Match: *`.
I originally read our gateway as evaluating that header check-then-act, and designed around
it. That was already out of date. `buildWriteCondition`
(`weed/s3api/s3api_object_routed_write.go`) reduces `If-None-Match: *` to a filer
`WriteCondition{IF_NOT_EXISTS}`, and `putToFiler` routes the create to the object's owner
filer, which evaluates the precondition under its per-path lock; when routing is not
available it falls back to the object write lock, which evaluates it under the lock too.
Either way it is atomic. Sixteen concurrent writers of the same fresh key get one 200 and
fifteen 412s, repeatedly.
So the store already has the primitive Lance needs, cluster-wide, for every conditional-PUT
client and not just this one.
### What that removed
An earlier draft of this design offered the catalog as an **external manifest store**:
`managed_versioning: true` plus `CreateTableVersion` and friends, with the reserve step as a
filer `CreateEntry` with `o_excl`. It was implemented, tested, and shipped behind a default-off
flag — and it should not exist.
- It solves a problem this store does not have. The spec offers that path for stores that
cannot order commits themselves.
- It moves a table's version history out of the dataset and into the catalog, so a reader
that does not go through this namespace no longer sees the whole picture. That is a real
cost paid for nothing.
- lance 10 cannot even use it past the first commit: `NamespaceManifestStore::put_if_not_exists`
answers "put_if_not_exists is not supported for namespace-backed stores", which is exactly
what a second `append` needs.
The version operations now answer `Unsupported` alongside the other operations the catalog
does not serve, and `managed_versioning` is answered `false`. The property they were
protecting is covered instead by a test that races eight writers at the manifest key through
S3 and asserts one wins — testing the path Lance actually takes.
## Credential vending
Iceberg needed a header (`X-Iceberg-Access-Delegation: vended-credentials`) and a bespoke
response shape. Lance has it in the spec: `vend_credentials: true` on the request,
`storage_options` on the response, with `expires_at_millis` as the well-known expiry key.
Reuse the existing vendor interface unchanged — `iceberg.CredentialVendor` /
`STSService.AssumeRoleForPrincipal` scoped to the table prefix (#10777) — and map its output
to the storage options Lance passes through to `object_store`:
```
aws_access_key_id, aws_secret_access_key, aws_session_token,
aws_region, aws_endpoint, allow_http, expires_at_millis
```
Those are the names `pylakekeeper` emits as `lance_storage_options`, which is the shape
Lakekeeper's tested S3 path actually feeds to Lance. `object_store` also accepts the
un-prefixed aliases (`endpoint`, `region`) that the directory catalog's `storage.` prefix
strips down to and that Gravitino's `lance.storage.endpoint` resolves to, but the `aws_`
forms are the ones with a tested integration behind them, so emit those. `aws_endpoint`
should come from `deriveS3AdvertisedEndpoint()`, the same source the Iceberg `FileIO` config
uses, and `allow_http` must be set when that endpoint is plain HTTP or every read fails with
a TLS error that looks like a credential problem — Lakekeeper vends both automatically for
exactly this reason, and calls out that there is then no per-vendor branch in client code.
We emit this server-side, in the `storage_options` field the Lance spec already defines,
which is strictly better than Lakekeeper's arrangement: no client library has to translate
anything, so vending works from any stock Lance client rather than only from theirs.
Guard the same way #10777 had to after review: bucket-scoped list grants need an `s3:prefix`
condition, and a location containing `*` or `?` must be refused rather than widened into a
resource pattern.
## Auth and authorization
Authentication reuses `S3Authenticator` and `CredentialValidator` as-is. The Lance spec maps
identity to headers — `api_key` to `x-api-key`, `auth_token` to `Authorization: Bearer` — and
SigV4 keeps working because it is the same authenticator the Iceberg catalog already fronts.
Authorization needs nothing new. A Lance table gets the same ARN shape,
`arn:aws:s3tables:...:bucket/B/table/NS/T`, so every existing table-bucket policy covers
Lance tables with no new policy language and no second permission model. Route it through
`s3tables/permissions.go` and inherit the `DefaultAllow` semantics the Iceberg server already
mirrors from the S3 port.
One spec quirk worth honoring: request context entries prefixed `header.` become request
headers, and every response header comes back as a `header.`-prefixed context entry. Echoing
`x-request-id` through it costs nothing and makes tracing work.
## What to take from Lakekeeper
Rejecting Lakekeeper's API shape does not mean rejecting what it learned building it.
**Deregister is soft-delete, so implement it as one.** Lakekeeper gives generic tables
soft-deletion with undrop and a `protected` flag that makes a drop require `force=true`.
Lance already has the concept — `DeregisterTable` preserves the data and hides the table —
so the `.lance-deregistered` marker is a soft-delete by another name, and a re-register is
an undrop. A protection flag on table-bucket entries is worth having regardless of Lance:
it is a few lines against the existing xattrs and it applies to Iceberg tables too.
**Enforce one identifier space across entry kinds.** Lakekeeper rejects a generic table
whose name collides with an Iceberg table or view in the same namespace. Our catalog entries
already share one filer directory and already carry `s3tables.entryType`, so this is
structurally true — but it has to be enforced deliberately on every path, or a Lance handler
happily loads an Iceberg table's directory and vice versa. That is the same crossover bug
class as the view/table rename authorization fixed in #10776; the `catalogEntryKind` pattern
from that change is the thing to reuse rather than re-derive.
**A re-vend path matters more than it looks.** Lakekeeper exposes `/credentials` separately
from load, because STS credentials expire in the middle of long jobs and re-loading the
whole table to refresh them is wasteful. In Lance the spec's answer is another
`DescribeTable` with `vend_credentials: true`, which is fine — but it means `DescribeTable`
must stay cheap when `load_detailed_metadata` is false, which is another reason not to open
the dataset on that path.
**Generic tables are a cheap orthogonal win.** Lakekeeper's real insight is that Delta,
Parquet, CSV, Vortex and Paimon all get governance for free once the catalog stops caring
what the format is. Our `Table.Format` field already exists and the only thing stopping it
is the hard `"ICEBERG"` check in `handler_table.go:48`. Loosening that and letting the S3
Tables API register a table with an arbitrary format and a location — no metadata, no
commits — is a small change that makes every format cataloguable. It is independent of this
design and probably worth doing first, since `Format: "LANCE"` is then just a value rather
than a special case.
**Skip remote signing.** It is Lakekeeper's fallback for S3-compatible stores with no STS,
and their own documentation notes that Lance will not use it — format libraries with their
own S3 client expect static credentials and do not implement the Iceberg signer protocol. We
have STS, so vended credentials are the path, and the signer is not worth building for a
client that cannot consume it.
## Errors
Lance uses `{code, error, detail, instance}` with numeric codes, not Iceberg's exception-type
strings. The mapping is mechanical:
| HTTP | code | when |
| --- | --- | --- |
| 400 | 13 InvalidInput | charset violations, malformed id, route/body mismatch |
| 401 | 16 Unauthenticated | |
| 403 | 15 PermissionDenied | |
| 404 | 1 NamespaceNotFound, 4 TableNotFound, 11 TableVersionNotFound | |
| 409 | 2/5 AlreadyExists, 3 NamespaceNotEmpty, 14 ConcurrentModification | |
| 501 | 0 Unsupported | every phase-3 data operation |
Route/body mismatch is a spec requirement, not a nicety: when the identifier appears in both
the path and the body and they disagree, the server must return 400. Cheap to get right at
the decode step, annoying to retrofit.
## Route surface
Phase 0 is not in this table: point a stock Lance client at the existing Iceberg catalog
with the Iceberg impl, see how far it gets, and land the maintenance guard either way. That
tells us what the native server actually has to beat.
Phase 1, the whole `lance-spark` and `lance-ray` contract:
```
POST /v1/namespace/{id}/create CreateNamespace mode: Create|ExistOk|Overwrite
GET /v1/namespace/{id}/list ListNamespaces
POST /v1/namespace/{id}/describe DescribeNamespace
POST /v1/namespace/{id}/drop DropNamespace mode: Fail|Skip, behavior: Restrict|Cascade
POST /v1/namespace/{id}/exists NamespaceExists
GET /v1/namespace/{id}/table/list ListTables ?include_declared, ?page_token, ?limit
GET /v1/table ListAllTables
POST /v1/table/{id}/declare DeclareTable
POST /v1/table/{id}/describe DescribeTable ?with_table_uri, ?load_detailed_metadata, ?check_declared
POST /v1/table/{id}/exists TableExists
POST /v1/table/{id}/register RegisterTable mode: Create|Overwrite
POST /v1/table/{id}/deregister DeregisterTable
POST /v1/table/{id}/drop DropTable
POST /v1/table/{id}/rename RenameTable
```
`DescribeTable` with `load_detailed_metadata=false` needs only `location`, which is the
common case and which we can answer from xattrs alone. With `load_detailed_metadata=true`
the spec wants `version`, `schema` and `stats`, which means reading the Lance manifest. For
phase 1, return the fields we can derive from the filer — `version` from the highest entry in
`_versions/`, given V2 naming is `{u64::MAX - version:020}.manifest` and V1 is
`{version}.manifest` — and omit `schema`/`stats` rather than fabricating them. The spec
tolerates a partial response here; it does not tolerate a wrong one.
Phase 2 was the five version operations plus `managed_versioning`; it was built and then
removed, for the reasons under Commit safety.
Phase 3 is the data plane: `CreateTable`, `InsertIntoTable`, `MergeInsertIntoTable`,
`UpdateTable`, `DeleteFromTable`, `QueryTable`, `CountTableRows`, and the index and tag
operations. These exchange Arrow IPC, and more to the point they require reading and writing
the Lance file format, for which no Go implementation exists. Return `Unsupported` (code 0)
and say so in the docs. `arrow-go/v18` is already an indirect dependency, so Arrow framing is
not the blocker — Lance is.
## Does a Lance table need maintenance?
Yes, and one part of it has no Iceberg equivalent. The client exposes three jobs:
- `optimize.compact_files()` — Lance writes a fragment per write batch, so a table fed by
small appends accumulates small files exactly the way an Iceberg table does.
- `optimize.optimize_indices()` — **rows written after an index was built are not covered by
it.** A vector search against a stale index silently misses recent data. That is a
correctness-shaped failure, not a slow query, and it is specific to what people use Lance
for.
- `cleanup_old_versions()` — every version is retained until something removes it. Lance can
do this itself: `optimize.enable_auto_cleanup()` sets it on the dataset, so this one need
not be an external job at all.
None of it can run in the Go worker. All three read and rewrite Lance files, which needs
Lance format code that does not exist in Go, and there is no useful subset either: deciding
which fragments an old version still references means parsing Lance manifests.
So the maintenance worker must not touch a Lance table, and it declines by reading the format
the catalog recorded rather than by failing to parse Iceberg metadata.
## The worker can be Rust, and it is not a sidecar
The Go worker is not the only worker. `weed/pb/plugin.proto` defines `PluginControlService`,
a language-agnostic gRPC stream that external maintenance workers connect on: the worker
opens `WorkerStream`, sends `WorkerHello` with the job types it can `detect` and `execute`,
answers `RequestConfigSchema` with a `JobTypeDescriptor`, replies to `RunDetectionRequest`
with `JobProposal`s and to `ExecuteJobRequest` with `JobProgressUpdate`s and `JobCompleted`.
`weed worker -admin=host:23646` is the Go reference implementation of exactly that contract,
from outside the admin process.
Nothing in it is Go-specific, and the Rust toolchain is already in the tree.
`seaweed-volume/build.rs` compiles protos straight out of `../weed/pb/` with `tonic_build`,
including `filer.proto`, on tonic 0.12 and prost 0.13. A Lance worker is that same build
with `plugin.proto` added and the `lance` crate as a dependency — the real one, no FFI and
no Python.
Three job types, one per real maintenance operation:
| Job type | Calls | Detected from |
| --- | --- | --- |
| `lance_compact` | `optimize.compact_files` | fragment count and sizes |
| `lance_optimize_indices` | `optimize.optimize_indices` | rows an index does not cover |
| `lance_cleanup_versions` | `cleanup_old_versions` | version count and age |
What the existing machinery then supplies for free is the part worth noticing. Scheduling,
retries, dedupe by `dedupe_key`, progress reporting, per-job concurrency limits and the
admin settings page all come from the protocol: a worker that answers `RequestConfigSchema`
with a descriptor gets its configuration form rendered in the admin UI without a line of Go
or templ. A Rust worker is a first-class maintenance worker, not an appendage.
The remaining wiring is small and mostly decided already. `RunDetectionRequest` carries a
`ClusterContext` with filer and S3 addresses plus a free-form `metadata` map, which is where
the Lance namespace URL goes; the worker lists Lance tables from the namespace, which is the
catalog of record and already filters by format. It gets at the data by asking
`DescribeTable` for `storage_options` with `vend_credentials`, so the worker is just another
client of the STS path rather than a component with its own credentials. And when it commits
a compaction it goes through `CreateTableVersion` like any other writer, which is what
managed versioning was for.
## The worker is also the only thing that can describe the table
Admin can render an Iceberg table because it can read Iceberg metadata. It cannot read
Lance: it knows the dataset's location and its format string, and that is the whole of it.
The details page showed a location and two empty panels, which is an honest answer and a
useless one.
The worker already knows. Detection opens every dataset to decide whether it needs
compacting, so at that moment it holds the schema, the row count, the fragment count and
the version count. It just had no way to say so — every message on the stream was about
work.
So `WorkerObservations` is a body on `WorkerToAdminMessage`: a repeated `ObjectObservation`
of `object_id`, `object_kind`, `format`, and a `ConfigValue` map the worker fills with
whatever it can cheaply say. Admin keeps the last observation per object and serves it back
with the time it was taken and the worker that took it. Nothing schedules from it, and it is
not authoritative — it is a cache with its staleness on the label, which is why the page
badges it rather than presenting it as metadata it read itself.
The keys are the worker's to choose, which keeps the protocol out of the business of knowing
what a Lance table is. A worker for any other format admin cannot parse describes itself the
same way.
## A bucket declares its format
Format was recorded per table, which is enough for the storage layer and not enough for
anything that has to answer a question about a bucket. The admin UI printed one Iceberg
endpoint for every bucket, including the ones holding Lance datasets, where that endpoint
serves nothing; an empty bucket had no format at all.
So `CreateTableBucket` takes an optional `format`, stored with the rest of the bucket
metadata. Empty means `ICEBERG` - what AWS S3 Tables serves, and therefore what an SDK
that has never heard of the field means. `CreateTable` refuses another format, and
`CreateView` refuses outright outside an Iceberg bucket, a view being Iceberg metadata.
The Lance namespace declares `LANCE` for the buckets it creates.
**Enforced rather than defaulted**, because the point of showing a format at all is the
endpoint that follows from it, and that endpoint is only truthful if the bucket holds one
format. **Buckets that already exist stay undeclared** and keep taking anything: nothing is
migrated, and the UI shows "unset" as a fact about the bucket's age rather than a fault.
That state is also the only way to hold both formats at once, which is what the
Iceberg-REST adapter path produces.
## Sample rows are fetched, not cached
The same asymmetry has a second half. Admin renders an Iceberg table's rows by
reading its Parquet files directly; for Lance it has nothing to read with, so the data
page offered a Browse Data button that led to an empty grid.
`RequestObjectPreview` / `ObjectPreviewResponse` mirror the config-schema round trip
already on the stream: admin asks, the worker scans the dataset and hands back rows it
has already rendered as text, because it is the only side that knows the types. Admin
picks the worker from the observation store, so the one that last described a table is
the one asked to read it.
The rows are deliberately not cached, and that is the line between the two channels. An
observation describes an object, so a copy with a timestamp on it is useful. Rows are the
object's contents: a copy held in admin would be stale, larger, and nobody's business.
The page fetches on load, bounded, or says why it cannot.
## The sidecar question
The data plane is a different problem, and this design previously conflated the two.
Maintenance rides the worker protocol; `QueryTable` and `InsertIntoTable` do not, because
they are synchronous REST operations on the namespace's own surface. Serving those means a
Rust process that answers HTTP, either behind the Go namespace as a proxy target or in front
of it. It would make SeaweedFS a store you can run vector search *in* rather than one you
read vectors *out of*, which is the larger prize and the reason to keep the option open.
Neither should gate phase 1. Phases 1 and 2 are pure Go over the filer and are worth
shipping on their own — they are what makes Spark and Ray work.
## Testing
Mirror the Iceberg package: `httptest` plus a fake filer client for the handler tests, in
`weed/s3api/lance`. Then an integration suite under `test/s3tables/catalog/` next to the
existing `pyiceberg_test.go`, driving the generated Python `lance-namespace` client against
a live gateway. Three things that suite must cover and unit tests cannot:
- the storage-options key names actually work, i.e. a client that gets `storage_options` from
`DescribeTable` can open the dataset;
- a table created through the catalog is visible to `lance.dataset()` by URI and to a V1
directory-catalog client rooted at the namespace prefix;
- concurrent writers do not lose a commit, which is the phase-2 acceptance test and the
thing that justifies the external manifest store.
Phase 1 is validated: `lance_namespace` 0.11.1 with `impl=rest` drives the namespace,
`lance.write_dataset` writes to the vended location with the vended `storage_options`, and
the rows read back. Note that this client version drops `check_declared` and
`include_declared` on the wire, so `is_only_declared` reads null through it however the
server behaves.
The commit path is validated at both levels. The mechanism: eight writers race the same
manifest key through S3 with `If-None-Match: *`, and exactly one wins. The property that
actually matters, which single-winner exclusivity does not by itself establish: eight
writers append to one dataset concurrently through lance, and afterwards every batch is
still there — the losers saw the conflict, rebased, and committed again. That second test
is also the sequence managed versioning could not complete at all, since its store answers
"put_if_not_exists is not supported" to the second commit.
One more that belongs in the Iceberg suite, not this one: a Lance dataset registered through
the Iceberg adapter must survive a full maintenance pass. Reading the code, that test should
fail today; it has not been run.
## Open questions
- Root-level `ListNamespaces` enumerating table buckets is convenient and is a listing
surface we do not have on the Iceberg side. Decide whether it is gated behind a flag.
- Whether the `.lance` directory suffix is worth the divergence from the Iceberg layout. I
think yes — it is what makes the catalog optional — but it means the two catalogs' tables
do not look alike on disk, and the admin UI has to know that.
- Names: our charsets are lowercase-only and Lance identifiers are arbitrary strings. Reject
and document, as Iceberg does, or case-fold. Rejecting is right, but see #10734 for how
case handling bites when only one side normalizes.
- Whether to land generic-format registration first. Dropping the `"ICEBERG"` check and
letting a table carry an arbitrary format plus a location is smaller than this whole
design, gets Delta and Parquet catalogued as a side effect, and turns `Format: "LANCE"`
into an ordinary value. The argument against is that it invites tables the maintenance
worker cannot service, so it needs a "catalog-only, no maintenance" marker to be honest.
+15 -11
View File
@@ -2,21 +2,25 @@ FROM ubuntu:22.04
LABEL author="Chris Lu"
# Use Azure's Ubuntu mirror — much faster than archive.ubuntu.com from GitHub-hosted runners,
# which have been hanging long enough on Ign:/retry to trip the 10-min step timeout.
# Note: This e2e test image intentionally runs as root for simplicity and compatibility.
# Production images (Dockerfile.go_build) use proper user isolation with su-exec.
# For testing purposes, running as root avoids permission complexities and dependency
# on Alpine-specific tools like su-exec (not available in Ubuntu repos).
# apt-install prefers Azure's mirror, which archive.ubuntu.com is slow enough from
# GitHub-hosted runners to justify, and falls through to archive.ubuntu.com when
# Azure is unreachable - which it periodically is, and Acquire::Retries against a
# single mirror just retries a dead host. Images built FROM this one install
# through it for the same reason.
COPY apt-install /usr/local/bin/apt-install
RUN chmod +x /usr/local/bin/apt-install && \
apt-install curl fio fuse ca-certificates && \
rm -rf /tmp/* /var/tmp/*
RUN sed -i 's|http://archive.ubuntu.com/ubuntu|http://azure.archive.ubuntu.com/ubuntu|g; s|http://security.ubuntu.com/ubuntu|http://azure.archive.ubuntu.com/ubuntu|g' /etc/apt/sources.list && \
apt-get -o Acquire::http::Timeout=15 update && \
DEBIAN_FRONTEND=noninteractive apt-get -o Acquire::http::Timeout=15 install -y \
--no-install-recommends \
--no-install-suggests \
curl \
fio \
fuse \
ca-certificates \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/* \
&& rm -rf /tmp/* \
&& rm -rf /var/tmp/*
RUN mkdir -p /etc/seaweedfs /data/filerldb2
COPY ./weed /usr/bin/
+1 -1
View File
@@ -1,4 +1,4 @@
FROM golang:1.26 AS builder
FROM golang:1.25 AS builder
RUN apt-get update && \
apt-get install -y build-essential wget ca-certificates && \
+3 -18
View File
@@ -1,6 +1,6 @@
# Pin the builder to the host arch and cross-compile the (CGO-free) Go binary,
# so arm64/arm/386 targets skip QEMU emulation of the whole compile.
FROM --platform=$BUILDPLATFORM golang:1.26-alpine AS builder
FROM --platform=$BUILDPLATFORM golang:1.25-alpine AS builder
RUN apk add git g++ fuse
RUN mkdir -p /go/src/github.com/seaweedfs/
ARG BRANCH=${BRANCH:-master}
@@ -29,7 +29,6 @@ FROM alpine:3.23 as rust_builder
ARG TARGETARCH
ARG TAGS
COPY weed-volume-prebuilt/ /prebuilt/
COPY weed-worker-prebuilt/ /prebuilt-worker/
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/seaweed-volume /build/seaweed-volume
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/weed /build/weed
WORKDIR /build/seaweed-volume
@@ -48,30 +47,16 @@ RUN if [ -f "/prebuilt/weed-volume-${TARGETARCH}" ]; then \
echo "Skipping Rust build for $TARGETARCH (unsupported)" && \
touch /weed-volume; \
fi
# The Rust maintenance worker is taken pre-built or not at all: the lance jobs
# it carries pull in arrow and datafusion, a far larger dependency tree than
# the image build can carry, so an architecture CI did not build for gets the
# same empty placeholder the entrypoint refuses to exec.
RUN if [ -f "/prebuilt-worker/weed-worker-${TARGETARCH}" ]; then \
echo "Using pre-built Rust worker for ${TARGETARCH}" && \
cp "/prebuilt-worker/weed-worker-${TARGETARCH}" /weed-worker; \
else \
echo "No pre-built Rust worker for ${TARGETARCH}" && \
touch /weed-worker; \
fi
# Pre-built binaries arrive via GitHub Actions artifacts, which drop the
# executable bit, so the copied file is 0644 and exec fails with "Permission
# denied". Restore it (no-op for the empty placeholders, which stay size 0).
RUN chmod 0755 /weed-volume /weed-worker
# denied". Restore it (no-op for the empty placeholder, which stays size 0).
RUN chmod 0755 /weed-volume
FROM alpine AS final
LABEL author="Chris Lu"
COPY --from=builder /go/bin/weed /usr/bin/
# Copy Rust volume server binary (real binary on amd64/arm64, empty placeholder on other platforms)
COPY --from=rust_builder /weed-volume /usr/bin/weed-volume
# Same for the Rust maintenance worker, which serves Lance table buckets
COPY --from=rust_builder /weed-worker /usr/bin/weed-worker
RUN mkdir -p /etc/seaweedfs
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/docker/filer.toml /etc/seaweedfs/filer.toml
COPY --from=builder /go/src/github.com/seaweedfs/seaweedfs/docker/entrypoint.sh /entrypoint.sh
+1 -1
View File
@@ -1,4 +1,4 @@
FROM golang:1.26 AS builder
FROM golang:1.25 AS builder
RUN apt-get update
RUN apt-get install -y build-essential libsnappy-dev zlib1g-dev libbz2-dev libgflags-dev liblz4-dev libzstd-dev
+1 -1
View File
@@ -1,4 +1,4 @@
FROM golang:1.26 AS builder
FROM golang:1.25 AS builder
RUN apt-get update
RUN apt-get install -y build-essential libsnappy-dev zlib1g-dev libbz2-dev libgflags-dev liblz4-dev libzstd-dev
-25
View File
@@ -1,25 +0,0 @@
#!/bin/sh
# Install packages, falling through to the next Ubuntu mirror when one is
# unreachable. See docker/Dockerfile.e2e.
set -e
# Every rewrite starts from the pristine list, so a mirror that just failed does
# not become the pattern the next rewrite has to match.
[ -f /etc/apt/sources.list.orig ] || cp /etc/apt/sources.list /etc/apt/sources.list.orig
for mirror in azure.archive.ubuntu.com archive.ubuntu.com; do
# Any archive host, so this works whether the pristine list came from the
# base image (archive.ubuntu.com) or a CI runner (azure.archive.ubuntu.com).
sed "s|http://[a-z0-9.]*archive\.ubuntu\.com/ubuntu|http://$mirror/ubuntu|g; s|http://security\.ubuntu\.com/ubuntu|http://$mirror/ubuntu|g" \
/etc/apt/sources.list.orig > /etc/apt/sources.list
if apt-get -o Acquire::Retries=3 -o Acquire::http::Timeout=15 update && \
DEBIAN_FRONTEND=noninteractive apt-get -o Acquire::Retries=3 -o Acquire::http::Timeout=15 install -y \
--no-install-recommends --no-install-suggests "$@"; then
apt-get clean
rm -rf /var/lib/apt/lists/*
exit 0
fi
echo "apt: $mirror unreachable, trying the next mirror" >&2
done
exit 1
-10
View File
@@ -90,16 +90,6 @@ case "$1" in
exec /usr/bin/weed-volume $ARGS $@
;;
'worker-rust')
shift
if [ ! -s /usr/bin/weed-worker ]; then
echo "Error: Rust maintenance worker is not available on this platform ($(uname -m))." >&2
echo "Use 'worker' for the Go maintenance worker instead." >&2
exit 1
fi
exec /usr/bin/weed-worker "$@"
;;
'server')
ARGS="-dir=/data -volume.max=0 -master.volumeSizeLimitMB=1024"
if isArgPassed "-volume.max" "$@"; then
+1 -5
View File
@@ -11,8 +11,4 @@ scrape_configs:
- 'master:9324'
- 'volume:9325'
- 'filer:9326'
- 's3:9327'
# A plugin worker publishes its own metrics when started with
# -metricsPort (Go) or --metrics-port (Rust). 9328 continues the
# series, since 9327 is already the S3 gateway's.
# - 'worker:9328'
- 's3:9327'
+52 -53
View File
@@ -1,10 +1,10 @@
module github.com/seaweedfs/seaweedfs
go 1.26
go 1.25.8
require (
cloud.google.com/go v0.123.0 // indirect
cloud.google.com/go/pubsub v1.51.1
cloud.google.com/go/pubsub v1.51.0
cloud.google.com/go/storage v1.64.0
github.com/Shopify/sarama v1.38.1
github.com/aws/aws-sdk-go v1.55.8
@@ -12,7 +12,7 @@ require (
github.com/bwmarrin/snowflake v0.3.0
github.com/cenkalti/backoff/v4 v4.3.0
github.com/coreos/go-semver v0.3.1 // indirect
github.com/coreos/go-systemd/v22 v22.7.0 // indirect
github.com/coreos/go-systemd/v22 v22.6.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/dustin/go-humanize v1.0.1
github.com/eapache/go-resiliency v1.6.0 // indirect
@@ -32,7 +32,7 @@ require (
github.com/google/btree v1.1.3
github.com/google/uuid v1.6.0
github.com/google/wire v0.7.0 // indirect
github.com/googleapis/gax-go/v2 v2.24.0 // indirect
github.com/googleapis/gax-go/v2 v2.23.0 // indirect
github.com/gorilla/mux v1.8.1
github.com/hashicorp/errwrap v1.1.0 // indirect
github.com/hashicorp/go-multierror v1.1.1 // indirect
@@ -49,7 +49,7 @@ require (
github.com/kurin/blazer v0.5.3
github.com/linxGnu/grocksdb v1.10.8
github.com/mailru/easyjson v0.9.2 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect
github.com/mattn/go-isatty v0.0.23 // indirect
github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd // indirect
github.com/modern-go/reflect2 v1.0.2 // indirect
github.com/olivere/elastic/v7 v7.0.32
@@ -64,7 +64,7 @@ require (
github.com/prometheus/procfs v0.21.1
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/seaweedfs/goexif v2.0.0+incompatible
github.com/seaweedfs/goexif v1.0.3
github.com/seaweedfs/raft v1.2.0
github.com/sirupsen/logrus v1.9.4 // indirect
github.com/spf13/afero v1.15.0 // indirect
@@ -84,34 +84,34 @@ require (
github.com/xdg-go/scram v1.2.0
github.com/xdg-go/stringprep v1.0.4 // indirect
github.com/youmark/pkcs8 v0.0.0-20240726163527-a2c0da244d78 // indirect
go.etcd.io/etcd/client/v3 v3.7.1
go.etcd.io/etcd/client/v3 v3.6.12
go.mongodb.org/mongo-driver v1.17.9
go.opencensus.io v0.24.0 // indirect
gocloud.dev v0.46.0
gocloud.dev/pubsub/natspubsub v0.46.0
gocloud.dev/pubsub/rabbitpubsub v0.46.0
golang.org/x/crypto v0.55.0
golang.org/x/crypto v0.54.0
golang.org/x/exp v0.0.0-20260709172345-9ea1abe57597
golang.org/x/image v0.44.0
golang.org/x/net v0.58.0
golang.org/x/net v0.57.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.47.0
golang.org/x/text v0.41.0 // indirect
golang.org/x/text v0.40.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.294.0
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/grpc v1.85.0-dev
google.golang.org/protobuf v1.36.12
google.golang.org/api v0.289.0
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94 // indirect
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6
google.golang.org/protobuf v1.36.11
gopkg.in/inf.v0 v0.9.1 // indirect
modernc.org/b v1.0.0 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
modernc.org/sqlite v1.57.0
modernc.org/sqlite v1.53.0
)
require (
cloud.google.com/go/kms v1.33.0
cloud.google.com/go/kms v1.31.0
github.com/Azure/azure-sdk-for-go/sdk/keyvault/azkeys v0.10.0
github.com/DATA-DOG/go-sqlmock v1.5.2
github.com/Jille/raft-grpc-transport v1.6.1
@@ -122,14 +122,14 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.45.1
github.com/aws/aws-sdk-go-v2 v1.43.2
github.com/aws/aws-sdk-go-v2/config v1.32.33
github.com/aws/aws-sdk-go-v2/credentials v1.19.34
github.com/aws/aws-sdk-go-v2/credentials v1.19.32
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
github.com/fluent/fluent-logger-golang v1.10.1
github.com/getsentry/sentry-go v0.48.0
github.com/getsentry/sentry-go v0.44.1
github.com/go-ldap/ldap/v3 v3.4.13
github.com/golang-jwt/jwt/v5 v5.3.1
github.com/google/flatbuffers/go v0.0.0-20230108230133-3b8644d32c50
@@ -141,23 +141,23 @@ require (
github.com/linkedin/goavro/v2 v2.15.0
github.com/minio/crc64nvme v1.1.1
github.com/orcaman/concurrent-map/v2 v2.0.1
github.com/parquet-go/parquet-go v0.32.0
github.com/parquet-go/parquet-go v0.30.1
github.com/pkg/sftp v1.13.11
github.com/rabbitmq/amqp091-go v1.14.0
github.com/rabbitmq/amqp091-go v1.11.0
github.com/rclone/rclone v1.75.0
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
github.com/schollz/progressbar/v3 v3.19.1
github.com/seaweedfs/go-fuse/v2 v2.9.4
github.com/shirou/gopsutil/v4 v4.26.7
github.com/shirou/gopsutil/v4 v4.26.6
github.com/tarantool/go-option v1.1.0
github.com/tarantool/go-tarantool/v3 v3.0.1
github.com/tarantool/go-tarantool/v3 v3.0.0
github.com/testcontainers/testcontainers-go v0.43.0
github.com/tikv/client-go/v2 v2.0.7
github.com/xeipuuv/gojsonschema v1.2.0
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
github.com/ydb-platform/ydb-go-sdk/v3 v3.146.3
go.etcd.io/etcd/client/pkg/v3 v3.6.12
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
golang.org/x/tools/godoc v0.1.0-deprecated
@@ -183,7 +183,7 @@ require (
github.com/antlr4-go/antlr/v4 v4.13.1 // indirect
github.com/apache/arrow-go/v18 v18.7.0 // indirect
github.com/apache/thrift v0.24.0 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.5.2 // indirect
github.com/bahlo/generic-list-go v0.2.0 // indirect
github.com/bazelbuild/rules_go v0.46.0 // indirect
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f // indirect
@@ -193,7 +193,7 @@ require (
github.com/cenkalti/backoff/v5 v5.0.3 // indirect
github.com/clipperhouse/uax29/v2 v2.7.0 // indirect
github.com/cockroachdb/apd/v3 v3.2.1 // indirect
github.com/cockroachdb/errors v1.14.0 // indirect
github.com/cockroachdb/errors v1.11.3 // indirect
github.com/cockroachdb/logtags v0.0.0-20241215232642-bb51bb14a506 // indirect
github.com/cockroachdb/redact v1.1.5 // indirect
github.com/cockroachdb/version v0.0.0-20250314144055-3860cd14adf2 // indirect
@@ -238,12 +238,12 @@ require (
github.com/lpar/calendar v0.2.0 // indirect
github.com/magiconair/properties v1.8.10 // indirect
github.com/moby/docker-image-spec v1.3.1 // indirect
github.com/moby/go-archive v0.3.0 // indirect
github.com/moby/go-archive v0.2.0 // indirect
github.com/moby/moby/api v1.54.2 // indirect
github.com/moby/moby/client v0.4.0 // indirect
github.com/moby/patternmatcher v0.6.1 // indirect
github.com/moby/sys/sequential v0.7.0 // indirect
github.com/moby/sys/user v0.4.1 // indirect
github.com/moby/sys/sequential v0.6.0 // indirect
github.com/moby/sys/user v0.4.0 // indirect
github.com/moby/sys/userns v0.1.0 // indirect
github.com/moby/term v0.5.2 // indirect
github.com/oklog/ulid/v2 v2.1.1 // indirect
@@ -260,7 +260,6 @@ require (
github.com/rclone/Proton-API-Bridge v1.0.4 // indirect
github.com/rclone/go-proton-api v1.0.3 // indirect
github.com/rogpeppe/go-internal v1.15.0 // indirect
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd // indirect
github.com/ryanuber/go-glob v1.0.0 // indirect
github.com/sasha-s/go-deadlock v0.3.1 // indirect
github.com/smarty/assertions v1.15.0 // indirect
@@ -292,11 +291,11 @@ require (
require (
cel.dev/expr v0.25.2 // indirect
cloud.google.com/go/auth v0.23.2 // indirect
cloud.google.com/go/auth v0.20.0 // indirect
cloud.google.com/go/auth/oauth2adapt v0.2.8 // indirect
cloud.google.com/go/compute/metadata v0.9.0 // indirect
cloud.google.com/go/iam v1.12.0 // indirect
cloud.google.com/go/monitoring v1.30.0 // indirect
cloud.google.com/go/iam v1.11.0 // indirect
cloud.google.com/go/monitoring v1.29.0 // indirect
filippo.io/edwards25519 v1.2.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.22.0
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.14.0
@@ -323,21 +322,21 @@ require (
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.33 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.33 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.33 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.34 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.14 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.33 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4
github.com/aws/smithy-go v1.28.1
github.com/aws/aws-sdk-go-v2/service/sso v1.33.2 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.2 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.45.2
github.com/aws/smithy-go v1.27.5
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
github.com/buengese/sgzip v0.1.1 // indirect
@@ -355,7 +354,7 @@ require (
github.com/d4l3k/messagediff v1.2.1 // indirect
github.com/dgryski/go-farm v0.0.0-20200201041132-a6ae2369ad13 // indirect
github.com/dropbox/dropbox-sdk-go-unofficial/v6 v6.4.0 // indirect
github.com/ebitengine/purego v0.10.2 // indirect
github.com/ebitengine/purego v0.10.1 // indirect
github.com/elastic/gosigar v0.14.3 // indirect
github.com/emersion/go-message v0.18.2 // indirect
github.com/emersion/go-vcard v0.0.0-20260618161152-d854b7e0e2d3 // indirect
@@ -384,12 +383,12 @@ require (
github.com/gogo/protobuf v1.3.2 // indirect
github.com/golang-jwt/jwt/v4 v4.5.2 // indirect
github.com/google/s2a-go v0.1.9 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.20 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.18 // indirect
github.com/gorilla/schema v1.4.1 // indirect
github.com/gorilla/securecookie v1.1.2 // indirect
github.com/gorilla/sessions v1.4.0
github.com/grpc-ecosystem/go-grpc-middleware v1.4.0 // indirect
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 // indirect
github.com/grpc-ecosystem/grpc-gateway/v2 v2.28.0 // indirect
github.com/hashicorp/go-cleanhttp v0.5.2 // indirect
github.com/hashicorp/go-hclog v1.6.3 // indirect
github.com/hashicorp/go-immutable-radix v1.3.1 // indirect
@@ -437,7 +436,7 @@ require (
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
github.com/pierrec/lz4/v4 v4.1.28
github.com/pierrec/lz4/v4 v4.1.27
github.com/pingcap/errors v0.11.5-0.20211224045212-9687c2b0f87c // indirect
github.com/pingcap/failpoint v0.0.0-20220801062533-2eaa32854a6c // indirect
github.com/pingcap/kvproto v0.0.0-20230403051650-e166ae588106 // indirect
@@ -475,7 +474,7 @@ require (
github.com/winfsp/cgofuse v1.6.1-0.20260126094232-f2c4fccdb286
github.com/xanzy/ssh-agent v0.3.3 // indirect
github.com/yandex-cloud/go-genproto v0.0.0-20211115083454-9ca41db5ed9e // indirect
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260810122915-65bfd5c4b705 // indirect
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260428144813-1c07baab7f7b // indirect
github.com/ydb-platform/ydb-go-yc v0.12.1 // indirect
github.com/ydb-platform/ydb-go-yc-metadata v0.6.1 // indirect
github.com/yunify/qingstor-sdk-go/v3 v3.2.0 // indirect
@@ -483,7 +482,7 @@ require (
github.com/zeebo/blake3 v0.2.4 // indirect
github.com/zeebo/errs v1.4.0 // indirect
go.etcd.io/bbolt v1.5.0 // indirect
go.etcd.io/etcd/api/v3 v3.7.1 // indirect
go.etcd.io/etcd/api/v3 v3.6.12 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.68.0 // indirect
@@ -497,13 +496,13 @@ require (
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.45.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 // indirect
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260715232425-e75dac1f907d // indirect
gopkg.in/natefinch/lumberjack.v2 v2.2.1 // indirect
gopkg.in/validator.v2 v2.0.1 // indirect
gopkg.in/yaml.v2 v2.4.0 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
modernc.org/libc v1.74.4 // indirect
modernc.org/libc v1.73.4 // indirect
moul.io/http2curl/v2 v2.3.0 // indirect
sigs.k8s.io/yaml v1.6.0 // indirect
storj.io/common v0.0.0-20260629224719-ba1bff0a7846 // indirect
+112 -114
View File
@@ -94,8 +94,8 @@ cloud.google.com/go/assuredworkloads v1.7.0/go.mod h1:z/736/oNmtGAyU47reJgGN+KVo
cloud.google.com/go/assuredworkloads v1.8.0/go.mod h1:AsX2cqyNCOvEQC8RMPnoc0yEarXQk6WEKkxYfL6kGIo=
cloud.google.com/go/assuredworkloads v1.9.0/go.mod h1:kFuI1P78bplYtT77Tb1hi0FMxM0vVpRC7VVoJC3ZoT0=
cloud.google.com/go/assuredworkloads v1.10.0/go.mod h1:kwdUQuXcedVdsIaKgKTp9t0UJkE5+PAVNhdQm4ZVq2E=
cloud.google.com/go/auth v0.23.2 h1:pxSCpfiji41hpzpPdMCftEUCezpgpqmmDdYiAjCKXxo=
cloud.google.com/go/auth v0.23.2/go.mod h1:4DhBRcqvtljQN3dJ57qtqbib5ZGCYE5f2crfiiC2EM0=
cloud.google.com/go/auth v0.20.0 h1:kXTssoVb4azsVDoUiF8KvxAqrsQcQtB53DcSgta74CA=
cloud.google.com/go/auth v0.20.0/go.mod h1:942/yi/itH1SsmpyrbnTMDgGfdy2BUqIKyd0cyYLc5Q=
cloud.google.com/go/auth/oauth2adapt v0.2.8 h1:keo8NaayQZ6wimpNSmW5OPc283g65QNIiLpZnkHRbnc=
cloud.google.com/go/auth/oauth2adapt v0.2.8/go.mod h1:XQ9y31RkqZCcwJWNSx2Xvric3RrU88hAYYbjDWYDL+c=
cloud.google.com/go/automl v1.5.0/go.mod h1:34EjfoFGMZ5sgJ9EoLsRtdPSNZLcfflJR39VbVNS2M0=
@@ -283,8 +283,8 @@ cloud.google.com/go/iam v0.7.0/go.mod h1:H5Br8wRaDGNc8XP3keLc4unfUUZeyH3Sfl9XpQE
cloud.google.com/go/iam v0.8.0/go.mod h1:lga0/y3iH6CX7sYqypWJ33hf7kkfXJag67naqGESjkE=
cloud.google.com/go/iam v0.11.0/go.mod h1:9PiLDanza5D+oWFZiH1uG+RnRCfEGKoyl6yo4cgWZGY=
cloud.google.com/go/iam v0.12.0/go.mod h1:knyHGviacl11zrtZUoDuYpDgLjvr28sLQaG0YB2GYAY=
cloud.google.com/go/iam v1.12.0 h1:Aki3bX9aHUDKPHfnRJfDcTdVedvy6quGBQcTqx3DRXk=
cloud.google.com/go/iam v1.12.0/go.mod h1:FEZ4lXpADAC2AIpQY7LANNjjwyQ2jK439CI2VaD+sLY=
cloud.google.com/go/iam v1.11.0 h1:KieQ9Pb+LLPak1O3Rv3GgCxhnmkYf7Xyh0P5HfF1jFM=
cloud.google.com/go/iam v1.11.0/go.mod h1:KP+nKGugNJW4LcLx1uEZcq1ok5sQHFaQehQNl4QDgV4=
cloud.google.com/go/iap v1.4.0/go.mod h1:RGFwRJdihTINIe4wZ2iCP0zF/qu18ZwyKxrhMhygBEc=
cloud.google.com/go/iap v1.5.0/go.mod h1:UH/CGgKd4KyohZL5Pt0jSKE4m3FR51qg6FKQ/z/Ix9A=
cloud.google.com/go/iap v1.6.0/go.mod h1:NSuvI9C/j7UdjGjIde7t7HBz+QTwBcapPE07+sSRcLk=
@@ -298,8 +298,8 @@ cloud.google.com/go/kms v1.4.0/go.mod h1:fajBHndQ+6ubNw6Ss2sSd+SWvjL26RNo/dr7uxs
cloud.google.com/go/kms v1.5.0/go.mod h1:QJS2YY0eJGBg3mnDfuaCyLauWwBJiHRboYxJ++1xJNg=
cloud.google.com/go/kms v1.6.0/go.mod h1:Jjy850yySiasBUDi6KFUwUv2n1+o7QZFyuUJg6OgjA0=
cloud.google.com/go/kms v1.9.0/go.mod h1:qb1tPTgfF9RQP8e1wq4cLFErVuTJv7UsSC915J8dh3w=
cloud.google.com/go/kms v1.33.0 h1:pG0X78m212b2pv9N4fdMoUO69LuZGQ9kSvn8sHBOFAo=
cloud.google.com/go/kms v1.33.0/go.mod h1:CSGvW6GnMQbY+1nOHcIzhMtHSbExXlOmCKjWtYVjcpA=
cloud.google.com/go/kms v1.31.0 h1:LS8N92OxFDgOLg5NCo3OmbvjtQAIVT5gUHVLKIDHaFE=
cloud.google.com/go/kms v1.31.0/go.mod h1:YIyXZym11R5uovJJt4oN5eUL3oPmirF3yKeIh6QAf4U=
cloud.google.com/go/language v1.4.0/go.mod h1:F9dRpNFQmJbkaop6g0JhSBXCNlO90e1KWx5iDdxbWic=
cloud.google.com/go/language v1.6.0/go.mod h1:6dJ8t3B+lUYfStgls25GusK04NLh3eDLQnWM3mdEbhI=
cloud.google.com/go/language v1.7.0/go.mod h1:DJ6dYN/W+SQOjF8e1hLQXMF21AkH2w9wiPzPCJa2MIE=
@@ -310,8 +310,8 @@ cloud.google.com/go/lifesciences v0.6.0/go.mod h1:ddj6tSX/7BOnhxCSd3ZcETvtNr8NZ6
cloud.google.com/go/lifesciences v0.8.0/go.mod h1:lFxiEOMqII6XggGbOnKiyZ7IBwoIqA84ClvoezaA/bo=
cloud.google.com/go/logging v1.6.1/go.mod h1:5ZO0mHHbvm8gEmeEUHrmDlTDSu5imF6MUP9OfilNXBw=
cloud.google.com/go/logging v1.7.0/go.mod h1:3xjP2CjkM3ZkO73aj4ASA5wRPGGCRrPIAeNqVNkzY8M=
cloud.google.com/go/logging v1.19.0 h1:NCqhdVUg3wQ8Cobdf16FDSuTGi3+6+hdSBHrY5TsR6Q=
cloud.google.com/go/logging v1.19.0/go.mod h1:i40NZCHC9Gqvod4yE+yQfDWwlgwW/SrshkkGibCHxcA=
cloud.google.com/go/logging v1.18.0 h1:KhzZq+1cSkPH9YUaKLLhLtQxIHitVayBmk0sGfoM9+k=
cloud.google.com/go/logging v1.18.0/go.mod h1:ZGKnpBaURITh+g/uom2VhbiFoFWvejcrHPDhxFtU/gI=
cloud.google.com/go/longrunning v0.1.1/go.mod h1:UUFxuDWkv22EuY93jjmDMFT5GPQKeFVJBIF6QlTqdsE=
cloud.google.com/go/longrunning v0.3.0/go.mod h1:qth9Y41RRSUE69rDcOn6DdK3HfQfsUI0YSmW3iIlLJc=
cloud.google.com/go/longrunning v0.4.1/go.mod h1:4iWDqhBZ70CvZ6BfETbvam3T8FMvLK+eFj0E6AaRQTo=
@@ -338,8 +338,8 @@ cloud.google.com/go/metastore v1.10.0/go.mod h1:fPEnH3g4JJAk+gMRnrAnoqyv2lpUCqJP
cloud.google.com/go/monitoring v1.7.0/go.mod h1:HpYse6kkGo//7p6sT0wsIC6IBDET0RhIsnmlA53dvEk=
cloud.google.com/go/monitoring v1.8.0/go.mod h1:E7PtoMJ1kQXWxPjB6mv2fhC5/15jInuulFdYYtlcvT4=
cloud.google.com/go/monitoring v1.12.0/go.mod h1:yx8Jj2fZNEkL/GYZyTLS4ZtZEZN8WtDEiEqG4kLK50w=
cloud.google.com/go/monitoring v1.30.0 h1:r/d+JUbyKmJ8b07iznuKfzVzrIXTWxHQ3lBRm3x2LlY=
cloud.google.com/go/monitoring v1.30.0/go.mod h1:htlUR0QWVMrjFzZmN4LGnMAve9xB/eduwjmINxVZ8RM=
cloud.google.com/go/monitoring v1.29.0 h1:AHhDsFaSax1/4k+qlIDX/SDGe6hggnfXJ9dkgD9qBPY=
cloud.google.com/go/monitoring v1.29.0/go.mod h1:72NOVjJXHY/HBfoLT0+qlCZBT059+9VXLeAnL2PeeVM=
cloud.google.com/go/networkconnectivity v1.4.0/go.mod h1:nOl7YL8odKyAOtzNX73/M5/mGZgqqMeryi6UPZTk/rA=
cloud.google.com/go/networkconnectivity v1.5.0/go.mod h1:3GzqJx7uhtlM3kln0+x5wyFvuVH1pIBJjhCpjzSt75o=
cloud.google.com/go/networkconnectivity v1.6.0/go.mod h1:OJOoEXW+0LAxHh89nXd64uGG+FbQoeH8DtxCHVOMlaM=
@@ -391,8 +391,8 @@ cloud.google.com/go/pubsub v1.3.1/go.mod h1:i+ucay31+CNRpDW4Lu78I4xXG+O1r/MAHgjp
cloud.google.com/go/pubsub v1.26.0/go.mod h1:QgBH3U/jdJy/ftjPhTkyXNj543Tin1pRYcdcPRnFIRI=
cloud.google.com/go/pubsub v1.27.1/go.mod h1:hQN39ymbV9geqBnfQq6Xf63yNhUAhv9CZhzp5O6qsW0=
cloud.google.com/go/pubsub v1.28.0/go.mod h1:vuXFpwaVoIPQMGXqRyUQigu/AX1S3IWugR9xznmcXX8=
cloud.google.com/go/pubsub v1.51.1 h1:R3G1wCOxBO7jRpL8x2pdZMv1GAJDF6ax/m2zPOtvTNE=
cloud.google.com/go/pubsub v1.51.1/go.mod h1:y2T0IKtW1iWwVvazYaRpqOAFO4gy2+O7dTDt9TWY/5U=
cloud.google.com/go/pubsub v1.51.0 h1:XOaCejsqX7EEtUdQz+WPag66wWsUUGliyCOfGPKfo90=
cloud.google.com/go/pubsub v1.51.0/go.mod h1:NERXf11sd82UV3VnflcUj8POIyQUXT/QwrKlxD8di/I=
cloud.google.com/go/pubsub/v2 v2.6.0 h1:8pjR0id+GTB+krKx5G6AGJoYrHog58w2Q89PCOrfM64=
cloud.google.com/go/pubsub/v2 v2.6.0/go.mod h1:4anqvV/w8Pcgu2tO0qr2XgsF3GXHowzryfQ5gOnVmWY=
cloud.google.com/go/pubsublite v1.5.0/go.mod h1:xapqNQ1CuLfGi23Yda/9l4bBCKz/wC3KIJ5gKcxveZg=
@@ -707,50 +707,50 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.45.1 h1:iIoG3NaLhV6UZpPXyPXlDj2I9oS8tV/nMcMnITCC6Ks=
github.com/aws/aws-sdk-go-v2 v1.45.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU=
github.com/aws/aws-sdk-go-v2 v1.43.2 h1:cl+IXwWb3qazClUcm08tGSsB6OiuV83JVJO9B0jQcPc=
github.com/aws/aws-sdk-go-v2 v1.43.2/go.mod h1:WEzLKBh/mEjXvx1FtQMWgSxMSTVqxQzjkRtk5fa3wkg=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 h1:3IZY0XAJquT3aHzbkHfPzy4ACPcEjVG0x87KOwtpqGY=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14/go.mod h1:zwM6veDkhGgQFqkBy+uT28AAYpLu+uFMlPl+rCg/73E=
github.com/aws/aws-sdk-go-v2/config v1.32.33 h1:M1m/Q6f0OKDEDGwhiNOqx1OjTdrewe3v+GDbHmKczWk=
github.com/aws/aws-sdk-go-v2/config v1.32.33/go.mod h1:fGj1iQj2QpIZzp7jE4aQQ+71TE8cd4z9K4+xCd6EqmE=
github.com/aws/aws-sdk-go-v2/credentials v1.19.34 h1:y6GkSmcv5myd1ngrYbGmiLlwQqB6TQhOuN/tbSSuWDY=
github.com/aws/aws-sdk-go-v2/credentials v1.19.34/go.mod h1:w3dTcnDVoQIewjo7JG45hduAToikiIFLC4FIO7fndvw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35 h1:+S7kbJoLDDQ5tE+lHrUBgMkzC8NLgsaioS2F3dVoFAE=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35/go.mod h1:Ak7xXviIARfFdNUJ9Etb0bdVDt/KAvKjMGJVLWXDzik=
github.com/aws/aws-sdk-go-v2/credentials v1.19.32 h1:eNE0JnIblBo1NCvd3tqEYuZz9XDefn69R74CHd3nT7U=
github.com/aws/aws-sdk-go-v2/credentials v1.19.32/go.mod h1:yYJu+6tqKUYZuJSYcpSGjz/6sV/SUaAaKIufnWKx2OU=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.33 h1:MobhiR6KIerWxmO74Zit5I3379+mSc2DOdZ3DeRFB9w=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.33/go.mod h1:xu02847OdZfNr/jAfZpHtyRk0b3v4d0kaoxNHxZGG/w=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 h1:Pn7OsMwBLbkZ6OnCxWHAjf0L/22H8cnhxZC0uPwtMtg=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34/go.mod h1:eToXR/Gk1uqpn04eSmdgVXwfS0WvH8aG4eBFr8ygbpU=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.2.3 h1:w5OoDiMN6x53ROmiIImGzmVcxXv2q1GXY+aKV4WAJYM=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.2.3/go.mod h1:dAhgYp776bX3LuWvnSCFwQEjNs6fuFg7YXIy5PXcP3Q=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35 h1:kzVuGlatQtYinwBJEEyLAbggepCoavosiaHHX9+fD+c=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35/go.mod h1:0yLx0yEI+SfqeJMPvOtIEFoZbiQYXMGszBueiutQyaI=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35 h1:WK6CjihTuLisCjSKKbildJ79sGZZgbBz3iNa7VsKIhU=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35/go.mod h1:KYleN57luLoe97R7vTnx8PMcVrr9gAcRECtOjl91DNg=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36 h1:jbGY4CXLzZElOXgGsexlC3Hi+3YM0rSmk4opFXKqg/k=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36/go.mod h1:uBu/9aKsS/UQGc72RAt3y54kjgYQxmhut8ZD2dXCDNE=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15 h1:JJLBQxwY+AFwuPAi5ivGc1ChnTdUt4cXMv7e76m2c/Y=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15/go.mod h1:lQknBIe78MVL0cQOQDlag8KGflMbMEVFx9mB6O8ENvk=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.33 h1:HAp1wLFZzch054uh3FK7rcVYg4v7J2FxVf3h3IGNZas=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.33/go.mod h1:mJk5fmqnF+WUlMdPG37pR2Fh3oh6r8F6ZGUgPKvzu0c=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.33 h1:0YA0aCKgsJyno6xkFfaIgjE3/wK08+Qxo9nQfe1UrWM=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.33/go.mod h1:UZqj4WIdTH+ga8Y/DgpAuy/8cGjM3h7gDCliJYGg2SE=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.34 h1:HQYnjFnXpX8EbPW5M1QT8mXzesRPwly0HEPTcFlS02Y=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.34/go.mod h1:tGzj56niKYZBbDIRhwPGDqrULzmWv5b6uBQGqyNaFZw=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.14 h1:SA43nfaY7+1jjMNIc2ywu99JLJLButtIdLP6j+bT870=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.14/go.mod h1:Du3llKcwbQvHsTXSLzTOGQz0DTDBMEzdg7DAGu7inrY=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23 h1:9Fjh6fi/U5JEStVZijmaMpUwE/gvBJj7x2B/PjbO9To=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23/go.mod h1:iMoT2f1tClxrWAAnKCXjZQ6LOmfLrMG14wmnWpM+F14=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35 h1:BBEElKh4a+rKshvjrfpajTe9CbpZvrbb4Jkg2PB7RzA=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35/go.mod h1:zaZk983w//8beSruBVec/mr4CmDwgZitW/qzGhAAX0g=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.33 h1:mqI7OrxN/DUH85F5OqVn3cIfuZ3+HVcebUm2N8mLlgQ=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.33/go.mod h1:eZ5jdEpvaaOU8nWWE4cTAJETSEA5FZoWxvNRao4piHY=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31 h1:uao4A3QZ5UmB326V6KF+qRpv9Tjz7IlnlnTbbANntlU=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31/go.mod h1:I/1+z0VwL1GhQyLgkoHDlygpUZ+iTAwOQ/NsftiUL2I=
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2 h1:5C00eQYpTrgQXnp6V3P6P7zPElna3AXvlukbANE6nJI=
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2/go.mod h1:zdmCoFO/dSI7GlrwsPqFJI+WlFnSU4Tc8TJnlXrM1Do=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4 h1:cOJELVNrq5Q3Udry2GLuHUM7MhwpeaQRdYaoa6GI/yI=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4/go.mod h1:f4LxzKBtaTxD7xh3PiVg3CE1tchQemfmghaJr+NbK2c=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.2 h1:EjI1CZzDcBxPkTa3j1BdtIrUDbqnOGssFMeyUS+6W0I=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.2/go.mod h1:vN3eb5H8MEAZ4dx0F5Wc9LT8eb3eW7bZZ5BjGJdbw9k=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 h1:p8WdWDh5AwSZdp19Haa3XMyPCICi9Z375a/Nu3IIEZY=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14/go.mod h1:NKVY7DER6VXHkt2I/ycmHakALNboi3Rqwt4eEf/1Cnk=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 h1:JP2wjWGmUp8lTCZb13Dv0Eciyc1jbO8pd0HZVMHFlrc=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24/go.mod h1:Ql9ziDutk8ERAN9HMaYANCW3lop451ppebkxEJMLCTM=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 h1:AMW7a7S8iQaHjBYZdU3PCq4GKRPijTPRAc7e6XtEThY=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4/go.mod h1:QQNsFV1DVXoXcZt18FS8lI8rtUrlDyAuWZLQ5shunv4=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 h1:AsbZcJAQPRmHDJG8K1N0pof/1zPWjVT8TFlTWuGLSvo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4/go.mod h1:6imqztH0//t0mKbl6yWl7swSEl7F/w32oAmqB3vP1ag=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4 h1:w/AryDYMjSUANSQ2uoZxJovUsMTwWJNTv3IMex30Y+4=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4/go.mod h1:WeBiAa67azG7Su9Vf+ChGDBLiAozJCXzdjXiPBUwtbc=
github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ=
github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.2 h1:zMP1FDFE08L7sM5f1QqkH/ZgKKg8Uc0Dz7KhSSYqWkw=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.2/go.mod h1:0LoIZSUKjdo2BleHfT1hv/jlD33LQS00IrBlzoUsoUQ=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.2 h1:9eTqUYl+SyVmaRPMyBXSO9wwqC6TRwZB82pKENK2hdQ=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.2/go.mod h1:DThweuz22kiLc7lGHop5vQ9c3bx5W6Azs/YqSHa2fu8=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.2 h1:EJd8vZO3E8SE6nmPqxuxlQ1NeSb8as50sf6eGdV4Saw=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.2/go.mod h1:OgpPvKzsO2Ranjpli/20djMkg6UrV5mw4W3pZpq1Mqo=
github.com/aws/smithy-go v1.27.5 h1:d1ro7KpYOYwP6m73YFa+Kc/A130VsAdX68SpsJwARMM=
github.com/aws/smithy-go v1.27.5/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
github.com/bahlo/generic-list-go v0.2.0/go.mod h1:2KvAjgMlE5NNynlg/5iLrrCCZ2+5xWbdbCW3pNTGyYg=
github.com/bazelbuild/rules_go v0.46.0 h1:CTefzjN/D3Cdn3rkrM6qMWuQj59OBcuOjyIp3m4hZ7s=
@@ -848,8 +848,8 @@ github.com/cncf/xds/go v0.0.0-20260202195803-dba9d589def2 h1:aBangftG7EVZoUb69Os
github.com/cncf/xds/go v0.0.0-20260202195803-dba9d589def2/go.mod h1:qwXFYgsP6T7XnJtbKlf1HP8AjxZZyzxMmc+Lq5GjlU4=
github.com/cockroachdb/apd/v3 v3.2.1 h1:U+8j7t0axsIgvQUqthuNm82HIrYXodOV2iWLWtEaIwg=
github.com/cockroachdb/apd/v3 v3.2.1/go.mod h1:klXJcjp+FffLTHlhIG69tezTDvdP065naDsHzKhYSqc=
github.com/cockroachdb/errors v1.14.0 h1:EfdVEJpN3z8rPMo43Yit59LxoiIa470fSXpZXuEs+ZI=
github.com/cockroachdb/errors v1.14.0/go.mod h1:xRa70jZ9sNBQmISt5KmJmAD++E4dQHm89oCRiZGEdq0=
github.com/cockroachdb/errors v1.11.3 h1:5bA+k2Y6r+oz/6Z/RFlNeVCesGARKuC6YymtcDrbC/I=
github.com/cockroachdb/errors v1.11.3/go.mod h1:m4UIW4CDjx+R5cybPsNrRbreomiFqt8o1h1wUVazSd8=
github.com/cockroachdb/logtags v0.0.0-20241215232642-bb51bb14a506 h1:ASDL+UJcILMqgNeV5jiqR4j+sTuvQNHdf2chuKj1M5k=
github.com/cockroachdb/logtags v0.0.0-20241215232642-bb51bb14a506/go.mod h1:Mw7HqKr2kdtu6aYGn3tPmAftiP3QPX63LdK/zcariIo=
github.com/cockroachdb/redact v1.1.5 h1:u1PMllDkdFfPWaNGMyLD1+so+aq3uUItthCFqzwPJ30=
@@ -885,8 +885,8 @@ github.com/containerd/typeurl/v2 v2.2.3 h1:yNA/94zxWdvYACdYO8zofhrTVuQY73fFU1y++
github.com/containerd/typeurl/v2 v2.2.3/go.mod h1:95ljDnPfD3bAbDJRugOiShd/DlAAsxGtUBhJxIn7SCk=
github.com/coreos/go-semver v0.3.1 h1:yi21YpKnrx1gt5R+la8n5WgS0kCrsPp33dmEyHReZr4=
github.com/coreos/go-semver v0.3.1/go.mod h1:irMmmIw/7yzSRPWryHsK7EYSg09caPQL03VsM8rvUec=
github.com/coreos/go-systemd/v22 v22.7.0 h1:LAEzFkke61DFROc7zNLX/WA2i5J8gYqe0rSj9KI28KA=
github.com/coreos/go-systemd/v22 v22.7.0/go.mod h1:xNUYtjHu2EDXbsxz1i41wouACIwT7Ybq9o0BQhMwD0w=
github.com/coreos/go-systemd/v22 v22.6.0 h1:aGVa/v8B7hpb0TKl0MWoAavPDmHvobFe5R5zn0bCJWo=
github.com/coreos/go-systemd/v22 v22.6.0/go.mod h1:iG+pp635Fo7ZmV/j14KUcmEyWF+0X7Lua8rrTWzYgWU=
github.com/cosmos/go-bip39 v1.0.0 h1:pcomnQdrdH22njcAatO0yWojsUnCO3y2tNoV1cb6hHY=
github.com/cosmos/go-bip39 v1.0.0/go.mod h1:RNJv0H/pOIVgxw6KS7QeX2a0Uo0aKUlfhZ4xuwvCdJw=
github.com/cpuguy83/dockercfg v0.3.2 h1:DlJTyZGBDlXqUZ2Dk2Q3xHs/FtnooJJVaad2S9GKorA=
@@ -956,8 +956,8 @@ github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 h1:Oy0F4A
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3/go.mod h1:YvSRo5mw33fLEx1+DlK6L2VV43tJt5Eyel9n9XBcR+0=
github.com/eapache/queue v1.1.0 h1:YOEu7KNc61ntiQlcEeUIoDTJ2o8mQznoNvUhiigpIqc=
github.com/eapache/queue v1.1.0/go.mod h1:6eCeP0CKFpHLu8blIFXhExK/dRa7WDZfr6jVFPTqq+I=
github.com/ebitengine/purego v0.10.2 h1:W809HbnvzAxgdm+aOvlSekrM16wGCdT/e76+9tS7gzE=
github.com/ebitengine/purego v0.10.2/go.mod h1:iIjxzd6CiRiOG0UyXP+V1+jWqUXVjPKLAI0mRfJZTmQ=
github.com/ebitengine/purego v0.10.1 h1:dewVBCBT2GaMu1SrNTYxQhgQBethzfhiwvZiLGP/qyY=
github.com/ebitengine/purego v0.10.1/go.mod h1:iIjxzd6CiRiOG0UyXP+V1+jWqUXVjPKLAI0mRfJZTmQ=
github.com/eiannone/keyboard v0.0.0-20220611211555-0d226195f203 h1:XBBHcIb256gUJtLmY22n99HaZTz+r2Z51xUPi01m3wg=
github.com/eiannone/keyboard v0.0.0-20220611211555-0d226195f203/go.mod h1:E1jcSv8FaEny+OP/5k9UxZVw9YFWGj7eI4KR/iOBqCg=
github.com/elastic/gosigar v0.14.3 h1:xwkKwPia+hSfg9GqrCUKYdId102m9qTJIIr7egmK/uo=
@@ -1031,8 +1031,8 @@ github.com/gabriel-vasile/mimetype v1.4.13 h1:46nXokslUBsAJE/wMsp5gtO500a4F3Nkz9
github.com/gabriel-vasile/mimetype v1.4.13/go.mod h1:d+9Oxyo1wTzWdyVUPMmXFvp4F9tea18J8ufA774AB3s=
github.com/geoffgarside/ber v1.2.0 h1:/loowoRcs/MWLYmGX9QtIAbA+V/FrnVLsMMPhwiRm64=
github.com/geoffgarside/ber v1.2.0/go.mod h1:jVPKeCbj6MvQZhwLYsGwaGI52oUorHoHKNecGT85ZCc=
github.com/getsentry/sentry-go v0.48.0 h1:FRZNr7Uk1C86ev1bSJmYlUkL9oyivQA6YOcdYfaaMmY=
github.com/getsentry/sentry-go v0.48.0/go.mod h1:E5UkA5wp1qR2+MDydNYlVeUiNN2xEdjYMidkgf0Qoss=
github.com/getsentry/sentry-go v0.44.1 h1:/cPtrA5qB7uMRrhgSn9TYtcEF36auGP3Y6+ThvD/yaI=
github.com/getsentry/sentry-go v0.44.1/go.mod h1:XDotiNZbgf5U8bPDUAfvcFmOnMQQceESxyKaObSssW0=
github.com/ghodss/yaml v1.0.0/go.mod h1:4dBDuWmgqj2HViK6kFavaiC9ZROes6MMH2rRYeMEF04=
github.com/gin-contrib/sse v1.1.0 h1:n0w2GMuUpWDVp7qSpvze6fAu9iRxJY4Hmj6AmBOU05w=
github.com/gin-contrib/sse v1.1.0/go.mod h1:hxRZ5gVpWMT7Z0B0gSNYqqsSCNIJMjzvm6fqCz9vjwM=
@@ -1236,8 +1236,8 @@ github.com/google/pprof v0.0.0-20210226084205-cbba55b83ad5/go.mod h1:kpwsk12EmLe
github.com/google/pprof v0.0.0-20210601050228-01bbb1931b22/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20210609004039-a478d1d731e9/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20210720184732-4bb14d4b1be1/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3 h1:LMLX+LgTNWpfvCBdFebv6EsYotImrt/Ppc5cXIriCSo=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3/go.mod h1:jl5iWTm0/hd5PjEYEOuwAJ57L/CibdZfrqZ5XA5GrCk=
github.com/google/pprof v0.0.0-20250317173921-a4b03ec1a45e h1:ijClszYn+mADRFY17kjQEVQ1XRhq2/JR1M3sGqeJoxs=
github.com/google/pprof v0.0.0-20250317173921-a4b03ec1a45e/go.mod h1:boTsfXsheKC2y+lKOCMpSfarhxDeIzfZG1jqGcPl3cA=
github.com/google/renameio v0.1.0/go.mod h1:KWCgfxg9yswjAJkECMjeO8J8rahYeXnNhOm40UhjYkI=
github.com/google/s2a-go v0.1.9 h1:LGD7gtMgezd8a/Xak7mEWL0PjoTQFvpRudN895yqKW0=
github.com/google/s2a-go v0.1.9/go.mod h1:YA0Ei2ZQL3acow2O62kdp9UlnvMmU7kA6Eutn0dXayM=
@@ -1255,8 +1255,8 @@ github.com/googleapis/enterprise-certificate-proxy v0.1.0/go.mod h1:17drOmN3MwGY
github.com/googleapis/enterprise-certificate-proxy v0.2.0/go.mod h1:8C0jb7/mgJe/9KK8Lm7X9ctZC2t60YyIpYEI16jx0Qg=
github.com/googleapis/enterprise-certificate-proxy v0.2.1/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.2.3/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.3.20 h1:t/xL64VUoN69MuMRQuJETqYGOw4Z9mSRJK9epIEtwFk=
github.com/googleapis/enterprise-certificate-proxy v0.3.20/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/enterprise-certificate-proxy v0.3.18 h1:hvVi34VucdrV1IIsiWuqYM8kutw/92MxNEFxCJZEh0k=
github.com/googleapis/enterprise-certificate-proxy v0.3.18/go.mod h1:rSEsBUemEBZEexP2y6jPp16LUmUbjmSbcPMQizR0o4k=
github.com/googleapis/gax-go/v2 v2.0.4/go.mod h1:0Wqv26UfaUD9n4G6kQubkQ+KchISgw+vpHVxEJEs9eg=
github.com/googleapis/gax-go/v2 v2.0.5/go.mod h1:DWXyrwAJ9X0FpwwEdw+IPEYBICEFu5mhpdKc/us6bOk=
github.com/googleapis/gax-go/v2 v2.1.0/go.mod h1:Q3nei7sK6ybPYH7twZdmQpAd1MKb7pfu6SK+H1/DsU0=
@@ -1267,8 +1267,8 @@ github.com/googleapis/gax-go/v2 v2.4.0/go.mod h1:XOTVJ59hdnfJLIP/dh8n5CGryZR2LxK
github.com/googleapis/gax-go/v2 v2.5.1/go.mod h1:h6B0KMMFNtI2ddbGJn3T3ZbwkeT6yqEF02fYlzkUCyo=
github.com/googleapis/gax-go/v2 v2.6.0/go.mod h1:1mjbznJAPHFpesgE5ucqfYEscaz5kMdcIDwU/6+DDoY=
github.com/googleapis/gax-go/v2 v2.7.0/go.mod h1:TEop28CZZQ2y+c0VxMUmu1lV+fQx57QpBWsYpwqHJx8=
github.com/googleapis/gax-go/v2 v2.24.0 h1:myMaPYyF9MecEmvQqMqomIwn9t/4KCZN9qnwsS76wlg=
github.com/googleapis/gax-go/v2 v2.24.0/go.mod h1:IaTHBDd7NHxSCiu0vEs8pQZu4dGZrWwuSoxCnk16OFM=
github.com/googleapis/gax-go/v2 v2.23.0 h1:Tchl7qkvE7Ip3y+ztvNufYFvkfqTe7NfLTYGIdJRLuE=
github.com/googleapis/gax-go/v2 v2.23.0/go.mod h1:rBQKOVJCdb8IFEzg+FCwlt1LP/xMDGuqUXhUG+XMXEg=
github.com/googleapis/go-type-adapters v1.0.0/go.mod h1:zHW75FOG2aur7gAO2B+MLby+cLsWGBF62rFAi7WjWO4=
github.com/googleapis/google-cloud-go-testing v0.0.0-20200911160855-bcd43fbb19e8/go.mod h1:dvDLG8qkwmyD9a/MJJN3XJcT3xFxOKAvTZGvuZmac9g=
github.com/gookit/assert v0.1.1 h1:lh3GcawXe/p+cU7ESTZ5Ui3Sm/x8JWpIis4/1aF0mY0=
@@ -1295,8 +1295,8 @@ github.com/grpc-ecosystem/grpc-gateway v1.16.0 h1:gmcG1KaJ57LophUzW0Hy8NmPhnMZb4
github.com/grpc-ecosystem/grpc-gateway v1.16.0/go.mod h1:BDjrQk3hbvj6Nolgz8mAMFbcEtjT1g+wF4CSlocrBnw=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.7.0/go.mod h1:hgWBS7lorOAVIJEQMi4ZsPv9hVvWI6+ch50m39Pf2Ks=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.11.3/go.mod h1:o//XUCC/F+yRGJoPO/VU0GSB0f8Nhgmxx0VIRUvaC0w=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 h1:5VipnvEpbqr2gA2VbM+nYVbkIF28c5ZQfqCBQ5g2xfk=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0/go.mod h1:Hyl3n6Twe1hvtd9XUXDec4pTvgMSEixRuQKPTMH2bNs=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.28.0 h1:HWRh5R2+9EifMyIHV7ZV+MIZqgz+PMpZ14Jynv3O2Zs=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.28.0/go.mod h1:JfhWUomR1baixubs02l85lZYYOm7LV6om4ceouMv45c=
github.com/hashicorp/errwrap v1.0.0/go.mod h1:YH+1FKiLXxHSkmPseP+kNlulaMuP3n2brvKWEqk/Jc4=
github.com/hashicorp/errwrap v1.1.0 h1:OxrOeh75EUXMY8TBjag2fzXGZ40LB6IKw45YeGUDY2I=
github.com/hashicorp/errwrap v1.1.0/go.mod h1:YH+1FKiLXxHSkmPseP+kNlulaMuP3n2brvKWEqk/Jc4=
@@ -1511,8 +1511,8 @@ github.com/mattn/go-colorable v0.1.15/go.mod h1:6LmQG8QLFO4G5z1gPvYEzlUgJ2wF+stg
github.com/mattn/go-isatty v0.0.12/go.mod h1:cbi8OIDigv2wuxKPP5vlRcQ1OAZbq2CE4Kysco4FUpU=
github.com/mattn/go-isatty v0.0.14/go.mod h1:7GGIvUiUoEMVVmxf/4nioHXj79iQHKdU27kJ6hsGG94=
github.com/mattn/go-isatty v0.0.16/go.mod h1:kYGgaQfpe5nmfYZH+SKPsOc2e4SrIfOl2e/yFXSvRLM=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/mattn/go-isatty v0.0.23 h1:cYwCQTQf3HB6xUC+BtyCLZNr7IzbOmoZbmssVNzSyiQ=
github.com/mattn/go-isatty v0.0.23/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/mattn/go-runewidth v0.0.3/go.mod h1:LwmH8dsx7+W8Uxz3IHJYH5QSwggIsqBzpuz5H//U1FU=
github.com/mattn/go-runewidth v0.0.13/go.mod h1:Jdepj2loyihRzMpdS35Xk/zdY8IAYHsh153qUoGf23w=
github.com/mattn/go-runewidth v0.0.24 h1:cpokDiIn0MGnhdHwuWnJBITySJ20QyNGnY2kR/ay2DU=
@@ -1543,8 +1543,8 @@ github.com/moby/buildkit v0.29.0 h1:wxLEFbCOJntEDjSNNN2YWd8zxltZxT5muDQ0LzpbtpU=
github.com/moby/buildkit v0.29.0/go.mod h1:Dmv2FeDe34t75QuzeU87rBoZpAAkcpT5zeu4hXzmASc=
github.com/moby/docker-image-spec v1.3.1 h1:jMKff3w6PgbfSa69GfNg+zN/XLhfXJGnEx3Nl2EsFP0=
github.com/moby/docker-image-spec v1.3.1/go.mod h1:eKmb5VW8vQEh/BAr2yvVNvuiJuY6UIocYsFu/DxxRpo=
github.com/moby/go-archive v0.3.0 h1:nos4BtzzUIqB406BgQnWGMI4qib9BZ8XUHU+ucv/n1c=
github.com/moby/go-archive v0.3.0/go.mod h1:Npdv43fFqlhZW7Xo8fbm3ZMYFvAGNviUPqX21VERbcE=
github.com/moby/go-archive v0.2.0 h1:zg5QDUM2mi0JIM9fdQZWC7U8+2ZfixfTYoHL7rWUcP8=
github.com/moby/go-archive v0.2.0/go.mod h1:mNeivT14o8xU+5q1YnNrkQVpK+dnNe/K6fHqnTg4qPU=
github.com/moby/locker v1.0.1 h1:fOXqR41zeveg4fFODix+1Ch4mj/gT0NE1XJbp/epuBg=
github.com/moby/locker v1.0.1/go.mod h1:S7SDdo5zpBK84bzzVlKr2V0hz+7x9hWbYC/kq7oQppc=
github.com/moby/moby/api v1.54.2 h1:wiat9QAhnDQjA7wk1kh/TqHz2I1uUA7M7t9SAl/JNXg=
@@ -1559,14 +1559,14 @@ github.com/moby/sys/capability v0.4.0 h1:4D4mI6KlNtWMCM1Z/K0i7RV1FkX+DBDHKVJpCnd
github.com/moby/sys/capability v0.4.0/go.mod h1:4g9IK291rVkms3LKCDOoYlnV8xKwoDTpIrNEE35Wq0I=
github.com/moby/sys/mountinfo v0.7.2 h1:1shs6aH5s4o5H2zQLn796ADW1wMrIwHsyJ2v9KouLrg=
github.com/moby/sys/mountinfo v0.7.2/go.mod h1:1YOa8w8Ih7uW0wALDUgT1dTTSBrZ+HiBLGws92L2RU4=
github.com/moby/sys/sequential v0.7.0 h1:ASQNGNROJSuOO6LL6bPHbKvuZu6NU8P4ldPWk31zj/8=
github.com/moby/sys/sequential v0.7.0/go.mod h1:NfSTAp6V3fw4tmkD62PEcOKeZKquXT8VKCkf7aVR79o=
github.com/moby/sys/sequential v0.6.0 h1:qrx7XFUd/5DxtqcoH1h438hF5TmOvzC/lspjy7zgvCU=
github.com/moby/sys/sequential v0.6.0/go.mod h1:uyv8EUTrca5PnDsdMGXhZe6CCe8U/UiTWd+lL+7b/Ko=
github.com/moby/sys/signal v0.7.1 h1:PrQxdvxcGijdo6UXXo/lU/TvHUWyPhj7UOpSo8tuvk0=
github.com/moby/sys/signal v0.7.1/go.mod h1:Se1VGehYokAkrSQwL4tDzHvETwUZlnY7S5XtQ50mQp8=
github.com/moby/sys/symlink v0.3.0 h1:GZX89mEZ9u53f97npBy4Rc3vJKj7JBDj/PN2I22GrNU=
github.com/moby/sys/symlink v0.3.0/go.mod h1:3eNdhduHmYPcgsJtZXW1W4XUJdZGBIkttZ8xKqPUJq0=
github.com/moby/sys/user v0.4.1 h1:RgjRlaDKi/Xmyrz4t8lyzXT6v2ooFeO/7xtchmhVWE0=
github.com/moby/sys/user v0.4.1/go.mod h1:E9QsW5WRe1kUAf7kW8hXKwu1uhsZEAdPLYHYSDudF4Y=
github.com/moby/sys/user v0.4.0 h1:jhcMKit7SA80hivmFJcbB1vqmw//wU61Zdui2eQXuMs=
github.com/moby/sys/user v0.4.0/go.mod h1:bG+tYYYJgaMtRKgEmuueC0hJEAZWwtIbZTB+85uoHjs=
github.com/moby/sys/userns v0.1.0 h1:tVLXkFOxVu9A64/yh59slHVv9ahO9UIev4JZusOLG/g=
github.com/moby/sys/userns v0.1.0/go.mod h1:IHUYgu/kao6N8YZlp9Cf444ySSvCmDlmzUcYfDHOl28=
github.com/moby/term v0.5.2 h1:6qk3FJAFDs6i/q3W/pQ97SX192qKfZgGjCQqfCJkgzQ=
@@ -1636,8 +1636,8 @@ github.com/parquet-go/bitpack v1.0.0 h1:AUqzlKzPPXf2bCdjfj4sTeacrUwsT7NlcYDMUQxP
github.com/parquet-go/bitpack v1.0.0/go.mod h1:XnVk9TH+O40eOOmvpAVZ7K2ocQFrQwysLMnc6M/8lgs=
github.com/parquet-go/jsonlite v1.0.0 h1:87QNdi56wOfsE5bdgas0vRzHPxfJgzrXGml1zZdd7VU=
github.com/parquet-go/jsonlite v1.0.0/go.mod h1:nDjpkpL4EOtqs6NQugUsi0Rleq9sW/OtC1NnZEnxzF0=
github.com/parquet-go/parquet-go v0.32.0 h1:NWDqTUHfrCS4cJP/Fj2HlxvqsrVedWG3sayMkf+znzM=
github.com/parquet-go/parquet-go v0.32.0/go.mod h1:navtkAYr2LGoJVp141oXPlO/sxLvaOe3la2JEoD8+rg=
github.com/parquet-go/parquet-go v0.30.1 h1:Oy6ganNrAdFiVwy7wNmWagfPTWA2X9Z3tVHBc7JtuX8=
github.com/parquet-go/parquet-go v0.30.1/go.mod h1:navtkAYr2LGoJVp141oXPlO/sxLvaOe3la2JEoD8+rg=
github.com/pascaldekloe/goe v0.1.0 h1:cBOtyMzM9HTpWjXfbbunk26uA6nG3a8n06Wieeh0MwY=
github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144TG7ZOy1lc=
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
@@ -1659,8 +1659,8 @@ github.com/phpdave11/gofpdf v1.4.2/go.mod h1:zpO6xFn9yxo3YLyMvW8HcKWVdbNqgIfOOp2
github.com/phpdave11/gofpdi v1.0.12/go.mod h1:vBmVV0Do6hSBHC8uKUQ71JGW+ZGQq74llk/7bXwjDoI=
github.com/phpdave11/gofpdi v1.0.13/go.mod h1:vBmVV0Do6hSBHC8uKUQ71JGW+ZGQq74llk/7bXwjDoI=
github.com/pierrec/lz4/v4 v4.1.15/go.mod h1:gZWDp/Ze/IJXGXf23ltt2EXimqmTUXEy0GFuRQyBid4=
github.com/pierrec/lz4/v4 v4.1.28 h1:pPEPwRJ4kybBTfGt28q7lQsRJQHhC08axprdLD5Ppio=
github.com/pierrec/lz4/v4 v4.1.28/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pierrec/lz4/v4 v4.1.27 h1:+PhzhWDrjRj89TH2sw43nE3+4+W8lSxIuQadEHZyjUk=
github.com/pierrec/lz4/v4 v4.1.27/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pierrre/compare v1.0.2 h1:k4IUsHgh+dbcAOIWCfxVa/7G6STjADH2qmhomv+1quc=
github.com/pierrre/compare v1.0.2/go.mod h1:8UvyRHH+9HS8Pczdd2z5x/wvv67krDwVxoOndaIIDVU=
github.com/pierrre/geohash v1.0.0 h1:f/zfjdV4rVofTCz1FhP07T+EMQAvcMM2ioGZVt+zqjI=
@@ -1749,8 +1749,8 @@ github.com/quic-go/qpack v0.6.0 h1:g7W+BMYynC1LbYLSqRt8PBg5Tgwxn214ZZR34VIOjz8=
github.com/quic-go/qpack v0.6.0/go.mod h1:lUpLKChi8njB4ty2bFLX2x4gzDqXwUpaO1DP9qMDZII=
github.com/quic-go/quic-go v0.59.0 h1:OLJkp1Mlm/aS7dpKgTc6cnpynnD2Xg7C1pwL6vy/SAw=
github.com/quic-go/quic-go v0.59.0/go.mod h1:upnsH4Ju1YkqpLXC305eW3yDZ4NfnNbmQRCMWS58IKU=
github.com/rabbitmq/amqp091-go v1.14.0 h1:RSaT7aOKt/OrkVUyswPDW29lnRz9psuGmfZFBmLqLek=
github.com/rabbitmq/amqp091-go v1.14.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rabbitmq/amqp091-go v1.11.0 h1:HxIctVm9Gid/Vtn706necmZ7Wj6pgGI2eqplRbEY8O8=
github.com/rabbitmq/amqp091-go v1.11.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rclone/Proton-API-Bridge v1.0.4 h1:uGQJRjQC1hVLd5kqLsXc6CWO6oqrVeLoKQYoHapEZDg=
github.com/rclone/Proton-API-Bridge v1.0.4/go.mod h1:VTPBYZotKAeDLlAzxU2O/s14NXk9FxUt9hn1jhH2iY8=
github.com/rclone/go-proton-api v1.0.3 h1:3gBTzR+j0dYiTwtj9yKIdN/aV3W2a8KIPKp0GArojyQ=
@@ -1790,8 +1790,6 @@ github.com/rs/zerolog v1.34.0 h1:k43nTLIwcTVQAncfCw4KZ2VY6ukYoZaBPNOE8txlOeY=
github.com/rs/zerolog v1.34.0/go.mod h1:bJsvje4Z08ROH4Nhs5iH600c3IkWhwp44iRc54W6wYQ=
github.com/ruudk/golang-pdf417 v0.0.0-20181029194003-1af4ab5afa58/go.mod h1:6lfFZQK844Gfx8o5WFuvpxWRwnSoipWe/p622j1v06w=
github.com/ruudk/golang-pdf417 v0.0.0-20201230142125-a7e3863a1245/go.mod h1:pQAZKsJ8yyVxGRWYNEm9oFB8ieLgKFnamEyDmSA0BRk=
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd h1:CmH9+J6ZSsIjUK3dcGsnCnO41eRBOnY12zwkn5qVwgc=
github.com/rwcarlsen/goexif v0.0.0-20190401172101-9e8deecbddbd/go.mod h1:hPqNNc0+uJM6H+SuU8sEs5K5IQeKccPqeSjfgcKGgPk=
github.com/ryanuber/go-glob v1.0.0 h1:iQh3xXAumdQ+4Ufa5b25cRpC5TYKlno6hsv6Cb3pkBk=
github.com/ryanuber/go-glob v1.0.0/go.mod h1:807d1WSdnB0XRJzKNil9Om6lcp/3a0v4qIHxIXzX/Yc=
github.com/sabhiram/go-gitignore v0.0.0-20210923224102-525f6e181f06 h1:OkMGxebDjyw0ULyrTYWeN0UNCCkmCWfjPnIA2W6oviI=
@@ -1810,8 +1808,8 @@ github.com/seaweedfs/cockroachdb-parser v0.0.0-20260225204133-2f342c5ea564 h1:Tg
github.com/seaweedfs/cockroachdb-parser v0.0.0-20260225204133-2f342c5ea564/go.mod h1:JSKCh6uCHBz91lQYFYHCyTrSVIPge4SUFVn28iwMNB0=
github.com/seaweedfs/go-fuse/v2 v2.9.4 h1:ACyloiuopdhRSjdLLeSWbsVaemMPskORaRF01TY6GyM=
github.com/seaweedfs/go-fuse/v2 v2.9.4/go.mod h1:zABdmWEa6A0bwaBeEOBUeUkGIZlxUhcdv+V1Dcc/U/I=
github.com/seaweedfs/goexif v2.0.0+incompatible h1:x8pckiT12QQhifwhDQpeISgDfsqmQ6VR4LFPQ64JRps=
github.com/seaweedfs/goexif v2.0.0+incompatible/go.mod h1:Oni780Z236sXpIQzk1XoJlTwqrJ02smEin9zQeff7Fk=
github.com/seaweedfs/goexif v1.0.3 h1:ve/OjI7dxPW8X9YQsv3JuVMaxEyF9Rvfd04ouL+Bz30=
github.com/seaweedfs/goexif v1.0.3/go.mod h1:Oni780Z236sXpIQzk1XoJlTwqrJ02smEin9zQeff7Fk=
github.com/seaweedfs/raft v1.2.0 h1:Ez4Hw9ifBbTT7wg54DvGHBjw1vRlTb4roH0TKl0Oj9Y=
github.com/seaweedfs/raft v1.2.0/go.mod h1:fgs/rAVEzjQ7e04XMzG3eJhwZZRmBW+2uRtjakeCGeU=
github.com/secure-systems-lab/go-securesystemslib v0.10.0 h1:l+H5ErcW0PAehBNrBxoGv1jjNpGYdZ9RcheFkB2WI14=
@@ -1822,8 +1820,8 @@ github.com/sergi/go-diff v1.2.0 h1:XU+rvMAioB0UC3q1MFrIQy4Vo5/4VsRDQQXHsEya6xQ=
github.com/sergi/go-diff v1.2.0/go.mod h1:STckp+ISIX8hZLjrqAeVduY0gWCT9IjLuqbuNXdaHfM=
github.com/shibumi/go-pathspec v1.3.0 h1:QUyMZhFo0Md5B8zV8x2tesohbb5kfbpTi9rBnKh5dkI=
github.com/shibumi/go-pathspec v1.3.0/go.mod h1:Xutfslp817l2I1cZvgcfeMQJG5QnU2lh5tVaaMCl3jE=
github.com/shirou/gopsutil/v4 v4.26.7 h1:IXzpHz/dkMRYAhKkOXr1HB6SuzWU3eoyyeWe7g3bNZc=
github.com/shirou/gopsutil/v4 v4.26.7/go.mod h1:5O9FjBiXoTDFatIWjZZosqj4pV0DRtLx598xGbBehzM=
github.com/shirou/gopsutil/v4 v4.26.6 h1:Mzr/npDtQC/xpeEuQKHZt8Zo9CmPvhTj8nkR8w5TLDs=
github.com/shirou/gopsutil/v4 v4.26.6/go.mod h1:LZ6ewCSkBqUpvSOf+LsTGnRinC6iaNUNMGBtDkJBaLQ=
github.com/sigstore/sigstore v1.10.4 h1:ytOmxMgLdcUed3w1SbbZOgcxqwMG61lh1TmZLN+WeZE=
github.com/sigstore/sigstore v1.10.4/go.mod h1:tDiyrdOref3q6qJxm2G+JHghqfmvifB7hw+EReAfnbI=
github.com/sigstore/sigstore-go v1.1.4 h1:wTTsgCHOfqiEzVyBYA6mDczGtBkN7cM8mPpjJj5QvMg=
@@ -1909,8 +1907,8 @@ github.com/tarantool/go-iproto v1.1.0 h1:HULVOIHsiehI+FnHfM7wMDntuzUddO09DKqu2Wn
github.com/tarantool/go-iproto v1.1.0/go.mod h1:LNCtdyZxojUed8SbOiYHoc3v9NvaZTB7p96hUySMlIo=
github.com/tarantool/go-option v1.1.0 h1:ShoOhNsdL41sRpm4hXCRDjV8H0WzPkd4UnKhLKbW//w=
github.com/tarantool/go-option v1.1.0/go.mod h1:hMr9z2JXOWlgdCBpCPSL2nwp8718GKYvNBJ+ZuzJbCo=
github.com/tarantool/go-tarantool/v3 v3.0.1 h1:vaUX4xmVmXh2dIJ/LqlX1MXK3iYqAqV6YiE54Wwl/qg=
github.com/tarantool/go-tarantool/v3 v3.0.1/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/tarantool/go-tarantool/v3 v3.0.0 h1:zsIXS4nvSSXqZWtN/f1Bfuq973j6J85O5U/8RYQCRcE=
github.com/tarantool/go-tarantool/v3 v3.0.0/go.mod h1:TXxLWhUCgdxXFfelnTSkq+goKRTRj660zxq4/WXPe8k=
github.com/testcontainers/testcontainers-go v0.43.0 h1:oEQx5MW2DGd9z3AeEQfB2lPM0eLs7ztyaGRu75bFo5A=
github.com/testcontainers/testcontainers-go v0.43.0/go.mod h1:+VxkT2NQnKOZPKi6praMuMKYHYyOGXr0XSBSlSMCzFo=
github.com/testcontainers/testcontainers-go/modules/compose v0.42.0 h1:+t1ZN31TD36cwxmeLqGwe7wIdvblBm0Z+vlj4SX8Mv0=
@@ -2032,14 +2030,14 @@ github.com/yandex-cloud/go-genproto v0.0.0-20211115083454-9ca41db5ed9e h1:9LPdmD
github.com/yandex-cloud/go-genproto v0.0.0-20211115083454-9ca41db5ed9e/go.mod h1:HEUYX/p8966tMUHHT+TsS0hF/Ca/NYwqprC5WXSDMfE=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20221215182650-986f9d10542f/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20230528143953-42c825ace222/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260810122915-65bfd5c4b705 h1:7VKlOrBIQ8L8acJ9wFCt6lWzLDbOhc05VLpL31743tU=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260810122915-65bfd5c4b705/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260428144813-1c07baab7f7b h1:xeiobG1riqe6dTMZuTcOvwOY0BbFidSk3gRLyVeLfJA=
github.com/ydb-platform/ydb-go-genproto v0.0.0-20260428144813-1c07baab7f7b/go.mod h1:Er+FePu1dNUieD+XTMDduGpQuCPssK5Q4BjF+IIXJ3I=
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2 h1:e2nGQPGC5OEPBWlMnLPFpdgyqNA8iTOPiZShCrv6944=
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2/go.mod h1:9YzkhlIymWaJGX6KMU3vh5sOf3UKbCXkG/ZdjaI3zNM=
github.com/ydb-platform/ydb-go-sdk/v3 v3.44.0/go.mod h1:oSLwnuilwIpaF5bJJMAofnGgzPJusoI3zWMNb8I+GnM=
github.com/ydb-platform/ydb-go-sdk/v3 v3.47.3/go.mod h1:bWnOIcUHd7+Sl7DN+yhyY1H/I61z53GczvwJgXMgvj0=
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1 h1:T+fB2ZDHpYIGC7DWjK+rzZHLoexBF0zG/n3Q9DGNskY=
github.com/ydb-platform/ydb-go-sdk/v3 v3.151.1/go.mod h1:dJXJ1u00IqO8Vsph8fWmYj8b1E7jyphnvJPqA69emBY=
github.com/ydb-platform/ydb-go-sdk/v3 v3.146.3 h1:0ybsLjWJ25M8XXGXk7NfjhO+s8bqBcGGWt8iB2DQiHQ=
github.com/ydb-platform/ydb-go-sdk/v3 v3.146.3/go.mod h1:b9NEO6mgaiqsnOMkS003uS82XsKh6GL+ZTFfPqXWz+c=
github.com/ydb-platform/ydb-go-yc v0.12.1 h1:qw3Fa+T81+Kpu5Io2vYHJOwcrYrVjgJlT6t/0dOXJrA=
github.com/ydb-platform/ydb-go-yc v0.12.1/go.mod h1:t/ZA4ECdgPWjAb4jyDe8AzQZB5dhpGbi3iCahFaNwBY=
github.com/ydb-platform/ydb-go-yc-metadata v0.6.1 h1:9E5q8Nsy2RiJMZDNVy0A3KUrIMBPakJ2VgloeWbcI84=
@@ -2075,12 +2073,12 @@ go.einride.tech/aip v0.83.0 h1:TI21IdeOnLTwZEJ3BxtImIZk6bsN2Q+sd0x99SLiQ+M=
go.einride.tech/aip v0.83.0/go.mod h1:E8+wdTApA70odnpFzJgsGogHozC2JCIhFJBKPr8bVig=
go.etcd.io/bbolt v1.5.0 h1:S7GAl7Fxv12yohbwFfIbQCGDWbQbtDGPET4P/bD4lxU=
go.etcd.io/bbolt v1.5.0/go.mod h1:mkltfYE5aUHQxUct9N9V+Kp7aSjFqjgrhcXIS70Lrdk=
go.etcd.io/etcd/api/v3 v3.7.1 h1:KJG0/DcWGfe3Y1otDf/fsBf0TSSgpxZ5RO/L8SFt73E=
go.etcd.io/etcd/api/v3 v3.7.1/go.mod h1:8bXIpCMeV7E3/XL0Ix123ATn3dB+0V7d9zklHbB0m78=
go.etcd.io/etcd/client/pkg/v3 v3.7.1 h1:rKYsj3pRkR0eK3yjT3XOgrhqfmIfj9pzNgxjh7mfFv4=
go.etcd.io/etcd/client/pkg/v3 v3.7.1/go.mod h1:cnzZGIUzSfjEwLC6UBVsSXlEK1eepS/JUD7wE6PLRT0=
go.etcd.io/etcd/client/v3 v3.7.1 h1:0PEMMC0KuZmVIN+RAbdqfkZ45pYTgKVtmBEbRCvZFUg=
go.etcd.io/etcd/client/v3 v3.7.1/go.mod h1:ffNqALa8tRCYhYo1F9oR489y23K39Gz+BSR3ApAGYq0=
go.etcd.io/etcd/api/v3 v3.6.12 h1:OLOZUKEuAA36TR48F0cIaa8FdzrWygjyfrJxXg4iDgs=
go.etcd.io/etcd/api/v3 v3.6.12/go.mod h1:p14EIQXHbuOQbVvL/WEes5uqKnxP9AgKJgpjbMVvzvE=
go.etcd.io/etcd/client/pkg/v3 v3.6.12 h1:36zzB+pQOdHbhN+kH2iJz/K8bJn0ZLtLfPPO7jozTDo=
go.etcd.io/etcd/client/pkg/v3 v3.6.12/go.mod h1:hh2+ZXtfLzs3o6mn92ntgNPBrTJJOvXqICM5g3L3DMY=
go.etcd.io/etcd/client/v3 v3.6.12 h1:kMSP6JcPZMqSJiX+TXdUIBU/4eXEZWBAaui4VihMbIc=
go.etcd.io/etcd/client/v3 v3.6.12/go.mod h1:CMs6fJWYiZQk4ytFjd4lE1diOvvRMmtbbn/alZXd3dQ=
go.mongodb.org/mongo-driver v1.17.9 h1:IexDdCuuNJ3BHrELgBlyaH9p60JXAvdzWR128q+U5tU=
go.mongodb.org/mongo-driver v1.17.9/go.mod h1:LlOhpH5NUEfhxcAwG0UEkMqwYcc4JU18gtCdGudk/tQ=
go.opencensus.io v0.21.0/go.mod h1:mSImk1erAIZhrmZN+AvHh14ztQfjbGwt4TtuofqLduU=
@@ -2184,8 +2182,8 @@ golang.org/x/crypto v0.6.0/go.mod h1:OFC/31mSvZgRz0V1QTNCzfAI1aIRzbiufJtkMIlEp58
golang.org/x/crypto v0.7.0/go.mod h1:pYwdfH91IfpZVANVyUOhSIPZaFoJGxTFbZhFTx+dXZU=
golang.org/x/crypto v0.13.0/go.mod h1:y6Z2r+Rw4iayiXXAIxJIDAJ1zMW4yaTpebo8fPOliYc=
golang.org/x/crypto v0.14.0/go.mod h1:MVFd36DqK4CsrnJYDkBA3VC4m2GkXAM0PvzMCn4JQf4=
golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M=
golang.org/x/crypto v0.55.0/go.mod h1:uq0V9dE/fzQuJtbnL+2EhWOE63vo164FY8xqEnV9xis=
golang.org/x/crypto v0.54.0 h1:YLIA59K4fiNzHzjnZt2tUJQjQtUWfWbeHBqKtk3eScw=
golang.org/x/crypto v0.54.0/go.mod h1:KWL8ny2AZdGR2cWmzeHrp2azQPGogOv+HeQaVEXC2dk=
golang.org/x/exp v0.0.0-20180321215751-8460e604b9de/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20180807140117-3d87b88a115f/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20190121172915-509febef88a4/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
@@ -2315,8 +2313,8 @@ golang.org/x/net v0.8.0/go.mod h1:QVkue5JL9kW//ek3r6jTKnTFis1tRmNAW2P1shuFdJc=
golang.org/x/net v0.10.0/go.mod h1:0qNGK6F8kojg2nk9dLZ2mShWaEBan6FAoqfSigmmuDg=
golang.org/x/net v0.15.0/go.mod h1:idbUs1IY1+zTqbi8yxTbhexhEEk5ur9LInksu6HrEpk=
golang.org/x/net v0.16.0/go.mod h1:NxSsAGuq816PNPmqtQdLE42eU2Fs7NoRIZrHJAlaCOE=
golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To=
golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU=
golang.org/x/net v0.57.0 h1:K5+3DljvIuDG9/Jv9rvyMywYNFCQ9RSUY6OOTTkT+tE=
golang.org/x/net v0.57.0/go.mod h1:KpXc8iv+r3XplLAG/f7Jsf9RPszJzdR0f58q9vGOuEU=
golang.org/x/oauth2 v0.0.0-20180821212333-d2e6202438be/go.mod h1:N/0e6XlmueqKjAGxoOufVs8QHGRruUQn6yWY3a++T0U=
golang.org/x/oauth2 v0.0.0-20190226205417-e64efc72b421/go.mod h1:gOpvHmFTYa4IltrdGE7lF6nIHvwfUNPOp7c8zoXwtLw=
golang.org/x/oauth2 v0.0.0-20190604053449-0f29369cfe45/go.mod h1:gOpvHmFTYa4IltrdGE7lF6nIHvwfUNPOp7c8zoXwtLw=
@@ -2502,8 +2500,8 @@ golang.org/x/text v0.8.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
@@ -2659,8 +2657,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.294.0 h1:8gASjJxdtcIieB3OqbkLcF0FfbXVNqKtU5iozD1ssvA=
google.golang.org/api v0.294.0/go.mod h1:02qB8+Ox1ZFzcaKFMguy1nQLJmSIyvV6Ff4txJEXtl4=
google.golang.org/api v0.289.0 h1:DmH0c6NigNFmsvsohM9bxv+MzVhag3aGHnojA5fFQjc=
google.golang.org/api v0.289.0/go.mod h1:weJZ3lldHFYI0DBFNKpJelUDNnusTt5YaOEgxvt8ci8=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
@@ -2794,12 +2792,12 @@ google.golang.org/genproto v0.0.0-20230209215440-0dfe4f8abfcc/go.mod h1:RGgjbofJ
google.golang.org/genproto v0.0.0-20230216225411-c8e22ba71e44/go.mod h1:8B0gmkoRebU8ukX6HP+4wrVQUY1+6PkQ44BSyIlflHA=
google.golang.org/genproto v0.0.0-20230222225845-10f96fb3dbec/go.mod h1:3Dl5ZL0q0isWJt+FVcfpQyirqemEuLAK/iFvg1UP1Hw=
google.golang.org/genproto v0.0.0-20230306155012-7f2fa6fef1f4/go.mod h1:NWraEVixdDnqcqQ30jipen1STv2r/n24Wb7twVTGR4s=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d h1:C9v1o0/4quuhOAfmRXA2j+we0PqZIp8traLdeogF3Ms=
google.golang.org/genproto v0.0.0-20260715232425-e75dac1f907d/go.mod h1:Wz2wFJntZFmLGo7pLDXZ3wYk5hyc0Mb+SkHhDDXT+lU=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d h1:QwnJwPte4XXAkhPu26LTDIahnsMSUV0kK8HkxbC+Pc4=
google.golang.org/genproto/googleapis/api v0.0.0-20260715232425-e75dac1f907d/go.mod h1:WRrQ7/7N19PypuT0fxLOL5Lq0waoiRri4FbtHDEKrGE=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688 h1:cYNAzI2sUwhmCcoj9TxvihSrqsxt6uIkj3rDRhSDmW4=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260819154853-08b0e4226688/go.mod h1:DjtHYE8FKJLivXcBEjGwndXfIC23G0VpXiXKqG179uA=
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94 h1:YJjbgu+dkp5kUJLfpMyCLfBIWZb/FcJyuLeo1gVBOuo=
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94/go.mod h1:RRHjglSYABVCWpQ7USCpdfhcd9t4PkajvVwyynZizTc=
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800 h1:admdQBe8jR3VWhBsUrAOaF2Qw6K/+p5pSm1GN8+6Fw4=
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800/go.mod h1:FPk7EXUKMtImne7AmknoYjT4QXqKIzzRbeQIXzLk6fQ=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260715232425-e75dac1f907d h1:Jkpk39hlTZOIp3RbfvNX9R8Hv+Sw0X89nlU/xFOErsc=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260715232425-e75dac1f907d/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/grpc v1.19.0/go.mod h1:mqu4LbDTu4XGKhr4mRzUsmM4RtVoemTSY81AxZiDr8c=
google.golang.org/grpc v1.20.1/go.mod h1:10oTOabMzJvdu6/UiuZezV6QK5dSlG84ov/aaiqXj38=
google.golang.org/grpc v1.21.1/go.mod h1:oYelfM1adQP15Ek0mdvEgi9Df8B9CZIaU1084ijfRaM=
@@ -2840,8 +2838,8 @@ google.golang.org/grpc v1.51.0/go.mod h1:wgNDFcnuBGmxLKI/qn4T+m5BtEBYXJPvibbUPsA
google.golang.org/grpc v1.52.0/go.mod h1:pu6fVzoFb+NBYNAvQL08ic+lvB2IojljRYuun5vorUY=
google.golang.org/grpc v1.53.0/go.mod h1:OnIrk0ipVdj4N5d9IUoFUx72/VlD7+jUsHwZgwSMQpw=
google.golang.org/grpc v1.55.0/go.mod h1:iYEXKGkEBhg1PjZQvoYEVPTDkHo1/bjTnfwTeGONTY8=
google.golang.org/grpc v1.85.0-dev h1:HxkDyKIIZPpFnroC56tQv5gNuKTmVvi0t7TzOf5zt7g=
google.golang.org/grpc v1.85.0-dev/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6 h1:HfjjkdGIa8u9sP9EW5WCygy0kQDuTI/Tax4j//t24Fo=
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6/go.mod h1:ljCht0DrxQrXBDRTZp52Qxh3Ffk8CdYm2sj4O2QN2C0=
google.golang.org/grpc/cmd/protoc-gen-go-grpc v1.1.0/go.mod h1:6Kw0yEErY5E/yWrBtf03jp27GLLJujG4z/JK95pnjjw=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6 h1:ExN12ndbJ608cboPYflpTny6mXSzPrDLh0iTaVrRrds=
google.golang.org/grpc/examples v0.0.0-20250407062114-b368379ef8f6/go.mod h1:6ytKWczdvnpnO+m+JiG9NjEDzR1FJfsnmJdG7B8QVZ8=
@@ -2863,8 +2861,8 @@ google.golang.org/protobuf v1.27.1/go.mod h1:9q0QmTI4eRPtz6boOQmLYwt+qCgq0jsYwAQ
google.golang.org/protobuf v1.28.0/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
google.golang.org/protobuf v1.28.1/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
google.golang.org/protobuf v1.30.0/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
google.golang.org/protobuf v1.36.12 h1:pJOKDDOyeXErUroCihFAd5LQuwXBSpVnKGrj5o/fwxc=
google.golang.org/protobuf v1.36.12/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
gopkg.in/alecthomas/kingpin.v2 v2.2.6/go.mod h1:FMv+mEhP44yOT+4EoQTLFTRgOQ1FBLkstjWtayDeSgw=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20180628173108-788fd7840127/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
@@ -2917,23 +2915,23 @@ modernc.org/b v1.0.0/go.mod h1:uZWcZfRj1BpYzfN9JTerzlNUnnPsV9O2ZA8JsRcubNg=
modernc.org/cc/v3 v3.36.0/go.mod h1:NFUHyPn4ekoC/JHeZFfZurN6ixxawE1BnVonP/oahEI=
modernc.org/cc/v3 v3.36.2/go.mod h1:NFUHyPn4ekoC/JHeZFfZurN6ixxawE1BnVonP/oahEI=
modernc.org/cc/v3 v3.36.3/go.mod h1:NFUHyPn4ekoC/JHeZFfZurN6ixxawE1BnVonP/oahEI=
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
modernc.org/cc/v4 v4.29.1/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/cc/v4 v4.28.4 h1:Hd/4Es+MBj+/7hSdZaisNyu6bv3V0Dp2MdllyfqaH+c=
modernc.org/cc/v4 v4.28.4/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/ccgo/v3 v3.0.0-20220428102840-41399a37e894/go.mod h1:eI31LL8EwEBKPpNpA4bU1/i+sKOwOrQy8D87zWUcRZc=
modernc.org/ccgo/v3 v3.0.0-20220430103911-bc99d88307be/go.mod h1:bwdAnOoaIt8Ax9YdWGjxWsdkPcZyRPHqrOvJxaKAKGw=
modernc.org/ccgo/v3 v3.16.4/go.mod h1:tGtX0gE9Jn7hdZFeU88slbTh1UtCYKusWOoCJuvkWsQ=
modernc.org/ccgo/v3 v3.16.6/go.mod h1:tGtX0gE9Jn7hdZFeU88slbTh1UtCYKusWOoCJuvkWsQ=
modernc.org/ccgo/v3 v3.16.8/go.mod h1:zNjwkizS+fIFDrDjIAgBSCLkWbJuHF+ar3QRn+Z9aws=
modernc.org/ccgo/v3 v3.16.9/go.mod h1:zNMzC9A9xeNUepy6KuZBbugn3c0Mc9TeiJO4lgvkJDo=
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
modernc.org/ccgo/v4 v4.34.4 h1:OVnSOWQjVKOYkFxoHYB+qQmSHK5gqMqARM+K9DpR/Ws=
modernc.org/ccgo/v4 v4.34.4/go.mod h1:qdKqE8FNIYyysougB1RX9MxCzp5oJOcQXSobANJ4TuE=
modernc.org/ccorpus v1.11.6/go.mod h1:2gEUTrWqdpH2pXsmTM1ZkjeSrUWDpjMu2T6m29L/ErQ=
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/gc/v3 v3.1.3 h1:6QAplYyVO+KdPW3pGnqmJDUxtkec8ooEWvks/hhU3lc=
modernc.org/gc/v3 v3.1.3/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
modernc.org/httpfs v1.0.6/go.mod h1:7dosgurJGp0sPaRanU53W4xZYKh14wfzX420oZADeHM=
@@ -2944,8 +2942,8 @@ modernc.org/libc v1.16.17/go.mod h1:hYIV5VZczAmGZAnG15Vdngn5HSF5cSkbvfz2B7GRuVU=
modernc.org/libc v1.16.19/go.mod h1:p7Mg4+koNjc8jkqwcoFBJx7tXkpj00G77X7A72jXPXA=
modernc.org/libc v1.17.0/go.mod h1:XsgLldpP4aWlPlsjqKRdHPqCxCjISdHfM/yeWC5GyW0=
modernc.org/libc v1.17.1/go.mod h1:FZ23b+8LjxZs7XtFMbSzL/EhPxNbfZbErxEHc7cbD9s=
modernc.org/libc v1.74.4 h1:fX1Omw4o2/1C2iRkkIsrQTasJQldLhRmuPreXLoWs9k=
modernc.org/libc v1.74.4/go.mod h1:eeQAS9W3sZeKYMFubydxJpII9ybHWshk+7or7bLG9co=
modernc.org/libc v1.73.4 h1:+ra4Ui8ngyt8HDcO1FTDPWlkAh6yOdaO2yAoh8MddQA=
modernc.org/libc v1.73.4/go.mod h1:DXZ3eO8qMCNn2SnmTNCiC71nJ9Rcq3PsnpU6Vc4rWK8=
modernc.org/mathutil v1.1.1/go.mod h1:mZW8CKdRPY1v87qxC/wUdX5O1qDzXMP5TH3wjfpga6E=
modernc.org/mathutil v1.2.2/go.mod h1:mZW8CKdRPY1v87qxC/wUdX5O1qDzXMP5TH3wjfpga6E=
modernc.org/mathutil v1.4.1/go.mod h1:mZW8CKdRPY1v87qxC/wUdX5O1qDzXMP5TH3wjfpga6E=
@@ -2964,8 +2962,8 @@ modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
modernc.org/sqlite v1.18.1/go.mod h1:6ho+Gow7oX5V+OiOQ6Tr4xeqbx13UZ6t+Fw9IRUG4d4=
modernc.org/sqlite v1.57.0 h1:qNQP6xnx5M0ISNtlnxoOX0+cD5bJ0/gr9aMmndFczzg=
modernc.org/sqlite v1.57.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
modernc.org/sqlite v1.53.0 h1:20WG8N9q4ji/dEqGk4uiI0c6OPjSeLTNYGFCc3+7c1M=
modernc.org/sqlite v1.53.0/go.mod h1:xoEpOIpGrgT48H5iiyt/YXPCZPEzlfmfFwtk8Lklw8s=
modernc.org/strutil v1.1.0/go.mod h1:lstksw84oURvj9y3tn8lGvRxyRC1S2+g5uuIzNfIOBs=
modernc.org/strutil v1.1.1/go.mod h1:DE+MQQ/hjKBZS2zNInV5hhcipt5rLPWkmpbGeW5mmdw=
modernc.org/strutil v1.1.3/go.mod h1:MEHNA7PdEnEwLvspRMtWTNnp2nnyvMfkimT1NKNAGbw=
+2 -84
View File
@@ -9,7 +9,7 @@
# curl -fsSL ... | bash -s -- --version 4.34 --dir /usr/local/bin
#
# Options:
# --component COMP Which binary to install: weed, volume-rust, worker-rust, all (default: weed)
# --component COMP Which binary to install: weed, volume-rust, all (default: weed)
# --version VER Release version tag (default: latest)
# --large-disk Use large disk variant (5-byte offset, 8TB max volume)
# --dir DIR Installation directory (default: /usr/local/bin)
@@ -22,7 +22,6 @@ COMPONENT="weed"
VERSION=""
LARGE_DISK=false
INSTALL_DIR="/usr/local/bin"
WORKER_INSTALLED=false
# Colors (if terminal supports them)
if [ -t 1 ]; then
@@ -129,33 +128,6 @@ rust_asset_name() {
fi
}
# Does a release carry this asset? Answers 0 for yes and 1 for a 404, and stops
# the installer on anything else: a rate limit or a network blip must not read
# as "this release predates the binary" and quietly skip it.
asset_exists() {
local url="$1" code=""
if command -v curl &>/dev/null; then
code="$(curl -sL -o /dev/null -I -w '%{http_code}' "$url" || true)"
elif command -v wget &>/dev/null; then
# --spider's exit status folds 404 in with every other server error, so
# read the status line itself; -S prints one per redirect hop.
code="$(wget -S --spider -q -O /dev/null "$url" 2>&1 | awk '/^ *HTTP\// {c=$2} END {print c}')"
fi
case "$code" in
200) return 0 ;;
404) return 1 ;;
*) error "Could not check ${url} (HTTP ${code:-none}). Retry, or install components one at a time." ;;
esac
}
# Build Rust maintenance worker asset name. No large-disk variant: the worker
# maintains tables through the namespace and never opens a volume file.
worker_asset_name() {
local os="$1" arch="$2"
echo "weed-worker_${os}_${arch}.tar.gz"
}
# Install a single component
install_component() {
local component="$1" os="$2" arch="$3"
@@ -232,48 +204,10 @@ install_component() {
ok "Installed weed-volume to ${INSTALL_DIR}/${dest_name}"
;;
worker-rust)
# Published for linux only: the worker runs beside the cluster it
# maintains, and its dependency tree makes every extra target an
# expensive build.
case "$os" in
linux) ;;
*) error "Rust maintenance worker is not available for ${os}. Supported: linux" ;;
esac
case "$arch" in
amd64|arm64) ;;
*) error "Rust maintenance worker is not available for ${arch}. Supported: amd64, arm64" ;;
esac
asset_name="$(worker_asset_name "$os" "$arch")"
download_url="https://github.com/${REPO}/releases/download/${VERSION}/${asset_name}"
download "$download_url" "${tmpdir}/${asset_name}"
info "Extracting ${asset_name}..."
tar xzf "${tmpdir}/${asset_name}" -C "$tmpdir"
local worker_bin
worker_bin="$(find "$tmpdir" -name 'weed-worker' -type f | head -1)"
if [ -z "$worker_bin" ]; then
error "Could not find weed-worker binary in archive"
fi
chmod +x "$worker_bin"
install_binary "$worker_bin" "weed-worker"
WORKER_INSTALLED=true
ok "Installed weed-worker to ${INSTALL_DIR}/weed-worker"
;;
*)
error "Unknown component: ${component}. Use: weed, volume-rust, worker-rust, all"
error "Unknown component: ${component}. Use: weed, volume-rust, all"
;;
esac
# The trap is per-process, so a later component's would replace this one and
# leave the earlier extraction behind. Clean up here and hand the trap back;
# the error paths above exit, which still fires it.
rm -rf "$tmpdir"
trap - EXIT
}
# Copy binary to install dir, using sudo if needed
@@ -317,16 +251,6 @@ main() {
all)
install_component "weed" "$os" "$arch"
install_component "volume-rust" "$os" "$arch"
# The worker is published for linux amd64/arm64 only, and only by
# releases new enough to carry it; skip either case rather than fail
# an install that has already put two binaries in place.
if [ "$os" != "linux" ] || { [ "$arch" != "amd64" ] && [ "$arch" != "arm64" ]; }; then
warn "Skipping the Rust maintenance worker: no build for ${os}/${arch}"
elif ! asset_exists "https://github.com/${REPO}/releases/download/${VERSION}/$(worker_asset_name "$os" "$arch")"; then
warn "Skipping the Rust maintenance worker: ${VERSION} does not carry one"
else
install_component "worker-rust" "$os" "$arch"
fi
;;
*)
install_component "$COMPONENT" "$os" "$arch"
@@ -341,17 +265,11 @@ main() {
if [ "$COMPONENT" = "volume-rust" ] || [ "$COMPONENT" = "all" ]; then
info " weed-volume: ${INSTALL_DIR}/weed-volume"
fi
if [ "$WORKER_INSTALLED" = true ]; then
info " weed-worker: ${INSTALL_DIR}/weed-worker"
fi
echo ""
info "Quick start:"
info " weed master # Start master server"
info " weed volume -mserver=localhost:9333 # Start Go volume server"
info " weed-volume -mserver localhost:9333 # Start Rust volume server"
if [ "$WORKER_INSTALLED" = true ]; then
info " weed-worker --admin localhost:23646 # Start the Rust maintenance worker"
fi
}
main
+2 -2
View File
@@ -1,6 +1,6 @@
apiVersion: v1
description: SeaweedFS
name: seaweedfs
appVersion: "4.45"
appVersion: "4.41"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.45.0
version: 4.41.0
+2 -28
View File
@@ -22,8 +22,8 @@ helm install --values=values.yaml seaweedfs seaweedfs/seaweedfs
## Info:
* master/filer/volume are stateful sets with anti-affinity on the hostname,
so your deployment will be spread/HA.
* leveldb2 is the default filer backend; a mysql-compatible database (memsql, ...) enables HA (multiple filer instances) and the backup/HA it can provide.
* with `filer.extraEnvironmentVars.WEED_MYSQL_ENABLED` set to `"true"`, mysql user/password are created in a k8s secret (default: `<release>-seaweedfs-db-secret`) and injected to the filer with ENV. On any other store neither the secret nor the `WEED_MYSQL_*` env, plain or secret-backed, is rendered.
* chart is using memsql(mysql) as the filer backend to enable HA (multiple filer instances) and backup/HA memsql can provide.
* mysql user/password are created in a k8s secret (default: `<release>-seaweedfs-db-secret`) and injected to the filer with ENV.
* cert config exists and can be enabled, but not been tested, requires cert-manager to be installed.
## Prerequisites
@@ -289,32 +289,6 @@ stringData:
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read"]}]}'
```
#### Source S3 credentials from an existing Secret
To keep the keys out of `values.yaml` while still letting the chart generate the
identities file, point an identity at an existing Secret:
```yaml
s3:
enabled: true
enableAuth: true
credentials:
admin:
existingSecret: minio-root
accessKeyKey: root-user
secretKeyKey: root-password
```
`accessKeyKey` and `secretKeyKey` default to the chart's own key names
(`admin_access_key_id`, `admin_secret_access_key`, and the `read_` pair). The
generated `seaweedfs_s3_config` references the keys as `${SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID}`
and the gateway resolves them from the environment, which the chart wires up
from the Secret. Nothing is read from the cluster at render time, so
`helm template`, `--dry-run` and an Argo CD diff all render what an install
applies. Rotating a key in the Secret takes effect on the next pod restart, as
with any other environment variable. The COSI driver parses the config itself
and does not resolve these references.
## Admin Component
The admin component provides a modern web-based administration interface for managing SeaweedFS clusters. It includes:
@@ -21,7 +21,6 @@ spec:
{{- $nodePorts = .Values.admin.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.admin.service }}
ports:
- name: "http"
port: {{ .Values.admin.port }}
@@ -69,9 +69,6 @@ spec:
{{- if .Values.admin.priorityClassName }}
priorityClassName: {{ .Values.admin.priorityClassName | quote }}
{{- end }}
{{- if .Values.admin.schedulerName }}
schedulerName: {{ .Values.admin.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.admin.serviceAccountName }}
serviceAccountName: {{ .Values.admin.serviceAccountName | quote }}
@@ -64,9 +64,6 @@ spec:
{{- if .Values.allInOne.priorityClassName }}
priorityClassName: {{ .Values.allInOne.priorityClassName | quote }}
{{- end }}
{{- if .Values.allInOne.schedulerName }}
schedulerName: {{ .Values.allInOne.schedulerName | quote }}
{{- end }}
{{- if .Values.allInOne.serviceAccountName }}
serviceAccountName: {{ .Values.allInOne.serviceAccountName | quote }}
{{- end }}
@@ -84,9 +81,6 @@ spec:
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
env:
{{- include "seaweedfs.licenseEnv" . | nindent 12 }}
{{- if and .Values.allInOne.s3.enabled (or .Values.allInOne.s3.enableAuth .Values.s3.enableAuth .Values.filer.s3.enableAuth) }}
{{- include "seaweedfs.s3.credentialEnv" . | nindent 12 }}
{{- end }}
{{- /* Determine default cluster alias and the corresponding env var keys to avoid conflicts */}}
{{- $mergedExtraEnvironmentVars := dict }}
{{- include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.allInOne "target" $mergedExtraEnvironmentVars) }}
@@ -21,7 +21,6 @@ spec:
{{- $nodePorts = .Values.allInOne.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.allInOne.service }}
internalTrafficPolicy: {{ .Values.allInOne.service.internalTrafficPolicy | default "Cluster" }}
{{- if and (semverCompare ">=1.31-0" .Capabilities.KubeVersion.GitVersion) .Values.allInOne.s3.trafficDistribution }}
trafficDistribution: {{ include "seaweedfs.trafficDistribution" (dict "value" .Values.allInOne.s3.trafficDistribution "Capabilities" .Capabilities) }}
@@ -57,9 +57,6 @@ spec:
{{- if .Values.cosi.priorityClassName }}
priorityClassName: {{ .Values.cosi.priorityClassName | quote }}
{{- end }}
{{- if .Values.cosi.schedulerName }}
schedulerName: {{ .Values.cosi.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
serviceAccountName: {{ include "seaweedfs.componentName" (list . "objectstorage-provisioner") }}
{{- if .Values.cosi.initContainers }}
@@ -10,6 +10,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- if .Values.filer.metricsPort }}
monitoring: "true"
{{- end }}
{{- if .Values.filer.annotations }}
annotations:
{{- toYaml .Values.filer.annotations | nindent 4 }}
@@ -10,9 +10,6 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- if .Values.filer.metricsPort }}
monitoring: "true"
{{- end }}
annotations:
service.alpha.kubernetes.io/tolerate-unready-endpoints: "true"
{{- if .Values.filer.annotations }}
@@ -30,7 +30,6 @@ spec:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
monitoring: "true"
{{- end }}
{{- end }}
{{- end }}
@@ -76,9 +76,6 @@ spec:
{{- if .Values.filer.priorityClassName }}
priorityClassName: {{ .Values.filer.priorityClassName | quote }}
{{- end }}
{{- if .Values.filer.schedulerName }}
schedulerName: {{ .Values.filer.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.filer.initContainers }}
initContainers:
@@ -92,7 +89,6 @@ spec:
image: {{ template "seaweedfs.filer.image" . }}
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
env:
{{- $mysqlEnabled := include "seaweedfs.filer.mysqlEnabled" . }}
- name: POD_IP
valueFrom:
fieldRef:
@@ -105,7 +101,6 @@ spec:
valueFrom:
fieldRef:
fieldPath: metadata.namespace
{{- if $mysqlEnabled }}
- name: WEED_MYSQL_USERNAME
valueFrom:
secretKeyRef:
@@ -118,21 +113,10 @@ spec:
name: {{ include "seaweedfs.fullname" . }}-db-secret
key: password
optional: true
{{- end }}
- name: SEAWEEDFS_FULLNAME
value: "{{ include "seaweedfs.fullname" . }}"
{{- if and .Values.filer.s3.enabled .Values.filer.s3.enableAuth }}
{{- include "seaweedfs.s3.credentialEnv" . | nindent 12 }}
{{- end }}
{{- $mergedExtraEnvironmentVars := dict }}
{{- include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.filer "target" $mergedExtraEnvironmentVars) }}
{{- if not $mysqlEnabled }}
{{- range $key := keys $mergedExtraEnvironmentVars }}
{{- if hasPrefix "WEED_MYSQL_" $key }}
{{- $_ := unset $mergedExtraEnvironmentVars $key }}
{{- end }}
{{- end }}
{{- end }}
{{- range $key := keys $mergedExtraEnvironmentVars | sortAlpha }}
{{- $value := index $mergedExtraEnvironmentVars $key }}
- name: {{ $key }}
@@ -145,12 +129,10 @@ spec:
{{- end }}
{{- if .Values.filer.secretExtraEnvironmentVars }}
{{- range $key, $value := .Values.filer.secretExtraEnvironmentVars }}
{{- if or $mysqlEnabled (not (hasPrefix "WEED_MYSQL_" $key)) }}
- name: {{ $key }}
valueFrom: {{ toYaml $value | nindent 16 }}
{{- end }}
{{- end }}
{{- end }}
command:
- "/bin/sh"
- "-ec"
@@ -69,9 +69,6 @@ spec:
{{- if .Values.master.priorityClassName }}
priorityClassName: {{ .Values.master.priorityClassName | quote }}
{{- end }}
{{- if .Values.master.schedulerName }}
schedulerName: {{ .Values.master.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
serviceAccountName: {{ .Values.master.serviceAccountName | default (include "seaweedfs.serviceAccountName" .) | quote }} # for deleting statefulset pods after migration
{{- if .Values.master.initContainers }}
@@ -61,9 +61,6 @@ spec:
{{- if .Values.s3.priorityClassName }}
priorityClassName: {{ .Values.s3.priorityClassName | quote }}
{{- end }}
{{- if .Values.s3.schedulerName }}
schedulerName: {{ .Values.s3.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.s3.serviceAccountName }}
serviceAccountName: {{ .Values.s3.serviceAccountName | quote }}
@@ -94,9 +91,6 @@ spec:
fieldPath: metadata.namespace
- name: SEAWEEDFS_FULLNAME
value: "{{ include "seaweedfs.fullname" . }}"
{{- if .Values.s3.enableAuth }}
{{- include "seaweedfs.s3.credentialEnv" . | nindent 12 }}
{{- end }}
{{- $mergedExtraEnvironmentVars := dict }}
{{- include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.s3 "target" $mergedExtraEnvironmentVars) }}
{{- range $key := keys $mergedExtraEnvironmentVars | sortAlpha }}
@@ -149,8 +143,6 @@ spec:
{{- if .Values.s3.icebergPort }}
-port.iceberg={{ .Values.s3.icebergPort }} \
{{- end }}
{{- /* rendered even when 0: an if would drop the flag and weed would serve its default */}}
-port.lance={{ .Values.s3.lancePort | default 0 }} \
{{- range .Values.s3.extraArgs }}
{{ . }} \
{{- end }}
@@ -200,10 +192,6 @@ spec:
- containerPort: {{ .Values.s3.icebergPort }}
name: swfs-iceberg
{{- end }}
{{- if .Values.s3.lancePort }}
- containerPort: {{ .Values.s3.lancePort }}
name: swfs-lance
{{- end }}
{{- if .Values.s3.metricsPort }}
- containerPort: {{ .Values.s3.metricsPort }}
name: metrics
@@ -1,61 +0,0 @@
{{- define "seaweedfs.s3.lance.ingress.paths" -}}
paths:
- path: {{ .Values.s3.lanceIngress.path | quote }}
pathType: {{ .Values.s3.lanceIngress.pathType | quote }}
backend:
{{- if semverCompare ">=1.19-0" .Capabilities.KubeVersion.GitVersion }}
service:
name: {{ include "seaweedfs.componentName" (list . "s3") }}
port:
number: {{ .Values.s3.lancePort }}
{{- else }}
serviceName: {{ include "seaweedfs.componentName" (list . "s3") }}
servicePort: {{ .Values.s3.lancePort }}
{{- end }}
{{- end -}}
{{- if and .Values.s3.enabled .Values.s3.lancePort .Values.s3.lanceIngress.enabled }}
{{- $hosts := list }}
{{- if kindIs "slice" .Values.s3.lanceIngress.host }}
{{- $hosts = .Values.s3.lanceIngress.host }}
{{- else if .Values.s3.lanceIngress.host }}
{{- $hosts = list .Values.s3.lanceIngress.host }}
{{- end }}
{{- if semverCompare ">=1.19-0" .Capabilities.KubeVersion.GitVersion }}
apiVersion: networking.k8s.io/v1
{{- else if semverCompare ">=1.14-0" .Capabilities.KubeVersion.GitVersion }}
apiVersion: networking.k8s.io/v1beta1
{{- else }}
apiVersion: extensions/v1beta1
{{- end }}
kind: Ingress
metadata:
name: ingress-{{ include "seaweedfs.fullname" . }}-s3-lance
namespace: {{ .Release.Namespace }}
{{- with .Values.s3.lanceIngress.annotations }}
annotations:
{{- toYaml . | nindent 4 }}
{{- end }}
labels:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
helm.sh/chart: {{ .Chart.Name }}-{{ .Chart.Version | replace "+" "_" }}
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3-lance
spec:
{{- if .Values.s3.lanceIngress.className }}
ingressClassName: {{ .Values.s3.lanceIngress.className | quote }}
{{- end }}
tls:
{{ .Values.s3.lanceIngress.tls | default list | toYaml | nindent 6}}
rules:
{{- if $hosts }}
{{- range $host := $hosts }}
- host: {{ $host | quote }}
http:
{{- include "seaweedfs.s3.lance.ingress.paths" $ | nindent 6 }}
{{- end }}
{{- else }}
- http:
{{- include "seaweedfs.s3.lance.ingress.paths" . | nindent 4 }}
{{- end }}
{{- end }}
@@ -1,7 +1,4 @@
{{- /* Mirrors the condition the all-in-one deployment mounts this secret under,
so the flags that make it mount are exactly the flags that create it. */}}
{{- $allInOneAuth := and .Values.allInOne.enabled .Values.allInOne.s3.enabled (or .Values.allInOne.s3.enableAuth .Values.s3.enableAuth .Values.filer.s3.enableAuth) (not (or .Values.allInOne.s3.existingConfigSecret .Values.s3.existingConfigSecret .Values.filer.s3.existingConfigSecret)) }}
{{- if or (and (or .Values.s3.enabled .Values.allInOne.enabled) .Values.s3.enableAuth (not .Values.s3.existingConfigSecret)) (and .Values.filer.s3.enabled .Values.filer.s3.enableAuth (not .Values.filer.s3.existingConfigSecret)) $allInOneAuth }}
{{- if or (and (or .Values.s3.enabled .Values.allInOne.enabled) .Values.s3.enableAuth (not .Values.s3.existingConfigSecret)) (and .Values.filer.s3.enabled .Values.filer.s3.enableAuth (not .Values.filer.s3.existingConfigSecret)) }}
{{- $secretName := printf "%s-s3-secret" (include "seaweedfs.fullname" .) }}
{{- $legacySecretName := "seaweedfs-s3-secret" }}
{{- $lookupName := $secretName }}
@@ -17,20 +14,14 @@
{{- $adminCreds := $creds.admin | default dict -}}
{{- $access_key_admin := $adminCreds.accessKey -}}
{{- $secret_key_admin := $adminCreds.secretKey -}}
{{- if $adminCreds.existingSecret -}}
{{- $access_key_admin = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "admin" "accessKey")) -}}
{{- $secret_key_admin = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "admin" "secretKey")) -}}
{{- else if not (and $access_key_admin $secret_key_admin) -}}
{{- if not (and $access_key_admin $secret_key_admin) -}}
{{- $access_key_admin = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "admin_access_key_id" "length" 20 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- $secret_key_admin = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "admin_secret_access_key" "length" 40 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- end -}}
{{- $readCreds := $creds.read | default dict -}}
{{- $access_key_read := $readCreds.accessKey -}}
{{- $secret_key_read := $readCreds.secretKey -}}
{{- if $readCreds.existingSecret -}}
{{- $access_key_read = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "read" "accessKey")) -}}
{{- $secret_key_read = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "read" "secretKey")) -}}
{{- else if not (and $access_key_read $secret_key_read) -}}
{{- if not (and $access_key_read $secret_key_read) -}}
{{- $access_key_read = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "read_access_key_id" "length" 20 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- $secret_key_read = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "read_secret_access_key" "length" 40 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- end -}}
@@ -50,16 +41,10 @@ metadata:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3
stringData:
{{- /* An identity read from an existing Secret keeps its keys there; the
config below names the environment variables carrying them. */}}
{{- if not $adminCreds.existingSecret }}
admin_access_key_id: {{ $access_key_admin }}
admin_secret_access_key: {{ $secret_key_admin }}
{{- end }}
{{- if not $readCreds.existingSecret }}
read_access_key_id: {{ $access_key_read }}
read_secret_access_key: {{ $secret_key_read }}
{{- end }}
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read"]}]}'
{{- if .Values.filer.s3.auditLogConfig }}
filer_s3_auditLogConfig.json: |
@@ -21,7 +21,6 @@ spec:
{{- $nodePorts = .Values.s3.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.s3.service }}
internalTrafficPolicy: {{ .Values.s3.internalTrafficPolicy | default "Cluster" }}
{{- $td := .Values.s3.trafficDistribution | default .Values.filer.s3.trafficDistribution }}
{{- if and (semverCompare ">=1.31-0" .Capabilities.KubeVersion.GitVersion) $td }}
@@ -44,15 +43,6 @@ spec:
{{- end }}
protocol: TCP
{{- end }}
{{- if and .Values.s3.enabled .Values.s3.lancePort }}
- name: "swfs-lance"
port: {{ .Values.s3.lancePort }}
targetPort: {{ .Values.s3.lancePort }}
{{- if $nodePorts.lance }}
nodePort: {{ $nodePorts.lance }}
{{- end }}
protocol: TCP
{{- end }}
{{- if and .Values.s3.enabled .Values.s3.httpsPort }}
- name: "swfs-s3-tls"
port: {{ .Values.s3.httpsPort }}
@@ -61,9 +61,6 @@ spec:
{{- if .Values.sftp.priorityClassName }}
priorityClassName: {{ .Values.sftp.priorityClassName | quote }}
{{- end }}
{{- if .Values.sftp.schedulerName }}
schedulerName: {{ .Values.sftp.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.sftp.serviceAccountName }}
serviceAccountName: {{ .Values.sftp.serviceAccountName | quote }}
@@ -21,7 +21,6 @@ spec:
{{- $nodePorts = .Values.sftp.service.nodePorts | default dict }}
{{- end }}
type: {{ $serviceType }}
{{- include "seaweedfs.service.loadBalancerFields" .Values.sftp.service }}
internalTrafficPolicy: {{ .Values.sftp.internalTrafficPolicy | default "Cluster" }}
ports:
- name: "swfs-sftp"
@@ -76,18 +76,6 @@ Inject extra environment vars in the format key:value, if populated
{{- end }}
{{- end -}}
{{/* Whether the mysql filer store is selected; a flag the chart cannot read counts as selected. */}}
{{- define "seaweedfs.filer.mysqlEnabled" -}}
{{- $merged := dict -}}
{{- $_ := include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.filer "target" $merged) -}}
{{- $enabled := index $merged "WEED_MYSQL_ENABLED" -}}
{{- if or (kindIs "map" $enabled) (hasKey (.Values.filer.secretExtraEnvironmentVars | default dict) "WEED_MYSQL_ENABLED") -}}
true
{{- else if and $enabled (eq (lower (toString $enabled)) "true") -}}
true
{{- end -}}
{{- end -}}
{{/* Return the proper filer image */}}
{{- define "seaweedfs.filer.image" -}}
{{- if .Values.filer.imageOverride -}}
@@ -148,15 +136,6 @@ true
{{- end -}}
{{- end -}}
{{/* Lance namespace URL the worker's Lance container maintains; empty when unreachable */}}
{{- define "seaweedfs.worker.lanceNamespaceUrl" -}}
{{- if .Values.worker.namespaceUrl -}}
{{- .Values.worker.namespaceUrl -}}
{{- else if and .Values.s3.enabled .Values.s3.lancePort -}}
{{- printf "http://%s.%s:%d" (include "seaweedfs.componentName" (list . "s3")) .Release.Namespace (int .Values.s3.lancePort) -}}
{{- end -}}
{{- end -}}
{{/* Return the proper volume image */}}
{{- define "seaweedfs.volume.image" -}}
{{- if .Values.volume.imageOverride -}}
@@ -517,39 +496,6 @@ true
{{- end }}
{{- end -}}
{{/* Name of the environment variable carrying one generated S3 credential
field, e.g. SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID. The generated identities file
names it in place of the key when the key lives in an existing Secret.
Usage: include "seaweedfs.s3.credentialEnvName" (list "admin" "accessKey") */}}
{{- define "seaweedfs.s3.credentialEnvName" -}}
{{- $identity := index . 0 -}}
{{- $field := index . 1 -}}
{{- printf "SEAWEEDFS_S3_%s_%s" (upper $identity) (ternary "ACCESS_KEY_ID" "SECRET_ACCESS_KEY" (eq $field "accessKey")) -}}
{{- end -}}
{{/* Environment for the S3 identities the chart generates from an existing
Secret. The gateway resolves the ${VAR} references the identities file
carries, so the keys never enter the rendered manifests and a dry run
renders the same as an install. */}}
{{- define "seaweedfs.s3.credentialEnv" -}}
{{- $creds := $.Values.s3.credentials | default dict -}}
{{- range $identity := list "admin" "read" -}}
{{- $identityCreds := index $creds $identity | default dict -}}
{{- if $identityCreds.existingSecret }}
- name: {{ include "seaweedfs.s3.credentialEnvName" (list $identity "accessKey") }}
valueFrom:
secretKeyRef:
name: {{ $identityCreds.existingSecret | quote }}
key: {{ default (printf "%s_access_key_id" $identity) $identityCreds.accessKeyKey | quote }}
- name: {{ include "seaweedfs.s3.credentialEnvName" (list $identity "secretKey") }}
valueFrom:
secretKeyRef:
name: {{ $identityCreds.existingSecret | quote }}
key: {{ default (printf "%s_secret_access_key" $identity) $identityCreds.secretKeyKey | quote }}
{{- end -}}
{{- end -}}
{{- end -}}
{{/* Generate a compatible trafficDistribution value due to "PreferClose" fast deprecation in k8s v1.35.
Accepts a dict with "value" (the trafficDistribution string) and "Capabilities". */}}
{{- define "seaweedfs.trafficDistribution" -}}
@@ -557,23 +503,3 @@ true
{{- and (eq .value "PreferClose") (semverCompare ">=1.35-0" .Capabilities.KubeVersion.GitVersion) | ternary "PreferSameZone" .value -}}
{{- end -}}
{{- end -}}
{{/*
Render LoadBalancer-specific service fields (loadBalancerClass, loadBalancerIP,
loadBalancerSourceRanges), only when the service type is LoadBalancer.
Usage: {{ include "seaweedfs.service.loadBalancerFields" .Values.s3.service }}
*/}}
{{- define "seaweedfs.service.loadBalancerFields" -}}
{{- if eq (.type | default "ClusterIP") "LoadBalancer" }}
{{- with .loadBalancerClass }}
loadBalancerClass: {{ . }}
{{- end }}
{{- with .loadBalancerIP }}
loadBalancerIP: {{ . }}
{{- end }}
{{- with .loadBalancerSourceRanges }}
loadBalancerSourceRanges:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- end }}
{{- end -}}
@@ -73,9 +73,6 @@
{{- if .Values.s3.icebergPort }}
{{- $ports = append $ports .Values.s3.icebergPort }}
{{- end }}
{{- if .Values.s3.lancePort }}
{{- $ports = append $ports .Values.s3.lancePort }}
{{- end }}
{{- if .Values.s3.metricsPort }}
{{- $ports = append $ports .Values.s3.metricsPort }}
{{- end }}
@@ -100,9 +97,6 @@
{{- if .Values.worker.metricsPort }}
{{- $ports = append $ports .Values.worker.metricsPort }}
{{- end }}
{{- if and .Values.worker.lanceMetricsPort (include "seaweedfs.worker.lanceNamespaceUrl" .) }}
{{- $ports = append $ports .Values.worker.lanceMetricsPort }}
{{- end }}
{{- $targets = append $targets (dict "component" "worker" "ports" $ports) }}
{{- end }}
@@ -1,4 +1,4 @@
{{- if and .Values.filer.enabled (include "seaweedfs.filer.mysqlEnabled" .) }}
{{- if .Values.filer.enabled }}
apiVersion: v1
kind: Secret
type: Opaque
@@ -73,9 +73,6 @@ spec:
{{- if $volume.priorityClassName }}
priorityClassName: {{ $volume.priorityClassName | quote }}
{{- end }}
{{- if $volume.schedulerName }}
schedulerName: {{ $volume.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
serviceAccountName: {{ $volume.serviceAccountName | default (include "seaweedfs.serviceAccountName" $) | quote }} # for deleting statefulset pods after migration
{{- $initContainers_exists := include "seaweedfs.volume.initContainers_exists" $ -}}
@@ -64,9 +64,6 @@ spec:
{{- if .Values.worker.priorityClassName }}
priorityClassName: {{ .Values.worker.priorityClassName | quote }}
{{- end }}
{{- if .Values.worker.schedulerName }}
schedulerName: {{ .Values.worker.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.worker.serviceAccountName }}
serviceAccountName: {{ .Values.worker.serviceAccountName | quote }}
@@ -223,81 +220,6 @@ spec:
{{- if .Values.worker.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.worker.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
{{- if include "seaweedfs.worker.lanceNamespaceUrl" . }}
- name: worker-lance
image: {{ template "seaweedfs.worker.image" . }}
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
env:
- name: POD_NAME
valueFrom:
fieldRef:
fieldPath: metadata.name
{{- /* the URL crosses a shell line; metacharacters must arrive as data, not syntax */}}
- name: LANCE_NAMESPACE_URL
value: {{ include "seaweedfs.worker.lanceNamespaceUrl" . | quote }}
command:
- "/bin/sh"
- "-ec"
- |
{{- /* the armv7/386 placeholder is empty; exec of it becomes the shell and exits 0 */}}
if [ ! -s /usr/bin/weed-worker ]; then
echo "the Rust worker is not available on this platform ($(uname -m)); it ships for amd64 and arm64" >&2
exit 1
fi
exec /usr/bin/weed-worker \
--id="$POD_NAME" \
{{- if .Values.worker.adminServer }}
--admin={{ .Values.worker.adminServer }} \
{{- else }}
--admin={{ template "seaweedfs.fullname" . }}-admin.{{ .Release.Namespace }}:{{ .Values.admin.port }}{{ if .Values.admin.grpcPort }}.{{ .Values.admin.grpcPort }}{{ end }} \
{{- end }}
--namespace="$LANCE_NAMESPACE_URL" \
{{- if .Values.global.seaweedfs.enableSecurity }}
--tls-ca=/usr/local/share/ca-certificates/ca/tls.crt \
--tls-cert=/usr/local/share/ca-certificates/worker/tls.crt \
--tls-key=/usr/local/share/ca-certificates/worker/tls.key \
{{- end }}
{{- if .Values.worker.lanceMetricsPort }}
--metrics-port={{ .Values.worker.lanceMetricsPort }} \
--metrics-ip=0.0.0.0 \
{{- end }}
--max-concurrency={{ .Values.worker.maxExecute }}
{{- if .Values.global.seaweedfs.enableSecurity }}
volumeMounts:
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
- name: worker-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/worker/
{{- end }}
{{- if .Values.worker.lanceMetricsPort }}
ports:
- containerPort: {{ .Values.worker.lanceMetricsPort }}
name: lance-metrics
livenessProbe:
httpGet:
path: /health
port: lance-metrics
initialDelaySeconds: 30
periodSeconds: 60
successThreshold: 1
failureThreshold: 5
timeoutSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: lance-metrics
initialDelaySeconds: 20
periodSeconds: 15
successThreshold: 1
failureThreshold: 3
timeoutSeconds: 10
{{- end }}
{{- if .Values.worker.containerSecurityContext.enabled }}
securityContext: {{- omit .Values.worker.containerSecurityContext "enabled" | toYaml | nindent 12 }}
{{- end }}
{{- end }}
{{- if .Values.worker.sidecars }}
{{- include "seaweedfs.tplvalues.render" (dict "value" .Values.worker.sidecars "context" $) | nindent 8 }}
{{- end }}
@@ -12,22 +12,13 @@ metadata:
app.kubernetes.io/component: worker
spec:
clusterIP: None # Headless service
{{- $lanceMetrics := and .Values.worker.lanceMetricsPort (include "seaweedfs.worker.lanceNamespaceUrl" .) }}
{{- if or .Values.worker.metricsPort $lanceMetrics }}
ports:
{{- if .Values.worker.metricsPort }}
ports:
- name: "metrics"
port: {{ .Values.worker.metricsPort }}
targetPort: {{ .Values.worker.metricsPort }}
protocol: TCP
{{- end }}
{{- if $lanceMetrics }}
- name: "lance-metrics"
port: {{ .Values.worker.lanceMetricsPort }}
targetPort: {{ .Values.worker.lanceMetricsPort }}
protocol: TCP
{{- end }}
{{- end }}
selector:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
app.kubernetes.io/instance: {{ .Release.Name }}
@@ -1,7 +1,6 @@
{{- include "seaweedfs.compat" . -}}
{{- if .Values.worker.enabled }}
{{- $lanceMetrics := and .Values.worker.lanceMetricsPort (include "seaweedfs.worker.lanceNamespaceUrl" .) }}
{{- if or .Values.worker.metricsPort $lanceMetrics }}
{{- if .Values.worker.metricsPort }}
{{- if .Values.global.seaweedfs.monitoring.enabled }}
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
@@ -23,16 +22,9 @@ metadata:
{{- end }}
spec:
endpoints:
{{- if .Values.worker.metricsPort }}
- interval: 30s
port: metrics
scrapeTimeout: 5s
{{- end }}
{{- if $lanceMetrics }}
- interval: 30s
port: lance-metrics
scrapeTimeout: 5s
{{- end }}
selector:
matchLabels:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
-77
View File
@@ -242,10 +242,6 @@ master:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to specify a custom scheduler for master pods
# ref: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -539,10 +535,6 @@ volume:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to specify a custom scheduler for volume pods
# ref: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -823,10 +815,6 @@ filer:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to specify a custom scheduler for filer pods
# ref: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -897,7 +885,6 @@ filer:
# extraEnvVars is a list of extra environment variables to set with the stateful set.
extraEnvironmentVars:
# the WEED_MYSQL_* keys and the db credential secret only render while this is "true"
WEED_MYSQL_ENABLED: "false"
WEED_MYSQL_HOSTNAME: "mysql-db-host"
WEED_MYSQL_PORT: "3306"
@@ -1006,9 +993,6 @@ s3:
# Iceberg catalog REST port (Apache Iceberg REST Catalog API)
# Set to a port number to enable, or 0/null to disable
icebergPort: null
# Lance Namespace port; weed serves 9101 by default, 0 disables it
# (and, unless worker.namespaceUrl points elsewhere, the worker's Lance container)
lancePort: 9101
loggingOverrideLevel: null
# enable user & permission to s3 (need to inject to all services)
enableAuth: false
@@ -1018,24 +1002,13 @@ s3:
# Optionally provide explicit credentials for the S3 gateway.
# When set, these are used in the generated s3 secret instead of
# auto-generating random credentials.
# An identity may instead name an existing Secret to read its keys from. The
# generated config then references the keys through environment variables, so
# nothing is looked up at render time and a dry run renders what an install
# applies. Note the COSI driver parses the config itself and does not resolve
# those references.
# credentials:
# admin:
# accessKey: ""
# secretKey: ""
# existingSecret: ""
# accessKeyKey: admin_access_key_id
# secretKeyKey: admin_secret_access_key
# read:
# accessKey: ""
# secretKey: ""
# existingSecret: ""
# accessKeyKey: read_access_key_id
# secretKeyKey: read_secret_access_key
auditLogConfig: {}
# You may specify buckets to be created during the install or upgrade process.
# Buckets may be exposed publicly by setting `anonymousRead` to `true`
@@ -1101,10 +1074,6 @@ s3:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to assign a custom scheduler to server pods
# ref: https://kubernetes.io/docs/tasks/extend-kubernetes/configure-multiple-schedulers/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -1187,16 +1156,11 @@ s3:
# Service settings
service:
type: ClusterIP
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
http: null
https: null
iceberg: null
lance: null
metrics: null
icebergIngress:
@@ -1208,15 +1172,6 @@ s3:
annotations: {}
tls: []
lanceIngress:
enabled: false
className: ""
host: "seaweedfs-lance.cluster.local"
path: "/"
pathType: Prefix
annotations: {}
tls: []
sftp:
enabled: false
imageOverride: null
@@ -1272,7 +1227,6 @@ sftp:
tolerations: ""
nodeSelector: ""
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
podSecurityContext: {}
containerSecurityContext: {}
@@ -1305,10 +1259,6 @@ sftp:
# Service settings
service:
type: ClusterIP
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
sftp: null
@@ -1401,7 +1351,6 @@ admin:
tolerations: ""
nodeSelector: ""
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
podSecurityContext: {}
containerSecurityContext: {}
@@ -1451,10 +1400,6 @@ admin:
service:
type: ClusterIP
annotations: {}
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
http: null
@@ -1473,15 +1418,6 @@ worker:
metricsPort: 9327
metricsIp: "" # If empty, defaults to 0.0.0.0
# The lance_* jobs run in their own container, /usr/bin/weed-worker.
# amd64/arm64 only; pin mixed clusters with worker.affinity/nodeSelector.
# Lance namespace URL override; empty derives it from s3.lancePort.
namespaceUrl: ""
# Metrics port for the Lance worker container; the Go worker keeps metricsPort
lanceMetricsPort: 9328
# Admin server to connect to
adminServer: ""
@@ -1557,7 +1493,6 @@ worker:
tolerations: ""
nodeSelector: ""
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
podSecurityContext: {}
containerSecurityContext: {}
@@ -1688,10 +1623,6 @@ allInOne:
annotations: {} # Annotations for the service
type: ClusterIP # Service type (ClusterIP, NodePort, LoadBalancer)
internalTrafficPolicy: Cluster # Internal traffic policy
# used only when type is LoadBalancer
loadBalancerClass: ""
loadBalancerIP: ""
loadBalancerSourceRanges: []
# fixed nodePorts, used only when type is NodePort or LoadBalancer
nodePorts:
master: null
@@ -1801,10 +1732,6 @@ allInOne:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# Used to assign a custom scheduler to pods
# ref: https://kubernetes.io/docs/tasks/extend-kubernetes/configure-multiple-schedulers/
schedulerName: ""
# Used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -1867,10 +1794,6 @@ cosi:
podSecurityContext: {}
containerSecurityContext: {}
# used to assign a custom scheduler to cosi pods
# ref: https://kubernetes.io/docs/tasks/extend-kubernetes/configure-multiple-schedulers/
schedulerName: ""
extraVolumes: ""
extraVolumeMounts: ""
-66
View File
@@ -1,66 +0,0 @@
<?xml version="1.0" encoding="UTF-8"?>
<svg id="Ebene_2" data-name="Ebene 2" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" viewBox="0 0 512 512">
<defs>
<style>
.cls-1 {
fill: url(#Unbenannter_Verlauf_20);
}
.cls-2 {
fill: url(#Unbenannter_Verlauf_32);
}
.cls-3 {
fill: url(#Unbenannter_Verlauf_42);
}
.cls-4 {
fill: #fff;
}
.cls-5 {
fill: url(#Unbenannter_Verlauf_12);
}
.cls-6 {
fill: url(#Unbenannter_Verlauf_28);
}
</style>
<linearGradient id="Unbenannter_Verlauf_12" data-name="Unbenannter Verlauf 12" x1="19.92" y1="19.92" x2="492.08" y2="492.08" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#0282e9"/>
<stop offset=".48" stop-color="#0162bf"/>
<stop offset=".87" stop-color="#01469b"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_42" data-name="Unbenannter Verlauf 42" x1="180.67" y1="151.78" x2="234.43" y2="384.66" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#23b5e6"/>
<stop offset=".41" stop-color="#138bcf"/>
<stop offset=".86" stop-color="#0158b3"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_28" data-name="Unbenannter Verlauf 28" x1="275.58" y1="46.33" x2="275.58" y2="466.33" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#9fecf6"/>
<stop offset=".49" stop-color="#62dff2"/>
<stop offset="1" stop-color="#0a6cb7"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_20" data-name="Unbenannter Verlauf 20" x1="124.99" y1="226.28" x2="235.06" y2="462.33" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#c9f9fd"/>
<stop offset=".29" stop-color="#9fe7f6"/>
<stop offset=".93" stop-color="#36bbe5"/>
<stop offset="1" stop-color="#2ab6e3"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_32" data-name="Unbenannter Verlauf 32" x1="358.38" y1="251.34" x2="277.45" y2="473.69" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#43cbee"/>
<stop offset=".42" stop-color="#29a1d6"/>
<stop offset="1" stop-color="#0161b2"/>
</linearGradient>
</defs>
<g id="seaweedfs">
<rect id="background" class="cls-5" width="512" height="512" rx="68" ry="68"/>
<circle class="cls-4" cx="366.58" cy="152.44" r="31.33"/>
<circle class="cls-4" cx="386.81" cy="223.33" r="18.67"/>
<circle class="cls-4" cx="383.69" cy="288" r="9.22"/>
<path class="cls-3" d="M177.81,152.44s22.5,8.02,30.17,28.02c7.67,20,3.83,61.31,9.5,81.98s19.33,29,21,45.67-1.67,76-1.67,76c0,0-24-31.89-29.5-64s-21.53-41.06-24.6-65.72c-3.07-24.67.77-32.61,2.1-54.61s-1.84-35.72-7-47.33Z"/>
<path class="cls-6" d="M288.58,46.33c-14.56,7.67-42.22,35.22-42.22,79.11s23.11,54.22,23.11,94.22-38.89,48.11-38.89,95.56,21.56,66,21.56,95.56-11.11,55.56-11.11,55.56c0,0,26.97-17.72,33.33-60s-6.89-54.44,7.22-81.78,39-48.22,39-92.89-36.78-89.22-40.89-108.44-2.33-43.22,8.89-76.89Z"/>
<path class="cls-1" d="M120.58,228.33s31.72,14.27,41.55,36.5c9.83,22.22,3.45,40.17,17.45,61.17s51,41,56,73.33-7.67,66.33-7.67,66.33c0,0-5.67-59-30.33-73.33s-46.33-35.67-54.33-59-1.33-50.67-5.33-69c-4-18.33-7.33-23.67-17.33-36Z"/>
<path class="cls-2" d="M352.25,249.11c-9,5.67-18.87,17.85-21.71,31.2s2.94,19.57-7.06,41.57-22.19,25.28-30.97,49.72.21,35.36-11.33,59.56-23.92,35.17-23.92,35.17c0,0,34.77-8.17,45.69-30.94s5.08-35.22,20.53-52.61,24.67-30.89,27.25-56.72-7.33-26.22,1.53-76.94Z"/>
</g>
</svg>

Before

Width:  |  Height:  |  Size: 3.5 KiB

+1 -954
View File
@@ -9263,959 +9263,6 @@
}
]
},
{
"type": "row",
"title": "Plugin Workers",
"collapsed": true,
"id": 229,
"gridPos": {
"h": 1,
"w": 24,
"x": 0,
"y": 27
},
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Plugin Workers Connected",
"description": "Workers with a live control stream. Unlike the Admin/Maintenance row, this counts plugin workers, which is what a Rust or Go plugin worker actually is.",
"type": "stat",
"id": 219,
"gridPos": {
"h": 4,
"w": 6,
"x": 0,
"y": 28
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "count(SeaweedFS_worker_connected{cluster=~\"$cluster\"} == 1) or vector(0)",
"instant": true,
"range": false,
"refId": "A"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Worker Job Failures (1h)",
"description": "Jobs that ended in failure across all workers in the last hour.",
"type": "stat",
"id": 220,
"gridPos": {
"h": 4,
"w": 6,
"x": 6,
"y": 28
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(increase(SeaweedFS_worker_jobs_total{cluster=~\"$cluster\",result=\"failed\"}[1h])) or vector(0)",
"instant": true,
"range": false,
"refId": "A"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Objects Seen vs Skipped",
"description": "A sweep that proposes nothing because there was nothing to do and one that proposes nothing because it could read nothing produce the same proposal count. Skips are the difference, and are the thing to alert on.",
"type": "timeseries",
"id": 221,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 32
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type) (rate(SeaweedFS_worker_objects_seen_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "seen {{job_type}}"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type, reason) (rate(SeaweedFS_worker_objects_skipped_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "B",
"legendFormat": "skipped {{job_type}} ({{reason}})"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Detections and Proposals",
"description": "Detection sweeps by outcome, and the work they handed to admin.",
"type": "timeseries",
"id": 222,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 32
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type, result) (rate(SeaweedFS_worker_detections_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "detections {{job_type}} ({{result}})"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type) (rate(SeaweedFS_worker_proposals_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "B",
"legendFormat": "proposals {{job_type}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Jobs Executed",
"description": "Jobs run by these workers, by outcome.",
"type": "timeseries",
"id": 223,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 40
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type, result) (rate(SeaweedFS_worker_jobs_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "{{job_type}} ({{result}})"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Job Duration p99",
"description": "How long a job takes, per job type.",
"type": "timeseries",
"id": 224,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 40
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "s",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "histogram_quantile(0.99, sum(rate(SeaweedFS_worker_job_seconds_bucket{cluster=~\"$cluster\"}[$__rate_interval])) by (le, job_type))",
"range": true,
"refId": "A",
"legendFormat": "{{job_type}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Worker Slots",
"description": "Capacity a worker advertised and how much of it is in use. Held slots are what admin schedules against.",
"type": "timeseries",
"id": 225,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 48
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (lane) (SeaweedFS_worker_slots_used{cluster=~\"$cluster\"})",
"range": true,
"refId": "A",
"legendFormat": "used {{lane}}"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (lane) (SeaweedFS_worker_slots_total{cluster=~\"$cluster\"})",
"range": true,
"refId": "B",
"legendFormat": "total {{lane}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Control Stream Events",
"description": "Connects, closes and failures. A worker that reconnects steadily is usually two workers sharing one id, evicting each other.",
"type": "timeseries",
"id": 226,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 48
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (event) (rate(SeaweedFS_worker_stream_events_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "{{event}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Lance Maintenance Reclaimed",
"description": "What the Lance jobs actually removed. A job that runs every minute and reclaims nothing is a different thing from a job that never runs.",
"type": "timeseries",
"id": 227,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 56
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_fragments_removed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "fragments merged"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_rows_indexed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "B",
"legendFormat": "rows indexed"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_versions_removed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "C",
"legendFormat": "versions removed"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Storage Reclaimed",
"description": "Bytes freed by version cleanup. Its own panel because bytes and counts do not belong on one axis.",
"type": "timeseries",
"id": 228,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 56
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "Bps",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_bytes_reclaimed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "reclaimed"
}
]
}
]
},
{
"type": "row",
"title": "Go Runtime",
@@ -10225,7 +9272,7 @@
"h": 1,
"w": 24,
"x": 0,
"y": 28
"y": 27
},
"panels": [
{
-65
View File
@@ -2929,70 +2929,6 @@ dependencies = [
"prost 0.13.5",
]
[[package]]
name = "protoc-bin-vendored"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d1c381df33c98266b5f08186583660090a4ffa0889e76c7e9a5e175f645a67fa"
dependencies = [
"protoc-bin-vendored-linux-aarch_64",
"protoc-bin-vendored-linux-ppcle_64",
"protoc-bin-vendored-linux-s390_64",
"protoc-bin-vendored-linux-x86_32",
"protoc-bin-vendored-linux-x86_64",
"protoc-bin-vendored-macos-aarch_64",
"protoc-bin-vendored-macos-x86_64",
"protoc-bin-vendored-win32",
]
[[package]]
name = "protoc-bin-vendored-linux-aarch_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c350df4d49b5b9e3ca79f7e646fde2377b199e13cfa87320308397e1f37e1a4c"
[[package]]
name = "protoc-bin-vendored-linux-ppcle_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a55a63e6c7244f19b5c6393f025017eb5d793fd5467823a099740a7a4222440c"
[[package]]
name = "protoc-bin-vendored-linux-s390_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1dba5565db4288e935d5330a07c264a4ee8e4a5b4a4e6f4e83fad824cc32f3b0"
[[package]]
name = "protoc-bin-vendored-linux-x86_32"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8854774b24ee28b7868cd71dccaae8e02a2365e67a4a87a6cd11ee6cdbdf9cf5"
[[package]]
name = "protoc-bin-vendored-linux-x86_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b38b07546580df720fa464ce124c4b03630a6fb83e05c336fea2a241df7e5d78"
[[package]]
name = "protoc-bin-vendored-macos-aarch_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "89278a9926ce312e51f1d999fee8825d324d603213344a9a706daa009f1d8092"
[[package]]
name = "protoc-bin-vendored-macos-x86_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "81745feda7ccfb9471d7a4de888f0652e806d5795b61480605d4943176299756"
[[package]]
name = "protoc-bin-vendored-win32"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "pxfm"
version = "0.1.28"
@@ -4593,7 +4529,6 @@ dependencies = [
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"protoc-bin-vendored",
"rand 0.10.2",
"redb",
"reed-solomon-erasure",
-4
View File
@@ -137,10 +137,6 @@ tempfile = "3"
[build-dependencies]
tonic-build = "0.12"
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry.
protoc-bin-vendored = "3"
[patch.crates-io]
reed-solomon-erasure = { path = "vendor/reed-solomon-erasure" }
+2 -10
View File
@@ -1,18 +1,10 @@
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Use the protoc that ships with protoc-bin-vendored rather than a system
// one, so the build needs no package manager and always sees the same
// version. An explicit PROTOC still wins, for packagers supplying their own.
if std::env::var_os("PROTOC").is_none() {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
}
let out_dir = std::path::PathBuf::from(std::env::var("OUT_DIR")?);
tonic_build::configure()
.build_server(true)
.build_client(true)
// filer.proto uses proto3 optional, which protoc rejects without this
// flag before 3.15. The vendored protoc is newer, but it still accepts
// the flag, so this keeps a build against an older PROTOC working.
// filer.proto uses proto3 optional; older protoc (e.g. 3.12 from ubuntu-22.04 apt)
// rejects it without this flag, and newer protoc still accepts the flag
.protoc_arg("--experimental_allow_proto3_optional")
.file_descriptor_set_path(out_dir.join("seaweed_descriptor.bin"))
.compile_protos(
-34
View File
@@ -27,8 +27,6 @@ service Seaweed {
}
rpc VolumeList (VolumeListRequest) returns (VolumeListResponse) {
}
rpc VolumeListStream (VolumeListRequest) returns (stream VolumeListStreamResponse) {
}
rpc LookupEcVolume (LookupEcVolumeRequest) returns (LookupEcVolumeResponse) {
}
rpc VacuumVolume (VacuumVolumeRequest) returns (VacuumVolumeResponse) {
@@ -408,44 +406,12 @@ message TopologyInfo {
map<string, DiskInfo> diskInfos = 3;
}
message VolumeListRequest {
// Empty and zero take everything. Only the volumes and ec shards listed
// under a disk are selected; the topology and its disk counters are always
// reported in full.
string collection = 1;
repeated uint32 volume_ids = 2;
// The one collection the empty string cannot name. A named collection wins.
bool default_collection_only = 3;
// Empty and zero take everything. Wildcards are supported.
string remote_storage_name = 4;
bool local_volume_only = 5;
// The topology, its disks and their counters alone, without the volumes
// and ec shards. Selecting volumes above contradicts this and is refused.
// A master that predates this field ignores it and answers in full.
bool topology_only = 6;
}
message VolumeListResponse {
TopologyInfo topology_info = 1;
uint64 volume_size_limit_mb = 2;
}
// VolumeListStream answers the same request as VolumeList without either end
// holding every volume in the cluster at once. At 800k volumes the reply is
// 36MB on the wire but 305MB as messages, which the master built in full
// before sending any of it.
message VolumeListStreamResponse {
// Sent once, first, listing no volumes: the topology, its disks and their
// counters. Every message after carries volumes for one of those disks.
VolumeListResponse header = 1;
// Which disk this batch is from. A disk arrives over as many batches as it
// takes, so append rather than assign.
string data_center = 2;
string rack = 3;
string data_node = 4;
string disk_type = 5;
repeated VolumeInformationMessage volume_infos = 6;
repeated VolumeEcShardInformationMessage ec_shard_infos = 7;
}
message LookupEcVolumeRequest {
uint32 volume_id = 1;
}
-14
View File
@@ -45,8 +45,6 @@ service VolumeServer {
}
rpc VolumeUnmount (VolumeUnmountRequest) returns (VolumeUnmountResponse) {
}
rpc VolumeConsolidateIndex (VolumeConsolidateIndexRequest) returns (VolumeConsolidateIndexResponse) {
}
rpc VolumeDelete (VolumeDeleteRequest) returns (VolumeDeleteResponse) {
}
rpc VolumeMarkReadonly (VolumeMarkReadonlyRequest) returns (VolumeMarkReadonlyResponse) {
@@ -246,12 +244,6 @@ message VolumeUnmountRequest {
message VolumeUnmountResponse {
}
message VolumeConsolidateIndexRequest {
uint32 volume_id = 1;
}
message VolumeConsolidateIndexResponse {
}
message VolumeDeleteRequest {
uint32 volume_id = 1;
bool only_empty = 2;
@@ -265,8 +257,6 @@ message VolumeDeleteResponse {
message VolumeMarkReadonlyRequest {
uint32 volume_id = 1;
bool persist = 2;
// reject writes but keep accepting deletes, so expiring data can drain the volume
bool can_delete = 3;
}
message VolumeMarkReadonlyResponse {
}
@@ -539,7 +529,6 @@ message VolumeEcShardsInfoResponse {
uint64 volume_size = 2;
uint64 file_count = 3;
uint64 file_deleted_count = 4;
EcShardConfig ec_shard_config = 10; // the layout this holder serves reads through; a binary predating a field reports it as unset
}
message EcShardInfo {
@@ -605,7 +594,6 @@ message VolumeInfo {
uint64 expire_at_sec = 6; // expiration time of ec volume
bool read_only = 7;
EcShardConfig ec_shard_config = 8; // EC shard configuration (optional, null = use default 10+4)
bool read_only_can_delete = 9; // with read_only: writes are rejected but deletes still land
}
// EcShardConfig specifies erasure coding shard configuration
@@ -613,7 +601,6 @@ message EcShardConfig {
uint32 data_shards = 1; // Number of data shards (e.g., 10)
uint32 parity_shards = 2; // Number of parity shards (e.g., 4)
int64 encode_ts_ns = 3; // encode time (unix nanos); a read served from a shard of a different encode run is rejected
int64 block_size = 4; // uniform block layout: each shard is a single contiguous block of this many bytes; 0 = legacy 1GiB/1MiB two-tier layout
}
// EcBitrotProtection is the entire content of a bitrot checksum sidecar
// (<base>.ecsum for the legacy generation, <base>.ecsum.v<N> for vacuum
@@ -716,7 +703,6 @@ enum VolumeScrubMode {
FULL = 2;
LOCAL = 3;
CHECKSUM = 4; // EC only: verify each local shard's raw bytes against the bitrot checksum sidecar
READS = 5; // like FULL, but EC intervals no shard can serve are reconstructed from parity
}
message ScrubVolumeRequest {
-47
View File
@@ -256,8 +256,6 @@ pub struct VolumeServerConfig {
pub https_client_ca_file: String,
pub grpc_cert_file: String,
pub grpc_key_file: String,
pub grpc_client_cert_file: String,
pub grpc_client_key_file: String,
pub grpc_ca_file: String,
pub grpc_allowed_wildcard_domain: String,
pub grpc_volume_allowed_common_names: Vec<String>,
@@ -807,8 +805,6 @@ fn resolve_config(cli: Cli) -> VolumeServerConfig {
https_client_ca_file: sec.https_client_ca_file,
grpc_cert_file: sec.grpc_cert_file,
grpc_key_file: sec.grpc_key_file,
grpc_client_cert_file: sec.grpc_client_cert_file,
grpc_client_key_file: sec.grpc_client_key_file,
grpc_ca_file: sec.grpc_ca_file,
grpc_allowed_wildcard_domain: sec.grpc_allowed_wildcard_domain,
grpc_volume_allowed_common_names: sec.grpc_volume_allowed_common_names,
@@ -841,8 +837,6 @@ pub struct SecurityConfig {
pub https_client_ca_file: String,
pub grpc_cert_file: String,
pub grpc_key_file: String,
pub grpc_client_cert_file: String,
pub grpc_client_key_file: String,
pub grpc_ca_file: String,
pub grpc_allowed_wildcard_domain: String,
pub grpc_volume_allowed_common_names: Vec<String>,
@@ -888,8 +882,6 @@ const SECURITY_CONFIG_FILE_NAME: &str = "security.toml";
/// [grpc.volume]
/// cert = "/path/to/cert.pem"
/// key = "/path/to/key.pem"
/// client_cert = "/path/to/client-cert.pem"
/// client_key = "/path/to/client-key.pem"
/// allowed_commonNames = "volume-a.internal,volume-b.internal"
/// ```
pub fn parse_security_config(path: &str) -> SecurityConfig {
@@ -1011,8 +1003,6 @@ pub fn parse_security_config(path: &str) -> SecurityConfig {
Section::GrpcVolume => match key {
"cert" => cfg.grpc_cert_file = value.to_string(),
"key" => cfg.grpc_key_file = value.to_string(),
"client_cert" => cfg.grpc_client_cert_file = value.to_string(),
"client_key" => cfg.grpc_client_key_file = value.to_string(),
// Go only reads CA from [grpc], not [grpc.volume]
"allowed_commonNames" => {
cfg.grpc_volume_allowed_common_names =
@@ -1144,12 +1134,6 @@ fn apply_env_overrides(cfg: &mut SecurityConfig) {
if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_KEY") {
cfg.grpc_key_file = v;
}
if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_CLIENT_CERT") {
cfg.grpc_client_cert_file = v;
}
if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_CLIENT_KEY") {
cfg.grpc_client_key_file = v;
}
if let Ok(v) = std::env::var("WEED_GRPC_CA") {
cfg.grpc_ca_file = v;
} else if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_CA") {
@@ -1247,8 +1231,6 @@ mod tests {
"WEED_HTTPS_CLIENT_CA",
"WEED_GRPC_VOLUME_CERT",
"WEED_GRPC_VOLUME_KEY",
"WEED_GRPC_VOLUME_CLIENT_CERT",
"WEED_GRPC_VOLUME_CLIENT_KEY",
"WEED_GRPC_CA",
"WEED_GRPC_VOLUME_CA",
"WEED_GRPC_ALLOWED_WILDCARD_DOMAIN",
@@ -1518,35 +1500,6 @@ key = "/etc/seaweedfs/volume-key.pem"
});
}
#[test]
fn test_parse_security_config_uses_grpc_volume_client_cert() {
let _guard = process_state_lock();
let tmp = tempfile::NamedTempFile::new().unwrap();
std::fs::write(
tmp.path(),
r#"
[grpc.volume]
cert = "/etc/seaweedfs/volume-cert.pem"
key = "/etc/seaweedfs/volume-key.pem"
client_cert = "/etc/seaweedfs/volume-client-cert.pem"
client_key = "/etc/seaweedfs/volume-client-key.pem"
"#,
)
.unwrap();
with_cleared_security_env(|| {
let cfg = parse_security_config(tmp.path().to_str().unwrap());
assert_eq!(
cfg.grpc_client_cert_file,
"/etc/seaweedfs/volume-client-cert.pem"
);
assert_eq!(
cfg.grpc_client_key_file,
"/etc/seaweedfs/volume-client-key.pem"
);
});
}
#[test]
fn test_parse_security_config_uses_grpc_peer_name_policy() {
let _guard = process_state_lock();
-10
View File
@@ -330,16 +330,6 @@ pub fn delete_collection_metrics(collection: &str) {
delete_partial_match_collection(&DISK_SIZE_GAUGE, collection);
}
/// Drop a collection's volume server series once its last volume leaves this
/// server. These gauges are only ever set for collections still present, so the
/// values from the heartbeat that saw the last volume would otherwise stand
/// until the process restarts.
pub fn delete_volume_server_collection_metrics(collection: &str) {
let _ = DISK_SIZE_GAUGE.remove_label_values(&[collection, DISK_SIZE_LABEL_NORMAL]);
let _ = DISK_SIZE_GAUGE.remove_label_values(&[collection, DISK_SIZE_LABEL_DELETED_BYTES]);
delete_partial_match_collection(&READ_ONLY_VOLUME_GAUGE, collection);
}
/// Remove all metric entries from a GaugeVec where the "collection" label matches.
/// This emulates Go's `DeletePartialMatch(prometheus.Labels{"collection": collection})`.
fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
@@ -17,7 +17,7 @@
//! narrow TOCTOU window (a hostname that resolves to a public IP here and then
//! flips to a blocked one when the SDK dials) remains as a follow-up.
use std::net::{IpAddr, Ipv4Addr, Ipv6Addr};
use std::net::{IpAddr, Ipv4Addr};
/// AWS/Azure/GCP IPv4 instance-metadata-service (IMDS) address. It is
/// link-local and thus already covered by [`is_link_local`], but is named
@@ -74,52 +74,9 @@ fn is_cgnat(ip: IpAddr) -> bool {
}
}
/// Returns the IPv4 address carried by an IPv6 transition address -- NAT64
/// 64:ff9b::/96 (RFC 6052), 6to4 2002::/16 (RFC 3056), Teredo 2001:0000::/32
/// (RFC 4380), and the deprecated IPv4-compatible ::/96 (RFC 4291) -- or None
/// when it is not one of those. IPv4-mapped ::ffff:0:0/96 is excluded; it is
/// already normalized via `to_ipv4_mapped`.
fn embedded_transition_ipv4(v6: Ipv6Addr) -> Option<Ipv4Addr> {
let o = v6.octets();
if o[0] == 0x00
&& o[1] == 0x64
&& o[2] == 0xff
&& o[3] == 0x9b
&& o[4..12].iter().all(|&b| b == 0)
{
return Some(Ipv4Addr::new(o[12], o[13], o[14], o[15]));
}
if o[0] == 0x20 && o[1] == 0x02 {
return Some(Ipv4Addr::new(o[2], o[3], o[4], o[5]));
}
if o[0] == 0x20 && o[1] == 0x01 && o[2] == 0x00 && o[3] == 0x00 {
// Teredo obfuscates the client IPv4 as its ones' complement.
return Some(Ipv4Addr::new(
o[12] ^ 0xff,
o[13] ^ 0xff,
o[14] ^ 0xff,
o[15] ^ 0xff,
));
}
if o[..12].iter().all(|&b| b == 0) {
// IPv4-compatible ::a.b.c.d; :: and ::1 are already handled by the
// unspecified / loopback checks before extraction runs.
return Some(Ipv4Addr::new(o[12], o[13], o[14], o[15]));
}
None
}
/// Returns an error if `ip` is not safe to dial from a server that can reach
/// cluster-internal hosts. Mirrors Go's `checkBlockedIP`.
pub fn check_blocked_ip(endpoint: &str, ip: IpAddr) -> Result<(), String> {
check_blocked_ip_policy(endpoint, ip, false)
}
/// Like [`check_blocked_ip`], but `allow_private` keeps RFC 1918 / CGNAT
/// reachable for callers whose target legitimately sits on an internal network
/// (peer volume servers), while still blocking loopback, link-local (IMDS) and
/// unspecified. Mirrors Go's `checkBlockedIPPolicy`.
pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool) -> Result<(), String> {
// Normalize IPv4-mapped IPv6 (`::ffff:a.b.c.d`) to its IPv4 form so the
// IPv4 deny rules apply. The OS routes these to the embedded IPv4 address,
// so without this `::ffff:127.0.0.1` / `::ffff:169.254.169.254` would slip
@@ -155,28 +112,17 @@ pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool)
endpoint, ip
));
}
if !allow_private {
if is_private(ip) {
return Err(format!(
"remote endpoint {:?} resolves to private address {}",
endpoint, ip
));
}
if is_cgnat(ip) {
return Err(format!(
"remote endpoint {:?} resolves to CGNAT address {}",
endpoint, ip
));
}
if is_private(ip) {
return Err(format!(
"remote endpoint {:?} resolves to private address {}",
endpoint, ip
));
}
// IPv6 transition addresses embed an IPv4 destination that routes to the
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
if is_cgnat(ip) {
return Err(format!(
"remote endpoint {:?} resolves to CGNAT address {}",
endpoint, ip
));
}
Ok(())
}
@@ -297,59 +243,6 @@ pub async fn validate_remote_endpoint(endpoint: &str) -> Result<(), String> {
}
}
/// Returns an error if `target` could redirect a replica upload away from a peer
/// volume server. The target must be a bare `host:port` -- a scheme, userinfo,
/// path, query or fragment can smuggle a different destination into the
/// formatted upload URL -- whose host is not loopback, link-local (IMDS) or
/// unspecified. Cluster peers legitimately sit on private networks, so RFC 1918
/// / CGNAT are allowed. Mirrors Go's `validateReplicaTarget`.
pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let trimmed = target.trim();
if trimmed.is_empty() {
return Err("replica target is empty".to_string());
}
if trimmed.contains("://") || trimmed.contains(['/', '?', '#', '@', '\\']) {
return Err(format!("replica target {:?} must be a bare host:port", target));
}
// Require an explicit host:port, handling `[IPv6]:port`. A bracketless IPv6
// literal (which carries its own colons) is rejected; peers are addressed as
// `[ipv6]:port`, matching Go's net.SplitHostPort.
let host = if let Some(rest) = trimmed.strip_prefix('[') {
match rest.split_once(']') {
Some((h, port)) if port.starts_with(':') && port.len() > 1 => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
}
} else {
match trimmed.rsplit_once(':') {
Some((h, port)) if !port.is_empty() && !h.contains(':') => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
}
};
if host.is_empty() {
return Err(format!("replica target {:?} has no host", target));
}
if is_blocked_imds_host(&host.to_ascii_lowercase()) {
return Err(format!(
"replica target {:?} targets instance metadata service",
target
));
}
if let Ok(ip) = host.parse::<IpAddr>() {
return check_blocked_ip_policy(target, ip, true);
}
let addrs = resolve_host(host).await?;
if addrs.is_empty() {
return Err(format!("resolve replica target host {:?}: no addresses", host));
}
for ip in addrs {
check_blocked_ip_policy(target, ip, true)?;
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
@@ -440,49 +333,6 @@ mod tests {
assert!(check_blocked_ip("e", ip("2606:4700:4700::1111")).is_ok());
}
#[test]
fn rejects_ipv6_transition_addresses() {
// Every transition form encoding an internal IPv4 must be blocked.
assert!(
check_blocked_ip("e", ip("64:ff9b::a9fe:a9fe")) // NAT64 -> IMDS
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip("e", ip("64:ff9b::7f00:1")) // NAT64 -> loopback
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("2002:a00:1::")) // 6to4 -> 10.0.0.1
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("2001:0:4136:e378:8000:63bf:80ff:fffe")) // Teredo -> loopback
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("::7f00:1")) // IPv4-compatible -> loopback
.unwrap_err()
.contains("loopback")
);
// A NAT64 address outside the 64:ff9b::/96 well-known prefix is not
// decoded (its embedded IPv4 lives elsewhere), so it is left as-is.
assert!(check_blocked_ip("e", ip("64:ff9b:1::a9fe:a9fe")).is_ok());
// Every transition form embedding a public IPv4 (8.8.8.8) still passes:
// NAT64, 6to4, Teredo, IPv4-compatible.
assert!(check_blocked_ip("e", ip("64:ff9b::808:808")).is_ok());
assert!(check_blocked_ip("e", ip("2002:808:808::")).is_ok());
assert!(check_blocked_ip("e", ip("2001::f7f7:f7f7")).is_ok());
assert!(check_blocked_ip("e", ip("::808:808")).is_ok());
// Bracketed transition literal via the full endpoint path.
assert!(precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata"));
}
#[test]
fn rejects_ipv4_mapped_ipv6() {
// IPv4-mapped IPv6 must be unmapped so the IPv4 rules catch it.
@@ -506,71 +356,4 @@ mod tests {
.unwrap_err()
.contains("loopback"));
}
#[test]
fn check_blocked_ip_policy_allows_private_peers() {
// The replica leg targets peer volume servers, which may be private.
assert!(check_blocked_ip_policy("e", ip("10.0.0.7"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("192.168.1.5"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("100.64.0.42"), true).is_ok());
// Loopback / IMDS / unspecified stay blocked even when private is allowed.
assert!(check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified"));
}
#[tokio::test]
async fn validate_replica_target_rejects_and_allows() {
// A path plus a trailing ?a= would otherwise swallow ?type=replicate.
assert!(validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("")
.await
.unwrap_err()
.contains("empty"));
// Legitimate peer volume servers on private networks pass.
assert!(validate_replica_target("10.0.0.7:8080").await.is_ok());
assert!(validate_replica_target("192.168.1.5:8080").await.is_ok());
assert!(validate_replica_target("[fd00::1]:8080").await.is_ok());
}
}
+1 -33
View File
@@ -7,7 +7,7 @@ pub mod endpoint_guard;
pub mod s3;
pub mod s3_tier;
pub use endpoint_guard::{validate_remote_endpoint, validate_replica_target};
pub use endpoint_guard::validate_remote_endpoint;
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
@@ -211,36 +211,4 @@ mod tests {
};
assert_eq!(s3_compatible_endpoint(&gcs), None);
}
#[test]
fn azure_endpoint_has_no_ssrf_path() {
// The Go volume server guards the caller-supplied azure endpoint against
// SSRF. This server has no azure backend, so there is nothing to dial:
// azure is not S3-compatible (the endpoint guard does not apply) and
// make_remote_storage_client rejects the type before building a client.
let azure = RemoteConf {
r#type: "azure".to_string(),
azure_endpoint: "https://169.254.169.254/".to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&azure), None);
assert!(make_remote_storage_client(&azure).is_err());
}
#[test]
fn gcs_credentials_have_no_ssrf_path() {
// The Go volume server accepts only static-key gcs credentials and puts
// their token endpoint behind the SSRF guard, because the SDK dials
// whatever url, file or executable the credentials name. This server has
// no gcs backend, so make_remote_storage_client rejects the type before
// any credentials are parsed. Anyone adding one must carry both guards
// over with it.
let gcs = RemoteConf {
r#type: "gcs".to_string(),
gcs_google_application_credentials: r#"{"type":"external_account","credential_source":{"url":"http://169.254.169.254/"}}"#.to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&gcs), None);
assert!(make_remote_storage_client(&gcs).is_err());
}
}
-7
View File
@@ -192,13 +192,6 @@ impl Guard {
self.is_write_active = !is_empty_whitelist || !self.signing_key.is_empty();
}
/// Whether any write-side control is configured: a non-empty whitelist or a
/// signing key. Mirrors Go's non-nil `guard` -- when this is false, every
/// caller is allowed and admin RPCs need not carry peer info.
pub fn is_write_active(&self) -> bool {
self.is_write_active
}
/// Check if a remote IP is in the whitelist.
/// Returns true if write security is inactive (no whitelist and no signing key),
/// if the whitelist is empty, or if the IP matches.
+7 -56
View File
@@ -33,31 +33,23 @@ impl Error for GrpcClientError {}
pub fn load_outgoing_grpc_tls(
config: &VolumeServerConfig,
) -> Result<Option<OutgoingGrpcTlsConfig>, GrpcClientError> {
// prefer a dedicated client certificate: CAs may issue certs with only one of the serverAuth/clientAuth EKUs
let (cert_file, key_file) = if !config.grpc_client_cert_file.is_empty()
&& !config.grpc_client_key_file.is_empty()
if config.grpc_cert_file.is_empty()
|| config.grpc_key_file.is_empty()
|| config.grpc_ca_file.is_empty()
{
(&config.grpc_client_cert_file, &config.grpc_client_key_file)
} else {
if !config.grpc_client_cert_file.is_empty() || !config.grpc_client_key_file.is_empty() {
tracing::warn!("grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key");
}
(&config.grpc_cert_file, &config.grpc_key_file)
};
if cert_file.is_empty() || key_file.is_empty() || config.grpc_ca_file.is_empty() {
return Ok(None);
}
let cert_pem = std::fs::read_to_string(cert_file).map_err(|e| {
let cert_pem = std::fs::read_to_string(&config.grpc_cert_file).map_err(|e| {
GrpcClientError(format!(
"Failed to read outgoing gRPC cert '{}': {}",
cert_file, e
config.grpc_cert_file, e
))
})?;
let key_pem = std::fs::read_to_string(key_file).map_err(|e| {
let key_pem = std::fs::read_to_string(&config.grpc_key_file).map_err(|e| {
GrpcClientError(format!(
"Failed to read outgoing gRPC key '{}': {}",
key_file, e
config.grpc_key_file, e
))
})?;
let ca_pem = std::fs::read_to_string(&config.grpc_ca_file).map_err(|e| {
@@ -244,8 +236,6 @@ mod tests {
https_client_ca_file: String::new(),
grpc_cert_file: String::new(),
grpc_key_file: String::new(),
grpc_client_cert_file: String::new(),
grpc_client_key_file: String::new(),
grpc_ca_file: String::new(),
grpc_allowed_wildcard_domain: String::new(),
grpc_volume_allowed_common_names: vec![],
@@ -275,45 +265,6 @@ mod tests {
assert!(load_outgoing_grpc_tls(&config).unwrap().is_none());
}
fn write_pem_files(dir: &tempfile::TempDir, config: &mut VolumeServerConfig) {
let write = |name: &str, content: &str| {
let path = dir.path().join(name);
std::fs::write(&path, content).unwrap();
path.to_str().unwrap().to_string()
};
config.grpc_cert_file = write("server.pem", "server-cert");
config.grpc_key_file = write("server.key", "server-key");
config.grpc_ca_file = write("ca.pem", "ca");
}
#[test]
fn test_load_outgoing_grpc_tls_prefers_client_cert() {
let dir = tempfile::TempDir::new().unwrap();
let mut config = sample_config();
write_pem_files(&dir, &mut config);
let client_cert = dir.path().join("client.pem");
let client_key = dir.path().join("client.key");
std::fs::write(&client_cert, "client-cert").unwrap();
std::fs::write(&client_key, "client-key").unwrap();
config.grpc_client_cert_file = client_cert.to_str().unwrap().to_string();
config.grpc_client_key_file = client_key.to_str().unwrap().to_string();
let tls = load_outgoing_grpc_tls(&config).unwrap().unwrap();
assert_eq!(tls.cert_pem, "client-cert");
assert_eq!(tls.key_pem, "client-key");
}
#[test]
fn test_load_outgoing_grpc_tls_falls_back_to_server_cert() {
let dir = tempfile::TempDir::new().unwrap();
let mut config = sample_config();
write_pem_files(&dir, &mut config);
let tls = load_outgoing_grpc_tls(&config).unwrap().unwrap();
assert_eq!(tls.cert_pem, "server-cert");
assert_eq!(tls.key_pem, "server-key");
}
#[test]
fn test_build_grpc_endpoint_without_tls_uses_http_scheme() {
let endpoint = build_grpc_endpoint("127.0.0.1:19333", None).unwrap();
+72 -524
View File
@@ -38,7 +38,6 @@ fn scrub_mode_label(mode: i32) -> &'static str {
2 => "FULL",
3 => "LOCAL",
4 => "CHECKSUM",
5 => "READS",
_ => "UNKNOWN",
}
}
@@ -92,57 +91,6 @@ pub fn load_state_file(
volume_server_pb::VolumeServerState::decode(data.as_slice()).ok()
}
/// One disk location's stake in an EC volume, as seen by the rebuild handler.
struct LocInfo {
dir: String,
idx_dir: String,
shard_count: usize,
has_ecx: bool,
}
/// Picks the location a rebuild should write into — the one holding an `.ecx`
/// and the most shards — and returns every other directory it may have to read
/// from. Shards are only half of what the rebuild needs: a split
/// `-dir`/`-dir.idx` layout keeps `.ecx`/`.ecj`/`.vif` with the INDEX, and on a
/// multi-disk server the chosen disk may hold nothing but shards while this
/// volume's `.vif` or generation-0 `.ecsum` sits on a sibling. Miss those and
/// the layout resolution falls back to 10+4 with the legacy striping and
/// reconstructs through the wrong matrix, so both directories of every other
/// location are listed. The rebuild's own two are passed separately by the
/// caller and dropped here, along with empties and duplicates.
///
/// Returns `None` when no location holds an `.ecx`, i.e. there is nothing to
/// rebuild from.
fn select_rebuild_location(loc_infos: &[LocInfo]) -> Option<(usize, Vec<String>)> {
let mut rebuild_loc_idx: Option<usize> = None;
let mut other_dirs: Vec<String> = Vec::new();
for (i, info) in loc_infos.iter().enumerate() {
let better = info.has_ecx
&& rebuild_loc_idx
.is_none_or(|prev| info.shard_count > loc_infos[prev].shard_count);
if better {
if let Some(prev) = rebuild_loc_idx {
other_dirs.push(loc_infos[prev].dir.clone());
other_dirs.push(loc_infos[prev].idx_dir.clone());
}
rebuild_loc_idx = Some(i);
} else {
other_dirs.push(info.dir.clone());
other_dirs.push(info.idx_dir.clone());
}
}
let rebuild_loc_idx = rebuild_loc_idx?;
let rebuild_dir = &loc_infos[rebuild_loc_idx].dir;
let rebuild_idx_dir = &loc_infos[rebuild_loc_idx].idx_dir;
other_dirs.retain(|d| !d.is_empty() && d != rebuild_dir && d != rebuild_idx_dir);
other_dirs.sort();
other_dirs.dedup();
Some((rebuild_loc_idx, other_dirs))
}
struct WriteThrottler {
bytes_per_second: i64,
last_size_counter: i64,
@@ -207,13 +155,6 @@ impl VolumeGrpcService {
/// SocketAddr; if it is somehow None we deny, matching "if we don't
/// know who the caller is, refuse."
fn check_grpc_admin_auth<T>(&self, request: &Request<T>) -> Result<(), Status> {
// Mirror Go's `if vs.guard == nil { return nil }`: with no whitelist and
// no signing key, write security is inactive and every caller is allowed.
// Real gRPC connections always carry peer info, so requiring it below only
// affects in-process callers (tests, upgrades) once a control is enabled.
if !self.state.guard.read().unwrap().is_write_active() {
return Ok(());
}
let remote = match request.remote_addr() {
Some(addr) => addr,
None => {
@@ -296,17 +237,12 @@ impl VolumeGrpcService {
Ok(())
}
/// Shared helper matching Go's `makeVolumeReadonly(ctx, v, canDelete, persist)`.
/// Shared helper matching Go's `makeVolumeReadonly(ctx, v, persist)`.
/// 1. Check maintenance mode
/// 2. Notify master (readonly=true)
/// 3. Mark local volume readonly
/// 4. Notify master again (cover heartbeat race)
async fn make_volume_readonly(
&self,
vid: VolumeId,
can_delete: bool,
persist: bool,
) -> Result<(), Status> {
async fn make_volume_readonly(&self, vid: VolumeId, persist: bool) -> Result<(), Status> {
self.state.check_maintenance()?;
let info = {
@@ -332,7 +268,7 @@ impl VolumeGrpcService {
{
let mut store = self.state.store.write().unwrap();
if let Some((_, vol)) = store.find_volume_mut(vid) {
vol.set_read_only_persist(can_delete, persist)
vol.set_read_only_persist(persist)
.map_err(|e| Status::internal(e.to_string()))?;
}
self.state.volume_state_notify.notify_one();
@@ -352,7 +288,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::BatchDeleteRequest>,
) -> Result<Response<volume_server_pb::BatchDeleteResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let mut results = Vec::new();
@@ -1004,22 +939,6 @@ impl VolumeServer for VolumeGrpcService {
Ok(Response::new(volume_server_pb::VolumeUnmountResponse {}))
}
async fn volume_consolidate_index(
&self,
request: Request<volume_server_pb::VolumeConsolidateIndexRequest>,
) -> Result<Response<volume_server_pb::VolumeConsolidateIndexResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let vid = VolumeId(request.into_inner().volume_id);
let mut store = self.state.store.write().unwrap();
store
.consolidate_volume_index(vid)
.map_err(|e| Status::internal(e.to_string()))?;
Ok(Response::new(
volume_server_pb::VolumeConsolidateIndexResponse {},
))
}
async fn volume_delete(
&self,
request: Request<volume_server_pb::VolumeDeleteRequest>,
@@ -1058,8 +977,7 @@ impl VolumeServer for VolumeGrpcService {
.find_volume(vid)
.ok_or_else(|| Status::not_found(format!("volume {} not found", vid)))?;
}
self.make_volume_readonly(vid, req.can_delete, req.persist)
.await?;
self.make_volume_readonly(vid, req.persist).await?;
Ok(Response::new(
volume_server_pb::VolumeMarkReadonlyResponse {},
))
@@ -1211,7 +1129,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::SetStateRequest>,
) -> Result<Response<volume_server_pb::SetStateResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
if let Some(new_state) = &req.state {
@@ -1264,7 +1181,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeCopyRequest>,
) -> Result<Response<Self::VolumeCopyStream>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -1820,14 +1736,11 @@ impl VolumeServer for VolumeGrpcService {
}
Some(store.locations[info.disk_id as usize].directory.clone())
} else {
// The mounted-volume refusal above means no disk
// holds an in-memory claim here.
store
.find_ec_shard_target_location(
&info.collection,
vid,
DATA_SHARDS_COUNT as u32,
&[],
)
.map(|i| store.locations[i].directory.clone())
};
@@ -2069,7 +1982,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::ReadAllNeedlesRequest>,
) -> Result<Response<Self::ReadAllNeedlesStream>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let state = self.state.clone();
@@ -2270,7 +2182,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeTailReceiverRequest>,
) -> Result<Response<volume_server_pb::VolumeTailReceiverResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2381,7 +2292,7 @@ impl VolumeServer for VolumeGrpcService {
// Write needle to local volume
let mut store = state.store.write().unwrap();
store
.write_volume_needle(vid, &mut n, false)
.write_volume_needle(vid, &mut n)
.map_err(|e| Status::internal(format!("write needle: {}", e)))?;
}
@@ -2396,7 +2307,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsGenerateRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsGenerateResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2437,21 +2347,13 @@ impl VolumeServer for VolumeGrpcService {
)
};
// Check existing .vif for EC shard config (matching Go's MaybeLoadVolumeInfo).
// The block size is recomputed by the encode for the current .dat, so
// only the ratio is carried over from a prior config.
let (data_shards, parity_shards, _) =
// Check existing .vif for EC shard config (matching Go's MaybeLoadVolumeInfo)
let (data_shards, parity_shards) =
crate::storage::erasure_coding::ec_volume::read_ec_shard_config(
&dir, &idx_dir, collection, vid,
)
.map_err(|e| {
tonic::Status::internal(format!(
"read ec shard config for volume {}: {}",
vid.0, e
))
})?;
);
let block_size = match crate::storage::erasure_coding::ec_encoder::write_ec_files(
if let Err(e) = crate::storage::erasure_coding::ec_encoder::write_ec_files(
&dir,
&idx_dir,
collection,
@@ -2459,19 +2361,16 @@ impl VolumeServer for VolumeGrpcService {
data_shards as usize,
parity_shards as usize,
) {
Ok(block_size) => block_size,
Err(e) => {
// Cleanup partially-created .ecNN and .ecx files on failure (matching Go defer)
let base = crate::storage::volume::volume_file_name(&dir, collection, vid);
let total_shards = data_shards + parity_shards;
for i in 0..total_shards {
let shard_path = format!("{}.ec{:02}", base, i);
let _ = std::fs::remove_file(&shard_path);
}
let _ = std::fs::remove_file(format!("{}.ecx", base));
return Err(Status::internal(e.to_string()));
// Cleanup partially-created .ecNN and .ecx files on failure (matching Go defer)
let base = crate::storage::volume::volume_file_name(&dir, collection, vid);
let total_shards = data_shards + parity_shards;
for i in 0..total_shards {
let shard_path = format!("{}.ec{:02}", base, i);
let _ = std::fs::remove_file(&shard_path);
}
};
let _ = std::fs::remove_file(format!("{}.ecx", base));
return Err(Status::internal(e.to_string()));
}
// Write .vif file with EC shard metadata
{
@@ -2490,7 +2389,6 @@ impl VolumeServer for VolumeGrpcService {
.duration_since(std::time::UNIX_EPOCH)
.unwrap_or_default()
.as_nanos() as i64,
block_size,
}),
..Default::default()
};
@@ -2509,7 +2407,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsRebuildRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsRebuildResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2524,6 +2421,13 @@ impl VolumeServer for VolumeGrpcService {
format!("{}_{}", collection, vid.0)
};
struct LocInfo {
dir: String,
idx_dir: String,
shard_count: usize,
has_ecx: bool,
}
let store = self.state.store.read().unwrap();
let mut loc_infos: Vec<LocInfo> = Vec::new();
@@ -2571,8 +2475,26 @@ impl VolumeServer for VolumeGrpcService {
));
}
let (rebuild_loc_idx, other_dirs) = match select_rebuild_location(&loc_infos) {
Some(picked) => picked,
// Pick rebuild location: has .ecx and most shards
let mut rebuild_loc_idx: Option<usize> = None;
let mut other_dirs: Vec<String> = Vec::new();
for (i, info) in loc_infos.iter().enumerate() {
if info.has_ecx
&& (rebuild_loc_idx.is_none()
|| info.shard_count > loc_infos[rebuild_loc_idx.unwrap()].shard_count)
{
if let Some(prev) = rebuild_loc_idx {
other_dirs.push(loc_infos[prev].dir.clone());
}
rebuild_loc_idx = Some(i);
} else {
other_dirs.push(info.dir.clone());
}
}
let rebuild_loc_idx = match rebuild_loc_idx {
Some(i) => i,
None => {
return Ok(Response::new(
volume_server_pb::VolumeEcShardsRebuildResponse {
@@ -2586,37 +2508,13 @@ impl VolumeServer for VolumeGrpcService {
let rebuild_idx_dir = loc_infos[rebuild_loc_idx].idx_dir.clone();
// Determine data/parity shard config from rebuild dir
// The encode-time .dat size resolves the row count the ecx rebuild
// de-stripes with; 0 leaves it to infer from the padded shard extent.
// Both lookups search the sibling disks too: the rebuild writes into one
// location, but a multi-disk server may keep this volume's .vif or its
// generation-0 .ecsum on another, and defaulting to 10+4 with the
// legacy layout would reconstruct through the wrong matrix.
let dat_file_size = crate::storage::erasure_coding::ec_volume::load_vif_info_across_dirs(
&rebuild_dir,
&rebuild_idx_dir,
&other_dirs,
collection,
vid,
)
.ok()
.flatten()
.map(|(v, _)| v.dat_file_size)
.unwrap_or(0);
let (data_shards, parity_shards, block_size) =
crate::storage::erasure_coding::ec_volume::read_ec_shard_config_across_dirs(
let (data_shards, parity_shards) =
crate::storage::erasure_coding::ec_volume::read_ec_shard_config(
&rebuild_dir,
&rebuild_idx_dir,
&other_dirs,
collection,
vid,
)
.map_err(|e| {
tonic::Status::internal(format!(
"read ec shard config for volume {}: {}",
vid.0, e
))
})?;
);
let total_shards = data_shards + parity_shards;
// Check which shards are missing (check rebuild dir and all other dirs)
@@ -2655,17 +2553,8 @@ impl VolumeServer for VolumeGrpcService {
// Rebuild missing shards, searching all locations for input shards.
// Pass other_dirs so shards on sibling disks are found even when the
// primary rebuild dir doesn't hold them. This one takes a single flat
// list — the shape Go's RebuildEcFiles uses — so unlike the resolvers
// above it cannot be handed the rebuild's own index directory
// separately, and a split -dir/-dir.idx location keeps its .ecx and
// .vif there. Go's additionalDirs carries that directory for the same
// reason.
let mut rebuild_search_dirs: Vec<String> = other_dirs.clone();
if !rebuild_idx_dir.is_empty() && rebuild_idx_dir != rebuild_dir {
rebuild_search_dirs.push(rebuild_idx_dir.clone());
}
let other_dir_refs: Vec<&str> = rebuild_search_dirs.iter().map(|s| s.as_str()).collect();
// primary rebuild dir doesn't hold them.
let other_dir_refs: Vec<&str> = other_dirs.iter().map(|s| s.as_str()).collect();
crate::storage::erasure_coding::ec_encoder::rebuild_ec_files(
&rebuild_dir,
collection,
@@ -2703,8 +2592,6 @@ impl VolumeServer for VolumeGrpcService {
collection,
vid,
data_shards as usize,
block_size,
dat_file_size,
&ecx_dir_refs,
)
.map_err(|e| Status::internal(format!("RebuildEcxFile: {}", e)))?;
@@ -2720,7 +2607,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsCopyRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsCopyResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2729,15 +2615,13 @@ impl VolumeServer for VolumeGrpcService {
// When disk_id > 0: use that specific location.
// When disk_id == 0 (unset): auto-select via
// find_ec_shard_target_location, which prefers a disk that
// already owns one of the shards being copied (a retried move
// must overwrite in place, not leave two disks of this server
// claiming the same shard), then a disk that already has the EC
// volume mounted, then a disk that owns the .ecx on disk (volume
// not yet mounted — relevant for ec.rebuild, where only the
// first shard carries .ecx and subsequent shards must land on
// the same disk; see #9212), then any HDD, then any disk. Pass
// the build's default data-shard count; the helper takes it as a
// parameter so custom-ratio builds can swap it.
// already has the EC volume mounted, then a disk that owns the
// .ecx on disk (volume not yet mounted — relevant for
// ec.rebuild, where only the first shard carries .ecx and
// subsequent shards must land on the same disk; see #9212),
// then any HDD, then any disk. Pass the build's default
// data-shard count; the helper takes it as a parameter so
// custom-ratio builds can swap it.
let (dest_dir, dest_idx_dir) = {
let store = self.state.store.read().unwrap();
let count = store.locations.len();
@@ -2753,27 +2637,10 @@ impl VolumeServer for VolumeGrpcService {
let loc = &store.locations[req.disk_id as usize];
(loc.directory.clone(), loc.idx_directory.clone())
} else {
// A batch whose requested shards are already owned by
// different local disks has no single correct destination:
// writing them all to one disk would duplicate the other
// disks' claims. Refuse so the caller splits the batch per
// shard (or chooses explicitly via disk_id).
let owners = store.ec_shard_owner_disks(vid, &req.shard_ids);
if owners.len() > 1 {
let dirs: Vec<&str> = owners
.iter()
.map(|&i| store.locations[i].directory.as_str())
.collect();
return Err(Status::failed_precondition(format!(
"volume {} shards {:?} are already owned by multiple local disks {:?}: no single destination; copy per shard or pass disk_id",
req.volume_id, req.shard_ids, dirs
)));
}
match store.find_ec_shard_target_location(
&req.collection,
vid,
DATA_SHARDS_COUNT as u32,
&req.shard_ids,
) {
Some(i) => {
let loc = &store.locations[i];
@@ -3110,27 +2977,6 @@ impl VolumeServer for VolumeGrpcService {
Status::internal(format!("mount {}.{}: {}", req.volume_id, shard_id, e))
})?;
}
// A delivery can bring the checksum manifest alongside the shards, but
// the receive path only writes the file. When this server already had
// the volume mounted, the EcVolume in memory keeps whatever protection
// state it resolved at mount — off, for a volume whose sidecar arrives
// now — until a remount. Re-resolve it here, where the shards it
// describes have just been added.
//
// Every per-disk runtime, not just the first: a vid mounts as one
// EcVolume per disk, the delivery lands the .ecsum on one of them, and
// the first-match lookup would leave the siblings reporting no
// protection. Each re-resolves against its own data and index
// directories, so a shared -dir.idx reaches all of them.
// Resolving across every EC metadata directory is what makes that
// reload mean something: startup mirroring gives each shard-bearing
// disk its own .ecx/.ecj/.vif but deliberately not the sidecar, so a
// runtime restricted to its own two directories would find nothing
// however often it reloaded. One delivered copy, reachable from all.
let ec_metadata_dirs = store.ec_metadata_dirs();
for ec_vol in store.find_all_ec_volumes_mut(vid) {
ec_vol.reload_bitrot_sidecar(&ec_metadata_dirs);
}
drop(store);
self.state.volume_state_notify.notify_one();
@@ -3143,7 +2989,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsUnmountRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsUnmountResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3307,7 +3152,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsToVolumeRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsToVolumeResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3465,8 +3309,6 @@ impl VolumeServer for VolumeGrpcService {
let ecx_dir = ec_vol.ecx_actual_dir().to_string();
let collection = ec_vol.collection.clone();
let vif_dat_file_size = ec_vol.dat_file_size;
let (large_block_size, small_block_size) =
(ec_vol.large_block_size(), ec_vol.small_block_size());
// shard_dirs[i] is guaranteed Some for i in 0..data_shards by
// the check above; collect concrete dirs for the decoder.
let per_shard_dirs: Vec<String> = shard_dirs[..data_shards]
@@ -3500,8 +3342,6 @@ impl VolumeServer for VolumeGrpcService {
vif_dat_file_size,
data_shards,
&per_shard_dirs,
large_block_size as usize,
small_block_size as usize,
)
.map_err(|e| Status::internal(format!("WriteDatFile: {}", e)))?;
@@ -3560,24 +3400,12 @@ impl VolumeServer for VolumeGrpcService {
.walk_ecx_stats()
.map_err(|e| Status::internal(e.to_string()))?;
// The layout this holder serves reads through, as Go reports it: a
// coordinator cannot otherwise tell a holder that understands the
// uniform block layout from one that dropped the unknown .vif field
// and mounted the volume as legacy.
let ec_shard_config = Some(volume_server_pb::EcShardConfig {
data_shards: ec_vol.data_shards,
parity_shards: ec_vol.parity_shards,
encode_ts_ns: 0,
block_size: ec_vol.block_size,
});
Ok(Response::new(
volume_server_pb::VolumeEcShardsInfoResponse {
ec_shard_infos: shard_infos,
volume_size,
file_count,
file_deleted_count,
ec_shard_config,
},
))
}
@@ -3590,7 +3418,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeTierMoveDatToRemoteRequest>,
) -> Result<Response<Self::VolumeTierMoveDatToRemoteStream>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3708,12 +3535,7 @@ impl VolumeServer for VolumeGrpcService {
modified_time: dat_modified_secs,
extension: ".dat".to_string(),
});
vol.refresh_remote_write_mode().map_err(|e| {
Status::internal(format!(
"volume {} failed to refresh write mode: {}",
vid, e
))
})?;
vol.refresh_remote_write_mode();
if let Err(e) = vol.save_volume_info() {
return Err(Status::internal(format!(
@@ -3762,7 +3584,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeTierMoveDatFromRemoteRequest>,
) -> Result<Response<Self::VolumeTierMoveDatFromRemoteStream>, Status> {
self.check_grpc_admin_auth(&request)?;
// Note: Go does NOT check maintenance mode for TierMoveDatFromRemote
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3899,40 +3720,10 @@ impl VolumeServer for VolumeGrpcService {
)));
}
// Snapshot the remote reference before dropping it: the
// refresh below can fail, and a half-applied transition
// leaves the volume claiming local while the remote backend
// is still attached and the on-disk .vif still says remote
// — a state a retry reads as "already on local disk" and
// refuses to finish.
let removed_remote = if vol.volume_info.files.is_empty() {
None
} else {
Some(vol.volume_info.files.remove(0))
};
// Swaps the read-only sorted map out before the volume is
// published as writable; without it the first write would
// append to the local .dat and then fail to index.
if let Err(e) = vol.refresh_remote_write_mode() {
if let Some(remote) = removed_remote {
vol.volume_info.files.insert(0, remote);
}
// Put the derived flags and the needle map back where
// the restored reference says they belong. Best effort:
// if even this fails the volume stays pinned read-only,
// which is the safe end of the transition.
if let Err(restore_err) = vol.refresh_remote_write_mode() {
tracing::warn!(
volume_id = vid.0,
error = %restore_err,
"tier-down rollback could not restore the remote write mode",
);
}
return Err(Status::internal(format!(
"volume {} failed to refresh write mode: {}",
vid, e
)));
if !vol.volume_info.files.is_empty() {
vol.volume_info.files.remove(0);
}
vol.refresh_remote_write_mode();
if let Err(e) = vol.save_volume_info() {
return Err(Status::internal(format!(
@@ -4053,7 +3844,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::FetchAndWriteNeedleRequest>,
) -> Result<Response<volume_server_pb::FetchAndWriteNeedleResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -4129,21 +3919,6 @@ impl VolumeServer for VolumeGrpcService {
.as_secs();
n.set_has_last_modified_date();
// Validate every replica target before writing anything, so a malformed
// or internal target fails the request instead of leaving a local write
// behind. Mirrors the Go volume server. NOTE: like the Rust S3 endpoint
// guard, this validates the up-front DNS answer but does not yet re-check
// at connect time, so a rebinding hostname remains a follow-up.
if !self.state.allow_untrusted_remote_endpoints {
for replica in &req.replicas {
crate::remote_storage::validate_replica_target(&replica.url)
.await
.map_err(|e| {
Status::invalid_argument(format!("reject replica target: {}", e))
})?;
}
}
// Run local write and replica writes concurrently (matches Go's WaitGroup)
let mut handles: Vec<tokio::task::JoinHandle<Result<(), String>>> = Vec::new();
@@ -4155,7 +3930,7 @@ impl VolumeServer for VolumeGrpcService {
let local_handle = tokio::task::spawn_blocking(move || {
let mut store = state_clone.store.write().unwrap();
store
.write_volume_needle(vid, &mut n_clone, false)
.write_volume_needle(vid, &mut n_clone)
.map(|_| ())
.map_err(|e| format!("local write needle {} size {}: {}", needle_id, size, e))
});
@@ -4230,7 +4005,7 @@ impl VolumeServer for VolumeGrpcService {
// Validate mode
let mode = req.mode;
match mode {
1 | 2 | 3 | 5 => {} // INDEX=1, FULL=2, LOCAL=3, READS=5 (FULL for regular volumes)
1 | 2 | 3 => {} // INDEX=1, FULL=2, LOCAL=3
_ => {
return Err(Status::invalid_argument(format!(
"unsupported volume scrub mode {}",
@@ -4292,7 +4067,7 @@ impl VolumeServer for VolumeGrpcService {
let mut errs: Vec<String> = Vec::new();
if req.mark_broken_volumes_readonly {
for vid in &broken_vids {
match self.make_volume_readonly(*vid, false, true).await {
match self.make_volume_readonly(*vid, true).await {
Ok(()) => {
details.push(format!("volume {} is now read-only", vid.0));
}
@@ -4324,13 +4099,12 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::ScrubEcVolumeRequest>,
) -> Result<Response<volume_server_pb::ScrubEcVolumeResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
// Validate mode
let mode = req.mode;
match mode {
1 | 2 | 3 | 4 | 5 => {} // INDEX=1, FULL=2, LOCAL=3, CHECKSUM=4, READS=5
1 | 2 | 3 | 4 => {} // INDEX=1, FULL=2, LOCAL=3, CHECKSUM=4
_ => {
return Err(Status::invalid_argument(format!(
"unsupported EC volume scrub mode {}",
@@ -4339,14 +4113,6 @@ impl VolumeServer for VolumeGrpcService {
}
}
// Only the modes that walk needles can be strict about deleted ones.
let force_deleted_needles_check = req.force_deleted_needles_check;
if force_deleted_needles_check && mode != 2 && mode != 5 {
return Err(Status::invalid_argument(
"deleted needle checks are only supported for FULL and READS scrubs",
));
}
// Collect the volume ids under a brief lock, then release it: FULL (mode 2)
// reads remote shards and must not hold the !Send store guard across .await.
let vids: Vec<VolumeId> = {
@@ -4388,8 +4154,8 @@ impl VolumeServer for VolumeGrpcService {
}
}
}
2 | 5 => {
// FULL/READS: Go-parity per-needle local+remote walk, PLUS a TEMPORARY
2 => {
// FULL: Go-parity per-needle local+remote walk, PLUS a TEMPORARY
// local Reed-Solomon parity check. The needle walk only reads
// DATA-shard intervals of LIVE needles, so on its own it can't
// catch silent bitrot in a PARITY shard or an unwalked cold
@@ -4421,13 +4187,8 @@ impl VolumeServer for VolumeGrpcService {
// (1) Per-needle local+remote walk (Go ScrubEcVolume parity).
let (files, mut shard_infos, mut errs) =
crate::server::store_ec::scrub_ec_volume_distributed(
&self.state,
vid,
force_deleted_needles_check,
mode == 5,
)
.await;
crate::server::store_ec::scrub_ec_volume_distributed(&self.state, vid, false)
.await;
total_files += files as u64; // count comes from the needle walk only
// (2) Local parity check, gated on all-shards-local. Blocking RS
@@ -4695,7 +4456,6 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeNeedleStatusRequest>,
) -> Result<Response<volume_server_pb::VolumeNeedleStatusResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
let needle_id = NeedleId(req.needle_id);
@@ -5218,110 +4978,6 @@ mod tests {
use tempfile::TempDir;
use tokio_stream::StreamExt;
fn loc(dir: &str, idx_dir: &str, shard_count: usize, has_ecx: bool) -> LocInfo {
LocInfo {
dir: dir.to_string(),
idx_dir: idx_dir.to_string(),
shard_count,
has_ecx,
}
}
// The rebuild reads its shards from one directory but resolves the volume's
// layout -- ratio and uniform block size -- from the .vif or the
// generation-0 .ecsum, which on a multi-disk server may sit anywhere. Every
// directory that could hold one has to be in the search list, or the
// resolution silently falls back to 10+4 with the legacy striping and
// reconstructs through the wrong matrix.
// The rebuild's own data and index directories are handed to the resolvers
// as their own arguments, so they are deliberately absent from this list --
// unlike Go, whose resolver takes a single directory list and therefore
// carries the rebuild's index directory inside it.
#[test]
fn select_rebuild_location_excludes_the_rebuilds_own_dirs() {
for infos in [
vec![loc("/data1", "/idx1", 3, true)],
vec![loc("/data1", "/data1", 3, true)],
] {
let (idx, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(idx, 0);
assert!(others.is_empty(), "got {:?}", others);
}
}
// The case two reviewers flagged: a sibling holding only shards while its
// index directory holds this volume's .vif.
#[test]
fn select_rebuild_location_searches_a_siblings_index_dir_not_just_its_data_dir() {
let infos = vec![
loc("/data1", "/data1", 5, true),
loc("/data2", "/idx2", 2, false),
];
let (idx, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(idx, 0);
assert_eq!(others, vec!["/data2".to_string(), "/idx2".to_string()]);
}
// Several disks pointed at one index directory is a normal -dir.idx
// deployment; the shared directory is worth searching but only once.
#[test]
fn select_rebuild_location_lists_a_shared_index_dir_once() {
let infos = vec![
loc("/data1", "/data1", 5, true),
loc("/data2", "/shared-idx", 2, false),
loc("/data3", "/shared-idx", 1, false),
];
let (_, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(
others,
vec![
"/data2".to_string(),
"/data3".to_string(),
"/shared-idx".to_string()
]
);
}
// When the shared index directory is the rebuild's own it drops out, since
// the caller passes it separately.
#[test]
fn select_rebuild_location_omits_a_shared_index_dir_it_rebuilds_into() {
let infos = vec![
loc("/data1", "/shared-idx", 5, true),
loc("/data2", "/shared-idx", 2, false),
];
let (_, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(others, vec!["/data2".to_string()]);
}
// The winner moves as a fuller location turns up; the one it displaces
// still has to be searched, index directory included.
#[test]
fn select_rebuild_location_keeps_the_displaced_winners_dirs() {
let infos = vec![
loc("/data1", "/idx1", 2, true),
loc("/data2", "/idx2", 9, true),
];
let (idx, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(idx, 1, "the fuller location wins");
assert_eq!(others, vec!["/data1".to_string(), "/idx1".to_string()]);
}
#[test]
fn select_rebuild_location_drops_empty_dirs() {
let infos = vec![loc("/data1", "", 3, true), loc("/data2", "", 1, false)];
let (_, others) = select_rebuild_location(&infos).expect("a location with .ecx");
assert_eq!(others, vec!["/data2".to_string()]);
}
// Nothing carries an .ecx: there is no index to rebuild the shards against,
// so the caller answers with an empty rebuild rather than guessing.
#[test]
fn select_rebuild_location_is_none_without_an_ecx() {
let infos = vec![loc("/data1", "/idx1", 3, false)];
assert!(select_rebuild_location(&infos).is_none());
}
#[test]
fn test_parse_grpc_address_with_explicit_grpc_port() {
// Format: "ip:port.grpcPort" — used by SeaweedFS for source_data_node
@@ -5503,7 +5159,7 @@ mod tests {
data_size: "remote-incremental-copy".len() as u32,
..Needle::default()
};
volume.write_needle(&mut needle, true, false).unwrap();
volume.write_needle(&mut needle, true).unwrap();
volume.sync_to_disk().unwrap();
(
std::fs::read(volume.file_name(".dat")).unwrap(),
@@ -5671,7 +5327,7 @@ mod tests {
data_size: b"ec-generate".len() as u32,
..Needle::default()
};
volume.write_needle(&mut needle, true, false).unwrap();
volume.write_needle(&mut needle, true).unwrap();
volume.sync_to_disk().unwrap();
}
@@ -5732,29 +5388,6 @@ mod tests {
(VolumeGrpcService { state }, tmp)
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn test_volume_consolidate_index_rpc() {
let (service, _tmp) = make_local_service_with_volume("consolidate_rpc", None);
// Carry a peer address so the admin gate sees a caller; the test guard
// has an empty whitelist, so any peer is accepted.
let mut request = Request::new(volume_server_pb::VolumeConsolidateIndexRequest {
volume_id: 1,
});
request.extensions_mut().insert(tonic::transport::server::TcpConnectInfo {
local_addr: None,
remote_addr: Some("127.0.0.1:65000".parse().unwrap()),
});
service.volume_consolidate_index(request).await.unwrap();
// Data and index share a directory here, so there is nothing to move;
// the volume stays mounted and keeps its needle.
let store = service.state.store.read().unwrap();
let (_, v) = store.find_volume(VolumeId(1)).unwrap();
assert_eq!(v.file_count(), 1);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn test_volume_incremental_copy_streams_remote_only_volume_data() {
let (service, _tmp, shutdown_tx, dat_bytes, super_block_size, _delete_count) =
@@ -5894,7 +5527,7 @@ mod tests {
data: payload,
..Needle::default()
};
v.write_needle(&mut needle, true, false).unwrap();
v.write_needle(&mut needle, true).unwrap();
v.sync_to_disk().unwrap();
}
let dat_path = {
@@ -6338,89 +5971,4 @@ mod tests {
assert!(vif.expire_at_sec >= before + ttl.to_seconds());
assert!(vif.expire_at_sec <= before + ttl.to_seconds() + 5);
}
async fn scrub_ec_volume_1(
service: &VolumeGrpcService,
mode: i32,
) -> volume_server_pb::ScrubEcVolumeResponse {
service
.scrub_ec_volume(Request::new(volume_server_pb::ScrubEcVolumeRequest {
mode,
volume_ids: vec![1],
force_deleted_needles_check: false,
}))
.await
.unwrap()
.into_inner()
}
#[tokio::test]
async fn test_scrub_ec_volume_reads_mode_reconstructs_missing_shard() {
let (service, _tmp) = make_local_service_with_volume("", None);
service
.volume_ec_shards_generate(Request::new(
volume_server_pb::VolumeEcShardsGenerateRequest {
volume_id: 1,
collection: String::new(),
},
))
.await
.unwrap();
service
.volume_ec_shards_mount(Request::new(
volume_server_pb::VolumeEcShardsMountRequest {
volume_id: 1,
collection: String::new(),
shard_ids: (0..14).collect(),
source_disk_type: String::new(),
recover_missing_index: false,
},
))
.await
.unwrap();
// Seed the shard-location cache so the scrub skips the master lookup.
// The explicit-gRPC-port form targets port 1, so remote reads fail fast
// with connection-refused instead of resolving to a live local port.
{
let store = service.state.store.read().unwrap();
let ecv = store.find_ec_volume(VolumeId(1)).unwrap();
let mut locs = ecv.shard_locations.write().unwrap();
for sid in 0u8..14 {
locs.insert(sid, vec!["127.0.0.1:255.1".to_string()]);
}
*ecv.shard_locations_refresh_time.lock().unwrap() =
Some(std::time::Instant::now());
}
// All shards local: FULL is clean.
let resp = scrub_ec_volume_1(&service, 2).await;
assert!(resp.broken_volume_ids.is_empty(), "{:?}", resp.details);
assert_eq!(resp.total_files, 1);
service
.volume_ec_shards_unmount(Request::new(
volume_server_pb::VolumeEcShardsUnmountRequest {
volume_id: 1,
shard_ids: vec![0],
encode_ts_ns: 0,
},
))
.await
.unwrap();
// With a shard unreadable, FULL flags it AND fails the needle...
let resp = scrub_ec_volume_1(&service, 2).await;
assert_eq!(resp.broken_volume_ids, vec![1]);
assert!(resp.broken_shard_infos.iter().any(|s| s.shard_id == 0));
assert!(!resp.details.is_empty());
// ...while READS still reports the shard broken but reconstructs the
// interval from the local survivors, so no needle errors surface.
let resp = scrub_ec_volume_1(&service, 5).await;
assert_eq!(resp.broken_volume_ids, vec![1]);
assert!(resp.broken_shard_infos.iter().any(|s| s.shard_id == 0));
assert!(resp.details.is_empty(), "{:?}", resp.details);
assert_eq!(resp.total_files, 1);
}
}
+6 -78
View File
@@ -1589,15 +1589,12 @@ async fn get_or_head_handler_inner(
}
/// Handle HTTP Range requests. Returns 206 Partial Content or 416 Range Not Satisfiable.
#[derive(Clone, Copy, Debug)]
#[derive(Clone, Copy)]
struct HttpRange {
start: i64,
length: i64,
}
// Returned when the first-byte-pos of every byte-range-spec is at or past the content size.
const RANGE_NO_OVERLAP: &str = "invalid range: failed to overlap";
fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str> {
if s.is_empty() {
return Ok(Vec::new());
@@ -1607,7 +1604,6 @@ fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str
return Err("invalid range");
}
let mut ranges = Vec::new();
let mut no_overlap = false;
for part in s[PREFIX.len()..].split(',') {
let part = part.trim();
if part.is_empty() {
@@ -1631,13 +1627,9 @@ fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str
r.length = size - r.start;
} else {
let i = start_str.parse::<i64>().map_err(|_| "invalid range")?;
if i < 0 {
if i > size || i < 0 {
return Err("invalid range");
}
if i >= size {
no_overlap = true;
continue;
}
r.start = i;
if end_str.is_empty() {
r.length = size - r.start;
@@ -1654,9 +1646,6 @@ fn parse_range_header(s: &str, size: i64) -> Result<Vec<HttpRange>, &'static str
}
ranges.push(r);
}
if no_overlap && ranges.is_empty() {
return Err(RANGE_NO_OVERLAP);
}
Ok(ranges)
}
@@ -1690,15 +1679,7 @@ fn handle_range_request(
let total = data.len() as i64;
let ranges = match parse_range_header(range_str, total) {
Ok(r) => r,
Err(msg) => {
if msg == RANGE_NO_OVERLAP {
headers.insert(
"Content-Range",
format!("bytes */{}", total).parse().unwrap(),
);
}
return range_error_response(headers, msg);
}
Err(msg) => return range_error_response(headers, msg),
};
// Go's ProcessRangeRequest returns nil (empty body) for empty or oversized ranges
@@ -1781,15 +1762,7 @@ fn handle_range_request_from_source(
let total = info.data_size as i64;
let ranges = match parse_range_header(range_str, total) {
Ok(r) => r,
Err(msg) => {
if msg == RANGE_NO_OVERLAP {
headers.insert(
"Content-Range",
format!("bytes */{}", total).parse().unwrap(),
);
}
return range_error_response(headers, msg);
}
Err(msg) => return range_error_response(headers, msg),
};
if ranges.is_empty() {
@@ -2607,17 +2580,11 @@ pub async fn post_handler(
n.set_has_name();
}
// A durable write flushes before it is acked. Read it the way Go's
// r.FormValue does, off the decoded fields, so a percent-encoded value is
// honored here too. ReplicatedWrite forwards the parameter, so a replica
// sees it the same way the primary did.
let fsync = form_value("fsync").as_deref() == Some("true");
let write_result = if let Some(wq) = state.write_queue.get() {
wq.submit(vid, n.clone(), fsync).await
wq.submit(vid, n.clone()).await
} else {
let mut store = state.store.write().unwrap();
store.write_volume_needle(vid, &mut n, fsync)
store.write_volume_needle(vid, &mut n)
};
// Replicate to remote volume servers if this volume has replicas.
@@ -3885,26 +3852,6 @@ fn parse_content_disposition_filename(value: &str) -> Option<String> {
mod tests {
use super::*;
/// The upload handler reads fsync off the decoded query fields rather than
/// matching the raw string, because Go's r.FormValue decodes and a raw
/// match would silently drop a percent-encoded value.
#[test]
fn test_encoded_query_field_decodes() {
let raw = "fsync=%74rue";
assert!(
!raw.split('&').any(|p| p == "fsync=true"),
"a raw match is exactly what misses this"
);
let fields: Vec<(String, String)> = serde_urlencoded::from_str(raw).unwrap();
assert_eq!(
fields
.iter()
.find(|(k, _)| k == "fsync")
.map(|(_, v)| v.as_str()),
Some("true")
);
}
#[test]
fn test_parse_url_path_comma() {
let (vid, nid, cookie) = parse_url_path("/3,01637037d6").unwrap();
@@ -3939,25 +3886,6 @@ mod tests {
assert!(parse_url_path("").is_none());
}
#[test]
fn test_parse_range_header_no_overlap() {
assert_eq!(
parse_range_header("bytes=10-", 10).unwrap_err(),
RANGE_NO_OVERLAP
);
assert_eq!(
parse_range_header("bytes=100-", 10).unwrap_err(),
RANGE_NO_OVERLAP
);
// 416 only when every range fails to overlap
let ranges = parse_range_header("bytes=10-,0-1", 10).unwrap();
assert_eq!(ranges.len(), 1);
assert_eq!((ranges[0].start, ranges[0].length), (0, 2));
// an end past the size is clamped, still satisfiable
let ranges = parse_range_header("bytes=5-100", 10).unwrap();
assert_eq!((ranges[0].start, ranges[0].length), (5, 5));
}
#[test]
fn test_extract_jwt_bearer() {
let mut headers = HeaderMap::new();
+46 -186
View File
@@ -398,7 +398,8 @@ async fn do_heartbeat(
// Keep track of what we sent, to generate delta updates
let (initial_hb, initial_volumes) = collect_heartbeat_with_snapshot(config, state);
let mut last_volumes: HashMap<u32, VolumeIdentity> = volume_identities(&initial_volumes);
let mut last_volumes: HashMap<u32, master_pb::VolumeInformationMessage> =
initial_volumes.iter().map(|v| (v.id, v.clone())).collect();
let mut last_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -465,7 +466,8 @@ async fn do_heartbeat(
if changed {
let (adjusted_hb, adjusted_volumes) =
collect_heartbeat_with_snapshot(config, state);
last_volumes = volume_identities(&adjusted_volumes);
last_volumes =
adjusted_volumes.iter().map(|v| (v.id, v.clone())).collect();
last_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -499,7 +501,7 @@ async fn do_heartbeat(
s.maybe_adjust_volume_max();
}
let (current_hb, current_volumes) = collect_heartbeat_with_snapshot(config, state);
last_volumes = volume_identities(&current_volumes);
last_volumes = current_volumes.iter().map(|v| (v.id, v.clone())).collect();
last_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -528,7 +530,7 @@ async fn do_heartbeat(
return Ok(None);
}
let held_volumes = collect_volume_snapshot(config, state);
let current_volumes = volume_identities(&held_volumes);
let current_volumes: HashMap<u32, _> = held_volumes.iter().map(|v| (v.id, v.clone())).collect();
let current_ec_shards = {
let store = state.store.read().unwrap();
collect_ec_shard_delta_messages(&store)
@@ -539,13 +541,29 @@ async fn do_heartbeat(
for (id, vol) in &current_volumes {
if !last_volumes.contains_key(id) {
new_vols.push(vol.to_short_message(*id));
new_vols.push(master_pb::VolumeShortInformationMessage {
id: *id,
collection: vol.collection.clone(),
version: vol.version,
replica_placement: vol.replica_placement,
ttl: vol.ttl,
disk_type: vol.disk_type.clone(),
disk_id: vol.disk_id,
});
}
}
for (id, vol) in &last_volumes {
if !current_volumes.contains_key(id) {
del_vols.push(vol.to_short_message(*id));
del_vols.push(master_pb::VolumeShortInformationMessage {
id: *id,
collection: vol.collection.clone(),
version: vol.version,
replica_placement: vol.replica_placement,
ttl: vol.ttl,
disk_type: vol.disk_type.clone(),
disk_id: vol.disk_id,
});
}
}
@@ -726,54 +744,6 @@ fn parse_bool_property(value: Option<&String>) -> bool {
.unwrap_or(true)
}
/// What a mount or unmount delta has to name, which is far less than the
/// information message the heartbeat carries. A server holding millions of
/// volumes cannot keep a whole message for each just to notice one leave; the
/// Go report state keeps the same fields for the same reason.
#[derive(Clone)]
struct VolumeIdentity {
collection: String,
disk_type: String,
version: u32,
replica_placement: u32,
ttl: u32,
disk_id: u32,
}
impl VolumeIdentity {
fn of(v: &master_pb::VolumeInformationMessage) -> Self {
Self {
collection: v.collection.clone(),
disk_type: v.disk_type.clone(),
version: v.version,
replica_placement: v.replica_placement,
ttl: v.ttl,
disk_id: v.disk_id,
}
}
fn to_short_message(&self, id: u32) -> master_pb::VolumeShortInformationMessage {
master_pb::VolumeShortInformationMessage {
id,
collection: self.collection.clone(),
version: self.version,
replica_placement: self.replica_placement,
ttl: self.ttl,
disk_type: self.disk_type.clone(),
disk_id: self.disk_id,
}
}
}
fn volume_identities(
volumes: &[master_pb::VolumeInformationMessage],
) -> HashMap<u32, VolumeIdentity> {
volumes
.iter()
.map(|v| (v.id, VolumeIdentity::of(v)))
.collect()
}
/// Collect volume information into a Heartbeat message.
fn collect_heartbeat_with_snapshot(
config: &HeartbeatConfig,
@@ -801,10 +771,6 @@ fn collect_volume_snapshot(
build_heartbeat_with_ec_status(config, &mut store, Vec::new(), true, false).1
}
/// The heartbeat alone, without the volume snapshot the send loop pairs it
/// with. Only the tests want it that way; the loop calls
/// collect_heartbeat_with_snapshot directly.
#[cfg(test)]
fn collect_heartbeat(
config: &HeartbeatConfig,
state: &Arc<VolumeServerState>,
@@ -872,7 +838,8 @@ fn build_heartbeat_with_ec_status(
// master can tell whether applying what it was sent leaves it current.
// Volumes skipped below -- quarantined, phantom, expired -- are in neither.
let mut volume_digest: u64 = 0;
let (send_full_list, report_generation, report_pass) = store.volume_report.begin();
let (send_full_list, report_generation) = store.volume_report.begin();
let mut reported_hashes: HashMap<VolumeReportKey, u64> = HashMap::new();
let mut changed_volumes = Vec::new();
let mut max_file_key = NeedleId(0);
let mut max_volume_counts: HashMap<String, u32> = HashMap::new();
@@ -970,14 +937,8 @@ fn build_heartbeat_with_ec_status(
let hash = report_hash(&volume_message);
volume_digest ^= hash;
let key: VolumeReportKey = (volume_message.disk_id, volume_message.id);
// A snapshot must leave the reporting state as it found it, so
// it asks rather than marks.
let is_news = if commit_report {
store.volume_report.record(key, hash, report_pass)
} else {
store.volume_report.changed(key, hash)
};
if send_full_list || is_news {
reported_hashes.insert(key, hash);
if send_full_list || store.volume_report.changed(key, hash) {
changed_volumes.push(volume_message.clone());
}
volumes.push(volume_message);
@@ -986,30 +947,24 @@ fn build_heartbeat_with_ec_status(
should_delete_volume = true;
}
// Track disk size by collection. A volume on its way out is left
// out: an entry here is also what says the collection is still on
// this server.
// Track disk size by collection
let entry = disk_sizes.entry(vol.collection.clone()).or_insert((0, 0));
if !should_delete_volume {
let entry = disk_sizes.entry(vol.collection.clone()).or_insert((0, 0));
entry.0 += volume_size;
entry.1 += vol.deleted_size();
}
// An entry here is what says the collection is still on this
// server, so a volume on its way out must not make one.
if !should_delete_volume {
let read_only = ro_counts.entry(vol.collection.clone()).or_default();
if vol.is_read_only() {
read_only.is_read_only += 1;
if vol.is_no_write_or_delete() {
read_only.no_write_or_delete += 1;
}
if vol.is_no_write_can_delete() {
read_only.no_write_can_delete += 1;
}
if loc.is_disk_space_low.load(Ordering::Relaxed) {
read_only.is_disk_space_low += 1;
}
let read_only = ro_counts.entry(vol.collection.clone()).or_default();
if !should_delete_volume && vol.is_read_only() {
read_only.is_read_only += 1;
if vol.is_no_write_or_delete() {
read_only.no_write_or_delete += 1;
}
if vol.is_no_write_can_delete() {
read_only.no_write_can_delete += 1;
}
if loc.is_disk_space_low.load(Ordering::Relaxed) {
read_only.is_disk_space_low += 1;
}
}
@@ -1043,17 +998,6 @@ fn build_heartbeat_with_ec_status(
.with_label_values(&[col, crate::metrics::READ_ONLY_LABEL_IS_DISK_SPACE_LOW])
.set(counts.is_disk_space_low as f64);
}
// ro_counts has an entry for every collection that kept a volume through
// this pass, including the ones counting zero read-only volumes.
{
let mut reported = store.reported_collections.lock().unwrap();
for col in reported.iter() {
if !ro_counts.contains_key(col) {
crate::metrics::delete_volume_server_collection_metrics(col);
}
}
*reported = ro_counts.keys().cloned().collect();
}
// Update max volumes gauge
let total_max: i64 = max_volume_counts.values().map(|v| *v as i64).sum();
crate::metrics::MAX_VOLUMES.set(total_max);
@@ -1061,7 +1005,7 @@ fn build_heartbeat_with_ec_status(
// Only when this heartbeat is going to be sent: marking volumes reported
// and then discarding the message would leave the master never told.
if commit_report {
store.volume_report.commit(report_pass, report_generation);
store.volume_report.commit(reported_hashes, report_generation);
}
// has_no_volumes says the server holds nothing, so it may only be derived
@@ -1129,14 +1073,6 @@ fn collect_live_ec_shards(
.with_label_values(&[col, crate::metrics::DISK_SIZE_LABEL_EC])
.set(*size as f64);
}
let mut reported = store.reported_ec_collections.lock().unwrap();
for col in reported.iter() {
if !ec_sizes.contains_key(col) {
let _ = crate::metrics::DISK_SIZE_GAUGE
.remove_label_values(&[col, crate::metrics::DISK_SIZE_LABEL_EC]);
}
}
*reported = ec_sizes.keys().cloned().collect();
}
ec_shards
@@ -1488,10 +1424,10 @@ mod tests {
store.volume_report.accept_deltas();
build_heartbeat(&test_config(), &mut store);
let (full, generation, pass) = store.volume_report.begin();
let (full, generation) = store.volume_report.begin();
assert!(!full);
store.volume_report.request_full_list();
store.volume_report.commit(pass, generation);
store.volume_report.commit(HashMap::new(), generation);
let heartbeat = build_heartbeat(&test_config(), &mut store);
assert_eq!(heartbeat.volumes.len(), 2);
@@ -1608,7 +1544,7 @@ mod tests {
let (_, volume) = store.find_volume_mut(VolumeId(17)).unwrap();
volume.set_read_only().unwrap();
volume.volume_info.files.push(Default::default());
volume.refresh_remote_write_mode().unwrap();
volume.refresh_remote_write_mode();
}
let heartbeat = build_heartbeat(&test_config(), &mut store);
@@ -1656,82 +1592,6 @@ mod tests {
);
}
fn collection_series(gauge: &prometheus::GaugeVec, collection: &str) -> usize {
use prometheus::core::Collector;
gauge
.collect()
.iter()
.flat_map(|family| family.get_metric().to_vec())
.filter(|metric| {
metric
.get_label()
.iter()
.any(|label| label.get_name() == "collection" && label.get_value() == collection)
})
.count()
}
// The per-collection gauges are only ever set for collections the heartbeat
// still finds on this server. A volume.balance that moves a collection's
// last volume off a server used to leave its read-only count - marked
// read-only for the move, moments before it went - standing on that server
// until a restart, with nothing in volume.list to match it.
#[test]
fn test_build_heartbeat_clears_metrics_of_departed_collection() {
let temp_dir = tempfile::tempdir().unwrap();
let dir = temp_dir.path().to_str().unwrap();
let collection = "heartbeat_departed_case";
let mut store = Store::new(NeedleMapKind::InMemory);
store
.add_location(
dir,
dir,
8,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
store
.add_volume(
VolumeId(21),
collection,
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
{
let (_, volume) = store.find_volume_mut(VolumeId(21)).unwrap();
volume.set_read_only().unwrap();
}
build_heartbeat(&test_config(), &mut store);
assert_eq!(
READ_ONLY_VOLUME_GAUGE
.with_label_values(&[collection, READ_ONLY_LABEL_IS_READ_ONLY])
.get(),
1.0
);
assert!(store.unmount_volume(VolumeId(21)));
build_heartbeat(&test_config(), &mut store);
assert_eq!(
collection_series(&READ_ONLY_VOLUME_GAUGE, collection),
0,
"read-only series left after the collection left the server"
);
assert_eq!(
collection_series(&DISK_SIZE_GAUGE, collection),
0,
"disk size series left after the collection left the server"
);
}
#[test]
fn test_build_heartbeat_reports_disk_bytes() {
let temp_dir = tempfile::tempdir().unwrap();
@@ -1974,7 +1834,7 @@ mod tests {
key: "volumes/71.dat".to_string(),
..Default::default()
});
volume.refresh_remote_write_mode().unwrap();
volume.refresh_remote_write_mode();
let heartbeat = build_heartbeat(&test_config(), &mut store);
-2
View File
@@ -163,8 +163,6 @@ mod tests {
https_client_ca_file: String::new(),
grpc_cert_file: String::new(),
grpc_key_file: String::new(),
grpc_client_cert_file: String::new(),
grpc_client_key_file: String::new(),
grpc_ca_file: String::new(),
grpc_allowed_wildcard_domain: String::new(),
grpc_volume_allowed_common_names: vec![],
+125 -334
View File
@@ -33,9 +33,7 @@ use std::sync::Arc;
use std::time::{Duration, Instant};
use futures::future::join_all;
use futures::stream::{self, StreamExt};
use reed_solomon_erasure::galois_8::ReedSolomon;
use tokio::sync::Semaphore;
use tonic::Request;
use crate::pb::master_pb::{self, seaweed_client::SeaweedClient, LookupEcVolumeRequest};
@@ -51,19 +49,6 @@ use crate::storage::store_ec_reconcile::EcVolumeMissingIndex;
use crate::storage::types::*;
use crate::storage::volume::volume_file_name;
/// Bounds the fan-out of a single needle read. Mirrors Go's
/// `ecIntervalReadConcurrency`.
const INTERVAL_READ_CONCURRENCY: usize = 8;
/// Bounds the bytes EC recovery holds in flight across every concurrent read.
/// Recovery is the one read path that multiplies the served bytes — it keeps an
/// interval-sized buffer per shard alive until Reed-Solomon runs — and a peer
/// that is slow to fail holds each of them for the whole gRPC timeout, so a
/// burst of reads during a network blip walked the server into an OOM. Mirrors
/// Go's `ecRecoverBudget`.
const EC_RECOVER_BUDGET: usize = 256 << 20;
static EC_RECOVER_SEM: Semaphore = Semaphore::const_new(EC_RECOVER_BUDGET);
/// One interval's data after Phase A.
enum IntervalResult {
/// Already read from a locally-mounted shard.
@@ -105,7 +90,7 @@ pub async fn read_ec_shard_needle_distributed(
// intervals, and read any locally-mounted shard intervals. We must
// not `.await` while holding this guard (std::sync::RwLockReadGuard
// is !Send).
let mut snapshot = match snapshot_under_lock(state, vid, needle_id)? {
let snapshot = match snapshot_under_lock(state, vid, needle_id)? {
Some(s) => s,
None => return Ok(None),
};
@@ -121,9 +106,7 @@ pub async fn read_ec_shard_needle_distributed(
let mut shard_locations = snapshot.cached_locations.clone();
if any_remote
&& claim_shard_locations_refresh(
state,
vid,
&& needs_refresh(
&shard_locations,
snapshot.cache_refreshed_at,
snapshot.data_shards as usize,
@@ -134,21 +117,17 @@ pub async fn read_ec_shard_needle_distributed(
Ok(fresh) => {
// A complete reply merges into the cache; an incomplete one
// (< data_shards) is left unwritten — keep the prior cache.
match write_back_shard_locations(state, vid, fresh, snapshot.data_shards as usize)
if let Some(merged) =
write_back_shard_locations(state, vid, fresh, snapshot.data_shards as usize)
{
Some(merged) => shard_locations = merged,
// An incomplete reply leaves the cache unwritten and its refresh
// time unadvanced, so the mark this refresh consumed goes back.
None => mark_shard_locations_stale(state, vid),
shard_locations = merged;
}
}
Err(e) => {
// Lookup failed — proceed with cached values. If cache
// is empty, the remote fetch below will fail and we
// surface a NotFound (matching Go's behavior when no
// locations are known). The mark goes back: nothing
// answered for it, and the map stays disproved.
mark_shard_locations_stale(state, vid);
// locations are known).
tracing::warn!(
"ec lookup failed for volume {}: {} — using cached locations ({} entries)",
vid.0,
@@ -159,56 +138,39 @@ pub async fn read_ec_shard_needle_distributed(
}
}
// Phase C — fetch missing intervals, reconstructing when the direct peer
// read fails. Blocks that follow each other in the .dat live on different
// shards, so a needle spanning several of them costs one round trip per
// block when fetched in sequence; `buffered` keeps the order while letting
// INTERVAL_READ_CONCURRENCY of them fly at once.
let data_shards = snapshot.data_shards as usize;
let parity_shards = snapshot.parity_shards as usize;
let encode_ts_ns = snapshot.encode_ts_ns;
let intervals = std::mem::take(&mut snapshot.intervals);
let fetched: Vec<io::Result<(Vec<u8>, bool)>> = stream::iter(intervals.into_iter().map(|res| {
let shard_locations = &shard_locations;
async move {
match res {
IntervalResult::Local(buf) => Ok((buf, false)),
IntervalResult::NeedRemote {
// Phase C — fetch missing intervals, reconstructing when the
// direct peer read fails.
let mut assembled: Vec<Vec<u8>> = Vec::with_capacity(snapshot.intervals.len());
for res in snapshot.intervals {
match res {
IntervalResult::Local(buf) => assembled.push(buf),
IntervalResult::NeedRemote {
shard_id,
shard_offset,
size,
} => {
let (buf, is_deleted) = fetch_one_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
} => {
fetch_one_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
encode_ts_ns,
)
.await
&shard_locations,
snapshot.data_shards as usize,
snapshot.parity_shards as usize,
snapshot.encode_ts_ns,
)
.await?;
// A peer reports the needle deleted (a cross-server window where the
// local index still shows it live): treat as not-found rather than
// serving zeros, mirroring Go's ErrorDeleted.
if is_deleted {
return Ok(None);
}
assembled.push(buf);
}
}
}))
.buffered(INTERVAL_READ_CONCURRENCY)
.collect()
.await;
let mut assembled: Vec<Vec<u8>> = Vec::with_capacity(fetched.len());
for res in fetched {
let (buf, is_deleted) = res?;
// A peer reports the needle deleted (a cross-server window where the
// local index still shows it live): treat as not-found rather than
// serving zeros, mirroring Go's ErrorDeleted.
if is_deleted {
return Ok(None);
}
assembled.push(buf);
}
// Phase D — assemble and parse the Needle. Mirrors the tail of
@@ -246,9 +208,7 @@ pub async fn read_ec_shard_needle_distributed(
/// without decoding (so genuine shard faults are reported rather than healed).
/// Mirrors Go's `Store.ScrubEcVolume`. Returns (rows walked, broken shards,
/// errors). `force_deleted_needles_check` disables the benign delete-state
/// size-mismatch suppression. `recover_unreadable` (READS mode) rebuilds an
/// unreadable interval from the surviving shards: the same shards are reported
/// broken, but only needles parity can no longer recover become errors.
/// size-mismatch suppression.
///
/// Shard locations are refreshed once up front. Each needle is then processed via
/// `scrub_snapshot_under_lock` + lock-drop + no-reconstruct `read_remote_ec_shard_interval`,
@@ -257,19 +217,10 @@ pub async fn scrub_ec_volume_distributed(
state: &Arc<VolumeServerState>,
vid: VolumeId,
force_deleted_needles_check: bool,
recover_unreadable: bool,
) -> (i64, Vec<crate::pb::volume_server_pb::EcShardInfo>, Vec<String>) {
// Phase A — under the Store read lock, run the index scrub and grab the
// paths/scalars + shard-location staleness; release the lock before any await.
let (
ecx_path,
collection,
seed_errs,
cached_locations,
cache_refreshed_at,
data_shards,
total_shards,
) = {
let (ecx_path, collection, seed_errs, cached_locations, cache_refreshed_at, data_shards, total_shards) = {
let store = state.store.read().unwrap();
let ecv = match store.find_ec_volume(vid) {
Some(v) => v,
@@ -304,18 +255,10 @@ pub async fn scrub_ec_volume_distributed(
// cachedLookupEcShardLocations). A partial reply (< data_shards locations, a
// master mid-recovery) or a failed lookup is a hard, retryable error — never
// overwrite a good cache with a partial map or storm a down master per needle.
if claim_shard_locations_refresh(
state,
vid,
&cached_locations,
cache_refreshed_at,
data_shards,
total_shards,
) {
if needs_refresh(&cached_locations, cache_refreshed_at, data_shards, total_shards) {
match cached_lookup_ec_shard_locations(state, vid).await {
Ok(fresh) => {
if write_back_shard_locations(state, vid, fresh, data_shards).is_none() {
mark_shard_locations_stale(state, vid);
return (
0,
Vec::new(),
@@ -327,12 +270,11 @@ pub async fn scrub_ec_volume_distributed(
}
}
Err(e) => {
mark_shard_locations_stale(state, vid);
return (
0,
Vec::new(),
vec![format!("failed to locate shard via master grpc: {}", e)],
);
)
}
}
}
@@ -398,9 +340,8 @@ pub async fn scrub_ec_volume_distributed(
}
};
// Read each interval local-then-remote. Neither read decodes: the point is to
// find shards that are themselves broken, not to heal around them. READS then
// rebuilds what it could not read. Locations refreshed above.
// Read each interval local-then-remote WITHOUT reconstructing: we verify
// the shards are valid, we do not heal them. Locations refreshed above.
let n_intervals = snapshot.intervals.len();
let mut data: Vec<u8> = Vec::with_capacity(snapshot.actual_size);
for (i, res) in snapshot.intervals.iter().enumerate() {
@@ -430,9 +371,11 @@ pub async fn scrub_ec_volume_distributed(
// -> the delete-state suppression (mirrors Go's pre-zeroed buffer).
Ok((_, true)) => data.resize(data.len() + *ssize, 0),
Ok((buf, false)) => data.extend_from_slice(&buf),
Err(read_err) => {
// The shard is broken whether or not the needle survives it,
// so report it either way.
Err(_) => {
errs.push(format!(
"failed to read EC shard {} for needle {} on volume {} (interval {}/{})",
shard_id, id.0, vid.0, i + 1, n_intervals
));
broken_shards.insert(
*shard_id,
crate::pb::volume_server_pb::EcShardInfo {
@@ -443,41 +386,7 @@ pub async fn scrub_ec_volume_distributed(
..Default::default()
},
);
if !recover_unreadable {
errs.push(format!(
"failed to read EC shard {} for needle {} on volume {} (interval {}/{}): {}",
shard_id, id.0, vid.0, i + 1, n_intervals, read_err
));
break;
}
match recover_one_remote_ec_shard_interval(
state,
vid,
id,
*shard_id,
*shard_offset,
*ssize,
&locations,
data_shards,
total_shards - data_shards,
snapshot.encode_ts_ns,
)
.await
{
// Same as the direct read above: a holder reporting the
// needle deleted is authoritative and answers with no
// bytes, so zero-fill and let the delete-state
// suppression have it.
Ok((_, true)) => data.resize(data.len() + *ssize, 0),
Ok((buf, false)) => data.extend_from_slice(&buf),
Err(e) => {
errs.push(format!(
"failed to recover EC shard {} for needle {} on volume {} (interval {}/{}): {}",
shard_id, id.0, vid.0, i + 1, n_intervals, e
));
break;
}
}
break;
}
}
}
@@ -599,7 +508,7 @@ fn read_local_intervals(
) -> Vec<IntervalResult> {
let mut interval_results = Vec::with_capacity(intervals.len());
for interval in intervals {
let (shard_id, shard_offset) = ecv.interval_to_shard_id_and_offset(interval);
let (shard_id, shard_offset) = interval.to_shard_id_and_offset(ecv.data_shards);
let buf_size = interval.size as usize;
let local = ecv.shards.get(shard_id as usize).and_then(|s| s.as_ref());
match local {
@@ -662,7 +571,6 @@ fn build_snapshot(
fn needs_refresh(
locations: &HashMap<ShardId, Vec<String>>,
refreshed_at: Option<Instant>,
stale: bool,
data_shards: usize,
total_shards: usize,
) -> bool {
@@ -672,53 +580,16 @@ fn needs_refresh(
None => return true,
};
let shard_count = locations.len();
// A complete map is trusted longest. One short of data_shards, or one a
// failed read has just disproved, is re-checked promptly: until it is, reads
// keep aiming at a location the shard has left.
let ttl = if stale || shard_count < data_shards {
Duration::from_secs(11)
} else if shard_count == total_shards {
Duration::from_secs(37 * 60)
} else {
Duration::from_secs(7 * 60)
};
age >= ttl
}
/// Mark the cached shard map for a prompt re-check after a read failed against
/// one of its locations. Go drops the entry outright in `forgetShardId`, which
/// costs it the direct read until the map is re-learned; here the entry stays
/// (a dead peer just fails fast on the next attempt) and only the freshness
/// window is cut, so a shard that has moved is picked up in seconds either way.
fn mark_shard_locations_stale(state: &Arc<VolumeServerState>, vid: VolumeId) {
let store = state.store.read().unwrap();
if let Some(ecv) = store.find_ec_volume(vid) {
*ecv.shard_locations_stale.lock().unwrap() = true;
if shard_count < data_shards && age < Duration::from_secs(11) {
return false;
}
}
/// Decide whether the cached map is due a master lookup and, when it is, consume
/// its stale mark in the same critical section. A mark raised from here on
/// belongs to the next refresh: the read that raised it has disproved the map
/// this lookup is about to install.
fn claim_shard_locations_refresh(
state: &Arc<VolumeServerState>,
vid: VolumeId,
locations: &HashMap<ShardId, Vec<String>>,
refreshed_at: Option<Instant>,
data_shards: usize,
total_shards: usize,
) -> bool {
let store = state.store.read().unwrap();
let Some(ecv) = store.find_ec_volume(vid) else {
return needs_refresh(locations, refreshed_at, false, data_shards, total_shards);
};
let mut stale = ecv.shard_locations_stale.lock().unwrap();
let refresh = needs_refresh(locations, refreshed_at, *stale, data_shards, total_shards);
if refresh {
*stale = false;
if shard_count == total_shards && age < Duration::from_secs(37 * 60) {
return false;
}
refresh
if shard_count >= data_shards && age < Duration::from_secs(7 * 60) {
return false;
}
true
}
async fn cached_lookup_ec_shard_locations(
@@ -854,9 +725,6 @@ async fn fetch_one_interval(
sources,
e
);
// Reconstruction below skips this very shard, so nothing else
// invalidates the location that just failed.
mark_shard_locations_stale(state, vid);
}
}
}
@@ -864,7 +732,7 @@ async fn fetch_one_interval(
// Reconstruct: fan-out reads to every other shard at the same
// (shard_offset, size). Mirrors `recoverOneRemoteEcShardInterval`.
recover_one_remote_ec_shard_interval(
let buf = recover_one_remote_ec_shard_interval(
state,
vid,
needle_id,
@@ -876,7 +744,8 @@ async fn fetch_one_interval(
parity_shards,
expected_encode_ts_ns,
)
.await
.await?;
Ok((buf, false))
}
async fn read_remote_ec_shard_interval(
@@ -1036,7 +905,7 @@ async fn recover_one_remote_ec_shard_interval(
data_shards: usize,
parity_shards: usize,
expected_encode_ts_ns: i64,
) -> io::Result<(Vec<u8>, bool)> {
) -> io::Result<Vec<u8>> {
let total_shards = data_shards + parity_shards;
let rs = ReedSolomon::new(data_shards, parity_shards).map_err(|e| {
io::Error::new(
@@ -1045,136 +914,89 @@ async fn recover_one_remote_ec_shard_interval(
)
})?;
// Charge the buffers this recovery is about to hold against the budget, so a
// burst of them queues here rather than on the heap. An interval whose
// fan-out outgrows the whole budget takes all of it and so runs alone,
// rather than waiting on permits that can never be granted.
let _permit = EC_RECOVER_SEM
.acquire_many((size * data_shards).min(EC_RECOVER_BUDGET) as u32)
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"ec recover budget for shard {}.{}: {}",
vid.0, shard_id_to_recover, e
),
)
})?;
let mut bufs: Vec<Option<Vec<u8>>> = vec![None; total_shards];
// Phase 0: seed bufs from LOCALLY mounted shards. If this node
// already holds enough sibling shards, reconstruction completes
// without any peer fan-out — and even with a cold/incomplete
// shard_locations cache or a failed master lookup, local
// survivors still contribute.
let mut available = 0usize;
// survivors still contribute. Mirrors Go's
// recoverOneRemoteEcShardInterval behaviour, which is implicitly
// local-aware because the Store fan-out targets ALL known
// locations (including the caller's own server address); the
// Rust port had been remote-only, so reconstructing with a cold
// cache failed even when enough siblings were on disk.
{
let store = state.store.read().unwrap();
for sid in 0..total_shards {
if available >= data_shards {
break;
}
if sid as ShardId == shard_id_to_recover {
continue;
}
// Resolve the shard together with the EcVolume on the disk that owns
// it: a reconciled volume has its shards split across data dirs. A
// shard from a different encode run must not be fed to Reed-Solomon;
// lenient only when the caller carries no identity (pre-upgrade).
// Mirrors Go's `readLocalEcShardInterval`.
let owner = match store.find_ec_volume_with_shard(vid, sid as u32) {
Some(ecv) if expected_encode_ts_ns == 0 || ecv.encode_ts_ns == expected_encode_ts_ns => ecv,
_ => continue,
};
if let Some(Some(shard)) = owner.shards.get(sid) {
let mut buf = vec![0u8; size];
if shard.read_at(&mut buf, shard_offset as u64).map(|n| n == size).unwrap_or(false) {
bufs[sid] = Some(buf);
available += 1;
if let Some(ecv) = store.find_ec_volume(vid) {
for sid in 0..total_shards {
if sid as ShardId == shard_id_to_recover {
continue;
}
}
}
}
// Phase 1: remote fan-out over the shard locations we DON'T already have
// locally and DON'T need to recover. Reconstruction consumes data_shards
// shards, so reading every remaining one holds a third more buffers than
// that and asks a third more of peers that may already be struggling: fetch
// what is still missing, and widen only if some of those reads fail.
let mut candidates: Vec<(ShardId, Vec<String>)> = shard_locations
.iter()
.filter(|(sid, locs)| {
**sid != shard_id_to_recover
&& (**sid as usize) < total_shards
&& !locs.is_empty()
&& bufs[**sid as usize].is_none()
})
.map(|(sid, locs)| (*sid, locs.clone()))
.collect();
let mut any_deleted = false;
while available < data_shards && !candidates.is_empty() {
let rest = candidates.split_off((data_shards - available).min(candidates.len()));
let wave = std::mem::replace(&mut candidates, rest);
let results = join_all(wave.into_iter().map(|(sid, locs)| {
let state = state.clone();
async move {
let res = read_remote_ec_shard_interval(
&state,
&locs,
vid,
needle_id,
sid,
shard_offset,
size,
expected_encode_ts_ns,
)
.await;
(sid, res)
}
}))
.await;
for (sid, res) in results {
match res {
// Exclude a deleted shard from reconstruction (Go gates on a full
// read): feeding the empty/zero buffer into Reed-Solomon would
// corrupt the recovered shard.
Ok((buf, is_deleted)) => {
if is_deleted {
any_deleted = true;
continue;
if let Some(Some(shard)) = ecv.shards.get(sid) {
let mut buf = vec![0u8; size];
if shard.read_at(&mut buf, shard_offset as u64).map(|n| n == size).unwrap_or(false) {
bufs[sid] = Some(buf);
}
bufs[sid as usize] = Some(buf);
available += 1;
}
Err(e) => {
tracing::debug!(
"recover: read {}.{} for needle {} failed: {}",
vid.0,
sid,
needle_id,
e
);
mark_shard_locations_stale(state, vid);
}
}
}
if any_deleted {
// every shard of a deleted needle answers deleted, so another wave cannot help
break;
}
// Phase 1: remote fan-out — one task per known shard location
// we DON'T already have locally and DON'T need to recover.
let mut tasks = Vec::new();
for (sid, locs) in shard_locations {
if *sid == shard_id_to_recover || locs.is_empty() {
continue;
}
if bufs[*sid as usize].is_some() {
continue;
}
let sid = *sid;
let locs = locs.clone();
let state = state.clone();
tasks.push(async move {
let res = read_remote_ec_shard_interval(
&state,
&locs,
vid,
needle_id,
sid,
shard_offset,
size,
expected_encode_ts_ns,
)
.await;
(sid, res)
});
}
let results = join_all(tasks).await;
for (sid, res) in results {
match res {
// Exclude a deleted shard from reconstruction (Go gates on a full
// read): feeding the empty/zero buffer into Reed-Solomon would
// corrupt the recovered shard.
Ok((buf, is_deleted)) => {
if !is_deleted && (sid as usize) < total_shards {
bufs[sid as usize] = Some(buf);
}
}
Err(e) => {
tracing::debug!(
"recover: read {}.{} for needle {} failed: {}",
vid.0,
sid,
needle_id,
e
);
}
}
}
let available = bufs.iter().filter(|b| b.is_some()).count();
if available < data_shards {
// A holder reporting the needle deleted is authoritative -- deletes are
// never invented and never undone -- so answer that rather than the
// failure to gather shards of a needle that is gone.
if any_deleted {
return Ok((Vec::new(), true));
}
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
@@ -1195,7 +1017,7 @@ async fn recover_one_remote_ec_shard_interval(
})?;
match bufs.into_iter().nth(shard_id_to_recover as usize).flatten() {
Some(buf) => Ok((buf, any_deleted)),
Some(buf) => Ok(buf),
None => Err(io::Error::new(
io::ErrorKind::Other,
format!(
@@ -1433,34 +1255,3 @@ async fn drain_copy_stream(
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
fn locations(count: usize) -> HashMap<ShardId, Vec<String>> {
(0..count)
.map(|sid| (sid as ShardId, vec!["127.0.0.1:8080".to_string()]))
.collect()
}
#[test]
fn needs_refresh_re_checks_a_map_a_failed_read_disproved() {
let just_now = Some(Instant::now());
let aged = Some(Instant::now() - Duration::from_secs(12));
// A complete map is trusted for a long time, and one shard short still
// outlasts a 12-second gap.
assert!(!needs_refresh(&locations(14), aged, false, 10, 14));
assert!(!needs_refresh(&locations(13), aged, false, 10, 14));
// Disproved by a read, the same maps are re-checked within seconds.
assert!(needs_refresh(&locations(14), aged, true, 10, 14));
assert!(needs_refresh(&locations(13), aged, true, 10, 14));
// But the mark buys one prompt re-check, not a lookup per read.
assert!(!needs_refresh(&locations(14), just_now, true, 10, 14));
// A map short of the data shards is re-checked promptly regardless.
assert!(needs_refresh(&locations(9), aged, false, 10, 14));
// An unrefreshed cache always looks up.
assert!(needs_refresh(&locations(0), None, false, 10, 14));
}
}
+14 -20
View File
@@ -1,10 +1,9 @@
//! Async batched write processing for the volume server.
//!
//! Instead of each upload handler directly calling `write_needle`, writes are
//! submitted to a queue. A background worker drains the queue in batches (up to
//! 128 entries), groups them by volume ID, and processes them together under a
//! single store lock. Requests that asked for `fsync` are flushed by
//! `write_needle` itself, one flush per durable write.
//! Instead of each upload handler directly calling `write_needle` and syncing,
//! writes are submitted to a queue. A background worker drains the queue in
//! batches (up to 128 entries), groups them by volume ID, processes them
//! together, and syncs once per volume for the entire batch.
use std::sync::Arc;
@@ -20,14 +19,10 @@ use super::volume_server::VolumeServerState;
/// Result of a single write operation: (offset, size, is_unchanged).
pub type WriteResult = Result<(u64, Size, bool), VolumeError>;
/// The needles queued for one volume, each with the durability it asked for.
type VolumeBatch = Vec<(Needle, bool, oneshot::Sender<WriteResult>)>;
/// A request to write a needle, submitted to the write queue.
pub struct WriteRequest {
pub volume_id: VolumeId,
pub needle: Needle,
pub fsync: bool,
pub response_tx: oneshot::Sender<WriteResult>,
}
@@ -59,12 +54,11 @@ impl WriteQueue {
/// Submit a write request and wait for the result.
///
/// Returns `Err` if the worker has shut down or the response channel was dropped.
pub async fn submit(&self, volume_id: VolumeId, needle: Needle, fsync: bool) -> WriteResult {
pub async fn submit(&self, volume_id: VolumeId, needle: Needle) -> WriteResult {
let (response_tx, response_rx) = oneshot::channel();
let request = WriteRequest {
volume_id,
needle,
fsync,
response_tx,
};
@@ -144,14 +138,14 @@ fn process_batch(state: Arc<VolumeServerState>, batch: Vec<WriteRequest>) {
// Group requests by volume ID for efficient processing.
// We use a Vec of (VolumeId, Vec<(Needle, Sender)>) to preserve order
// and avoid requiring Hash on VolumeId.
let mut groups: Vec<(VolumeId, VolumeBatch)> = Vec::new();
let mut groups: Vec<(VolumeId, Vec<(Needle, oneshot::Sender<WriteResult>)>)> = Vec::new();
for req in batch {
let vid = req.volume_id;
if let Some(group) = groups.iter_mut().find(|(v, _)| *v == vid) {
group.1.push((req.needle, req.fsync, req.response_tx));
group.1.push((req.needle, req.response_tx));
} else {
groups.push((vid, vec![(req.needle, req.fsync, req.response_tx)]));
groups.push((vid, vec![(req.needle, req.response_tx)]));
}
}
@@ -159,8 +153,8 @@ fn process_batch(state: Arc<VolumeServerState>, batch: Vec<WriteRequest>) {
let mut store = state.store.write().unwrap();
for (vid, entries) in groups {
for (mut needle, fsync, response_tx) in entries {
let result = store.write_volume_needle(vid, &mut needle, fsync);
for (mut needle, response_tx) in entries {
let result = store.write_volume_needle(vid, &mut needle);
// Send result back; ignore error if receiver dropped.
let _ = response_tx.send(result);
}
@@ -245,7 +239,7 @@ mod tests {
..Needle::default()
};
let result = queue.submit(VolumeId(999), needle, false).await;
let result = queue.submit(VolumeId(999), needle).await;
assert!(result.is_err());
match result {
Err(VolumeError::NotFound) => {} // expected
@@ -270,7 +264,7 @@ mod tests {
data_size: 10,
..Needle::default()
};
q.submit(VolumeId(1), needle, false).await
q.submit(VolumeId(1), needle).await
}));
}
@@ -299,7 +293,7 @@ mod tests {
data_size: 4,
..Needle::default()
};
q.submit(VolumeId(42), needle, false).await
q.submit(VolumeId(42), needle).await
}));
}
@@ -333,7 +327,7 @@ mod tests {
data_size: 0,
..Needle::default()
};
let result = queue2.submit(VolumeId(1), needle, false).await;
let result = queue2.submit(VolumeId(1), needle).await;
assert!(result.is_err()); // NotFound is fine -- the point is it doesn't panic
}
}
+14 -126
View File
@@ -843,8 +843,7 @@ impl DiskLocation {
// .ecx open error, .ecj create error, malformed .vif) would
// have to panic via unwrap(). Build the EcVolume up front and
// propagate the error to the caller.
let created = !self.ec_volumes.contains_key(&vid);
if created {
if !self.ec_volumes.contains_key(&vid) {
let ec_vol = EcVolume::new(&dir, idx_dir, collection, vid)
.map_err(VolumeError::Io)?;
self.ec_volumes.insert(vid, ec_vol);
@@ -867,27 +866,9 @@ impl DiskLocation {
}
for &shard_id in shard_ids {
// A mount retry re-listing a shard this volume already holds:
// keep the existing registration (mirrors Go's AddEcVolumeShard
// added=false) — re-adding would replace a serving fd and bump
// the ec_shards gauge without growing the mounted count.
if ec_vol.has_shard(shard_id as u8) {
continue;
}
let mut shard = EcVolumeShard::new(&dir, collection, vid, shard_id as u8);
shard.disk_type = ec_vol.disk_type.clone();
if let Err(e) = ec_vol.add_shard(shard) {
// The shard was dropped (its descriptors closed) inside the
// failed add. If this call just created the EcVolume and it
// holds nothing, remove it too — a zero-shard registration
// would advertise a mount that serves no data while pinning
// its descriptors.
let now_empty = ec_vol.shard_count() == 0;
if created && now_empty {
self.ec_volumes.remove(&vid);
}
return Err(VolumeError::Io(e));
}
ec_vol.add_shard(shard).map_err(VolumeError::Io)?;
crate::metrics::VOLUME_GAUGE
.with_label_values(&[collection, "ec_shards"])
.inc();
@@ -940,32 +921,27 @@ impl DiskLocation {
// double-count for those filenames.
let mut seen: HashSet<String> = HashSet::new();
let mut entries: Vec<String> = Vec::new();
// Keep only the shard and index files this scan acts on: a disk of
// regular volumes has millions of .dat/.idx/.vif names that would
// otherwise each cost a String here and a slot in the sort below.
let mut collect = |dir: &str| -> io::Result<()> {
for ent in fs::read_dir(dir)? {
for ent in fs::read_dir(&self.directory)? {
let ent = ent?;
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
if seen.insert(name.clone()) {
entries.push(name);
}
}
if self.idx_directory != self.directory {
for ent in fs::read_dir(&self.idx_directory)? {
let ent = ent?;
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let ext = &name[dot..];
if parse_ec_shard_extension(ext).is_none() && ext != ".ecx" {
continue;
}
if seen.insert(name.clone()) {
entries.push(name);
}
}
Ok(())
};
collect(&self.directory)?;
if self.idx_directory != self.directory {
collect(&self.idx_directory)?;
}
entries.sort();
@@ -1798,94 +1774,6 @@ mod tests {
}
}
/// A refused shard (a 0-byte file beside an index with entries) on the
/// FIRST mount of a volume must not leave the just-created zero-shard
/// EcVolume registered — it would advertise a mount serving no data,
/// pin the .ecx/.ecj descriptors, and make placement's mounted tier
/// prefer this disk. A volume that already holds shards keeps them
/// (the mount RPC's pre-existing first-error-aborts contract).
#[test]
fn test_mount_ec_shards_refused_shard_removes_created_empty_volume() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
// An index with one 16-byte entry and a 0-byte shard: the mount is
// refused, and the EcVolume created for it must be unregistered.
std::fs::write(format!("{}/pics_9.ecx", dir), [0u8; 16]).unwrap();
std::fs::write(format!("{}/pics_9.ec00", dir), b"").unwrap();
let err = loc
.mount_ec_shards(VolumeId(9), "pics", &[0], "")
.expect_err("a 0-byte shard beside an index with entries must refuse the mount");
assert!(
err.to_string().contains("empty (0 bytes)"),
"want the empty-shard refusal, got: {}",
err
);
assert!(
loc.find_ec_volume(VolumeId(9)).is_none(),
"a refused first mount must not leave a zero-shard EcVolume registered",
);
// With a valid shard mounted, a later refused shard keeps the
// existing registration intact.
std::fs::write(format!("{}/pics_9.ec01", dir), b"good bytes").unwrap();
loc.mount_ec_shards(VolumeId(9), "pics", &[1], "").unwrap();
loc.mount_ec_shards(VolumeId(9), "pics", &[0], "")
.expect_err("the 0-byte shard stays refused");
assert_eq!(
loc.find_ec_volume(VolumeId(9)).map(|v| v.shard_count()),
Some(1),
"an existing volume keeps its valid shards when a later shard is refused",
);
}
/// A mount retry re-listing an already mounted shard must keep the
/// existing registration and not bump the ec_shards gauge — the Rust
/// twin of Go's AddEcVolumeShard added=false handling.
#[test]
fn test_mount_ec_shards_duplicate_keeps_registration_and_gauge() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut loc = DiskLocation::new(
dir,
dir,
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
// A collection name unique to this test: the gauge is process-global
// and sibling tests running in parallel touch other labels.
std::fs::write(format!("{}/dupmount_11.ec00", dir), b"shard bytes").unwrap();
let gauge = crate::metrics::VOLUME_GAUGE.with_label_values(&["dupmount", "ec_shards"]);
let before = gauge.get();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "").unwrap();
loc.mount_ec_shards(VolumeId(11), "dupmount", &[0], "")
.expect("a duplicate mount must succeed as a no-op");
assert_eq!(
loc.find_ec_volume(VolumeId(11)).map(|v| v.shard_count()),
Some(1),
);
assert_eq!(
gauge.get(),
before + 1.0,
"the duplicate mount must not bump the ec_shards gauge",
);
}
#[test]
fn test_disk_location_persists_directory_uuid_and_tags() {
let tmp = TempDir::new().unwrap();
@@ -465,28 +465,6 @@ pub fn resolve_status(
}
}
/// Whether a generation-matching sidecar agrees with the geometry the volume is
/// mounted with. Both files record the layout the generation was encoded with,
/// so a disagreement means one of them is wrong and reads through the other
/// would land at the wrong shard offsets — the caller fails the mount rather
/// than merely dropping protection. A sidecar that records no EC config has
/// nothing to contradict.
pub fn geometry_matches(
prot: &EcBitrotProtection,
data_shards: usize,
parity_shards: usize,
block_size: i64,
) -> bool {
match &prot.ec_shard_config {
None => true,
Some(cfg) => {
cfg.data_shards as usize == data_shards
&& cfg.parity_shards as usize == parity_shards
&& cfg.block_size == block_size
}
}
}
/// Returns the [`EcShardChecksums`] entry for a shard id, or `None`.
pub fn shard_checksums(prot: &EcBitrotProtection, shard_id: u32) -> Option<&EcShardChecksums> {
prot.shards.iter().find(|s| s.shard_id == shard_id)
@@ -561,12 +539,11 @@ fn read_full_at(f: &File, buf: &mut [u8], offset: u64) -> io::Result<()> {
/// Builds the `EcShardConfig` proto for the given layout. The bitrot sidecar
/// carries its own top-level encode_uuid, so the nested config leaves it empty.
pub fn ec_shard_config(data_shards: u32, parity_shards: u32, block_size: i64) -> EcShardConfig {
pub fn ec_shard_config(data_shards: u32, parity_shards: u32) -> EcShardConfig {
EcShardConfig {
data_shards,
parity_shards,
encode_ts_ns: 0,
block_size,
}
}
@@ -590,7 +567,6 @@ mod tests {
data_shards: 10,
parity_shards: 4,
encode_ts_ns: 0,
block_size: 0,
}),
shards: vec![
EcShardChecksums {
@@ -736,7 +712,7 @@ mod tests {
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
generation: 0,
ec_shard_config: Some(ec_shard_config(10, 4, 0)),
ec_shard_config: Some(ec_shard_config(10, 4)),
shards: vec![EcShardChecksums {
shard_id: 0,
covered_size: covered,
@@ -774,7 +750,7 @@ mod tests {
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
generation: 0,
ec_shard_config: Some(ec_shard_config(10, 4, 0)),
ec_shard_config: Some(ec_shard_config(10, 4)),
shards: vec![EcShardChecksums {
shard_id: 0,
covered_size: 5,
@@ -818,7 +794,7 @@ mod tests {
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
generation: 0,
ec_shard_config: Some(ec_shard_config(10, 4, 0)),
ec_shard_config: Some(ec_shard_config(10, 4)),
shards,
encode_uuid: vec![0u8; 16],
}
@@ -80,8 +80,6 @@ pub fn write_dat_file_from_shards(
dat_file_size: i64,
encoded_dat_file_size: i64,
data_shards: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let dirs: Vec<String> = (0..data_shards).map(|_| dir.to_string()).collect();
write_dat_file_from_shards_with_dirs(
@@ -92,8 +90,6 @@ pub fn write_dat_file_from_shards(
encoded_dat_file_size,
data_shards,
&dirs,
large_block_size,
small_block_size,
)
}
@@ -117,10 +113,7 @@ pub fn write_dat_file_from_shards(
/// boundary, and deriving the layout from the shrunk extent would read
/// the shards in the wrong block order. Pass zero when the .vif does
/// not record the encode-time size to infer the layout from the shard
/// size. `large_block_size`/`small_block_size` are the volume's shard
/// block layout, e.g. `EcVolume::large_block_size()` /
/// `small_block_size()` from its .vif EC config.
#[allow(clippy::too_many_arguments)]
/// size.
pub fn write_dat_file_from_shards_with_dirs(
dat_dir: &str,
collection: &str,
@@ -129,8 +122,6 @@ pub fn write_dat_file_from_shards_with_dirs(
encoded_dat_file_size: i64,
data_shards: usize,
shard_dirs: &[String],
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
write_dat_file(
dat_dir,
@@ -140,8 +131,8 @@ pub fn write_dat_file_from_shards_with_dirs(
encoded_dat_file_size,
data_shards,
shard_dirs,
large_block_size,
small_block_size,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
)
}
@@ -409,7 +400,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.write_needle(&mut n, true).unwrap();
}
v.sync_to_disk().unwrap();
let original_dat_size = v.dat_file_size().unwrap();
@@ -421,9 +412,7 @@ mod tests {
// Encode to EC
let data_shards = 10;
let parity_shards = 4;
let block_size =
ec_encoder::write_ec_files(dir, dir, "", VolumeId(1), data_shards, parity_shards)
.unwrap();
ec_encoder::write_ec_files(dir, dir, "", VolumeId(1), data_shards, parity_shards).unwrap();
// Delete original .dat and .idx
std::fs::remove_file(format!("{}/1.dat", dir)).unwrap();
@@ -437,8 +426,6 @@ mod tests {
original_dat_size as i64,
original_dat_size as i64,
data_shards,
block_size as usize,
block_size as usize,
)
.unwrap();
write_idx_file_from_ec_index(dir, "", VolumeId(1)).unwrap();
@@ -485,16 +472,7 @@ mod tests {
let dir = tmp.path().to_str().unwrap();
// No shard files exist, so de-striping must fail and publish nothing:
// neither the final .dat nor a partial .dat.tmp may remain.
let res = write_dat_file_from_shards(
dir,
"",
VolumeId(7),
100,
100,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
);
let res = write_dat_file_from_shards(dir, "", VolumeId(7), 100, 100, 10);
assert!(res.is_err());
assert!(!std::path::Path::new(&format!("{}/7.dat", dir)).exists());
assert!(!std::path::Path::new(&format!("{}/7.dat.tmp", dir)).exists());
@@ -547,7 +525,6 @@ mod tests {
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
)
@@ -639,7 +616,6 @@ mod tests {
&mut builders,
data_shards,
parity_shards,
SMALL,
LARGE,
SMALL,
)
@@ -25,10 +25,6 @@ use crate::storage::volume::volume_file_name;
///
/// Creates .ec00-.ec13 files in the same directory.
/// Also creates a sorted .ecx index from the .idx file.
///
/// Always encodes with the uniform block layout, sized for this .dat, and
/// returns the block size so the caller can persist it to .vif. Mirrors Go's
/// WriteEcFiles.
pub fn write_ec_files(
dir: &str,
idx_dir: &str,
@@ -36,7 +32,7 @@ pub fn write_ec_files(
volume_id: VolumeId,
data_shards: usize,
parity_shards: usize,
) -> io::Result<i64> {
) -> io::Result<()> {
let base = volume_file_name(dir, collection, volume_id);
let dat_path = format!("{}.dat", base);
let idx_base = volume_file_name(idx_dir, collection, volume_id);
@@ -70,7 +66,7 @@ pub fn write_ec_files(
.map(|_| ShardChecksumBuilder::new(DEFAULT_BITROT_BLOCK_SIZE as i64))
.collect();
let block_size = uniform_block_size(dat_size, data_shards);
// Encode in large blocks, then small blocks
encode_dat_file(
&dat_file,
dat_size,
@@ -79,9 +75,8 @@ pub fn write_ec_files(
&mut builders,
data_shards,
parity_shards,
ENCODE_BUFFER_SIZE,
block_size as usize,
block_size as usize,
ERASURE_CODING_LARGE_BLOCK_SIZE,
ERASURE_CODING_SMALL_BLOCK_SIZE,
)?;
// Close all shards
@@ -108,7 +103,6 @@ pub fn write_ec_files(
ec_shard_config: Some(ec_bitrot::ec_shard_config(
data_shards as u32,
parity_shards as u32,
block_size,
)),
shards: shard_checksums,
encode_uuid: ec_bitrot::new_encode_uuid(),
@@ -126,19 +120,7 @@ pub fn write_ec_files(
);
}
Ok(block_size)
}
/// uniform_block_size returns the per-shard block size of the uniform layout
/// for a .dat of the given size: ceil(dat_file_size/data_shards) rounded up to
/// a whole small block. For every input this equals the legacy layout's padded
/// shard size, so only the byte placement differs between the two layouts,
/// never the shard length. Mirrors Go's UniformBlockSize.
pub fn uniform_block_size(dat_file_size: i64, data_shards: usize) -> i64 {
let small = ERASURE_CODING_SMALL_BLOCK_SIZE as i64;
let per_shard = (dat_file_size + data_shards as i64 - 1) / data_shards as i64;
let blocks = ((per_shard + small - 1) / small).max(1);
blocks * small
Ok(())
}
/// Rebuild missing EC shard files from existing shards using Reed-Solomon reconstruct.
@@ -388,7 +370,7 @@ pub fn verify_ec_shards(
}
/// Write sorted .ecx index from .idx file.
pub(crate) fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::Result<()> {
fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::Result<()> {
if !std::path::Path::new(idx_path).exists() {
return Err(io::Error::new(
io::ErrorKind::NotFound,
@@ -439,8 +421,6 @@ pub fn rebuild_ecx_file(
collection: &str,
volume_id: VolumeId,
data_shards: usize,
block_size: i64,
dat_file_size: i64,
additional_dirs: &[&str],
) -> io::Result<()> {
use crate::storage::needle::needle::get_actual_size;
@@ -483,42 +463,10 @@ pub fn rebuild_ecx_file(
// Determine total logical data size from shard sizes
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
let total_data_size = shard_size as i64 * data_shards as i64;
// The volume's shard block layout: the .vif-recorded uniform block size,
// or the legacy two-tier sizes when 0. The row count comes from the shard
// length; -1 disambiguates a legacy shard that is an exact large-block
// multiple (mirrors the ecdFileSize-1 fallback in the read path).
let (large_block, small_block) = if block_size > 0 {
(block_size, block_size)
} else {
(
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
)
};
// The row count the de-stripe walks with. The encode-time .dat size is the
// authority — the same value the read path divides by data_shards — and the
// padded extent is only a fallback: under the legacy layout a shard that is
// an exact large-block multiple reads as one row too many, which
// re-interprets its last large row as small blocks and scrambles the
// recovered offsets. Subtracting one keeps that fallback on the safe side of
// the boundary, exactly as the read path's own fallback does.
let locate_shard_size = if dat_file_size > 0 {
dat_file_size / data_shards as i64
} else {
(shard_size as i64 - 1).max(0)
};
// Read version from superblock (first byte of logical data)
let mut sb_buf = [0u8; SUPER_BLOCK_SIZE];
read_from_data_shards(
&shards,
&mut sb_buf,
0,
data_shards,
locate_shard_size,
large_block,
small_block,
)?;
read_from_data_shards(&shards, &mut sb_buf, 0, data_shards)?;
let version = Version(sb_buf[0]);
// Walk needles starting after superblock
@@ -527,30 +475,10 @@ pub fn rebuild_ecx_file(
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
while offset + header_size as i64 <= total_data_size {
// Read needle header (cookie + needle_id + size = 16 bytes).
// A read failure is NOT the end of the data — every offset in
// range maps into the shards, so an error means a truncated or
// unreadable shard. Publishing the entries collected so far as
// a successful .ecx would hand out a silently incomplete
// recovery index; propagate instead. (The scan still ends
// normally on the zero-cookie tail below.)
// Read needle header (cookie + needle_id + size = 16 bytes)
let mut header_buf = [0u8; NEEDLE_HEADER_SIZE];
if let Err(e) = read_from_data_shards(
&shards,
&mut header_buf,
offset as u64,
data_shards,
locate_shard_size,
large_block,
small_block,
) {
for s in &mut shards {
s.close();
}
return Err(io::Error::new(
e.kind(),
format!("scan needle header at offset {}: {}", offset, e),
));
if read_from_data_shards(&shards, &mut header_buf, offset as u64, data_shards).is_err() {
break;
}
let cookie = Cookie::from_bytes(&header_buf[..COOKIE_SIZE]);
@@ -604,83 +532,58 @@ pub fn rebuild_ecx_file(
Ok(())
}
/// Read bytes from EC data shards at a logical offset in the .dat file,
/// resolving the shard/offset through the volume's block layout via
/// locate_data — the same mapping the read path uses.
#[allow(clippy::too_many_arguments)]
/// Read bytes from EC data shards at a logical offset in the .dat file.
fn read_from_data_shards(
shards: &[EcVolumeShard],
buf: &mut [u8],
logical_offset: u64,
data_shards: usize,
locate_shard_size: i64,
large_block_size: i64,
small_block_size: i64,
) -> io::Result<()> {
let intervals = crate::storage::erasure_coding::ec_locate::locate_data(
logical_offset as i64,
Size(buf.len() as i32),
locate_shard_size,
data_shards as u32,
large_block_size,
small_block_size,
);
let mut bytes_read = 0usize;
for interval in &intervals {
let (shard_id, shard_offset) =
interval.to_shard_id_and_offset(data_shards as u32, large_block_size, small_block_size);
if shard_id as usize >= data_shards {
let small_block = ERASURE_CODING_SMALL_BLOCK_SIZE as u64;
let data_shards_u64 = data_shards as u64;
let mut bytes_read = 0u64;
let mut remaining = buf.len() as u64;
let mut current_offset = logical_offset;
while remaining > 0 {
// Determine which shard and at what shard-offset this logical offset maps to.
// The data is interleaved: large blocks first, then small blocks.
// For simplicity, use the small block size for all calculations since
// large blocks are multiples of small blocks.
let row_size = small_block * data_shards_u64;
let row_index = current_offset / row_size;
let row_offset = current_offset % row_size;
let shard_index = (row_offset / small_block) as usize;
let shard_offset = row_index * small_block + (row_offset % small_block);
if shard_index >= data_shards {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
"shard index out of range",
));
}
let to_read = interval.size as usize;
let dest = &mut buf[bytes_read..bytes_read + to_read];
// Exact-read semantics: read_at may legally return fewer bytes
// than requested, and treating a short read as complete leaves
// the tail of `dest` as whatever the buffer held before. Loop
// until filled; zero bytes inside the mapped range means the
// shard is truncated — an error, not an end.
let mut filled = 0usize;
while filled < to_read {
let n = shards[shard_id as usize]
.read_at(&mut dest[filled..], shard_offset as u64 + filled as u64)?;
if n == 0 {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"short read from data shard {}: {} of {} bytes at offset {}",
shard_id, filled, to_read, shard_offset
),
));
}
filled += n;
}
bytes_read += to_read;
}
if bytes_read != buf.len() {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
"short read from data shards",
));
// How many bytes can we read from this position in this shard block
let bytes_left_in_block = small_block - (row_offset % small_block);
let to_read = remaining.min(bytes_left_in_block) as usize;
let dest = &mut buf[bytes_read as usize..bytes_read as usize + to_read];
shards[shard_index].read_at(dest, shard_offset)?;
bytes_read += to_read as u64;
remaining -= to_read as u64;
current_offset += to_read as u64;
}
Ok(())
}
/// Buffer size for one encode sub-batch per shard, mirroring Go's 256KB
/// bufferSize in WriteEcFiles. A block is processed in block_size/buffer_size
/// sub-batches, so memory stays at total_shards * 256KB no matter how large
/// the uniform block is.
const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
/// Encode the .dat file data into shard files.
///
/// Uses a two-phase approach matching Go's ec_encoder.go:
/// 1. Process as many large blocks as possible
/// 2. Process remaining data with small blocks
///
/// `buffer_size` must divide both block sizes.
/// 1. Process as many large blocks (1GB) as possible
/// 2. Process remaining data with small blocks (1MB)
#[allow(clippy::too_many_arguments)]
pub(crate) fn encode_dat_file(
dat_file: &File,
@@ -690,50 +593,44 @@ pub(crate) fn encode_dat_file(
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
buffer_size: usize,
large_block_size: usize,
small_block_size: usize,
) -> io::Result<()> {
let total_shards = data_shards + parity_shards;
let mut buffers: Vec<Vec<u8>> = (0..total_shards)
.map(|_| vec![0u8; buffer_size])
.collect();
let mut remaining = dat_size;
let mut offset: u64 = 0;
// Phase 1: process whole large-block rows while enough data remains
// Phase 1: Process large blocks (1GB each) while enough data remains
let large_row_size = large_block_size * data_shards;
while remaining >= large_row_size as i64 {
encode_data(
encode_one_batch(
dat_file,
offset,
large_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
parity_shards,
)?;
offset += large_row_size as u64;
remaining -= large_row_size as i64;
}
// Phase 2: process remaining data with small blocks
// Phase 2: Process remaining data with small blocks (1MB each)
let small_row_size = small_block_size * data_shards;
while remaining > 0 {
let to_process = remaining.min(small_row_size as i64);
encode_data(
encode_one_batch(
dat_file,
offset,
small_block_size,
rs,
&mut buffers,
shards,
builders,
data_shards,
parity_shards,
)?;
offset += to_process as u64;
remaining -= to_process;
@@ -742,71 +639,61 @@ pub(crate) fn encode_dat_file(
Ok(())
}
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
/// arbitrarily large blocks never require block-sized allocations. Mirrors
/// Go's encodeData.
#[allow(clippy::too_many_arguments)]
fn encode_data(
dat_file: &File,
row_offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
) -> io::Result<()> {
let buffer_size = buffers[0].len();
if block_size % buffer_size != 0 {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"unexpected block size {} buffer size {}",
block_size, buffer_size
),
));
}
let batch_count = block_size / buffer_size;
for b in 0..batch_count {
encode_one_batch(
dat_file,
row_offset + (b * buffer_size) as u64,
block_size,
rs,
buffers,
shards,
builders,
data_shards,
)?;
}
Ok(())
}
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
/// this row. Mirrors Go's encodeDataOneBatch.
#[allow(clippy::too_many_arguments)]
/// Encode one batch (row) of data.
fn encode_one_batch(
dat_file: &File,
offset: u64,
block_size: usize,
rs: &ReedSolomon,
buffers: &mut [Vec<u8>],
shards: &mut [EcVolumeShard],
builders: &mut [ShardChecksumBuilder],
data_shards: usize,
parity_shards: usize,
) -> io::Result<()> {
// Read data shards from the .dat file, zero-filling past EOF — the buffers
// are reused across batches, so the tail must be cleared explicitly.
let total_shards = data_shards + parity_shards;
// Each batch allocates block_size * total_shards bytes.
// With large blocks (1 GiB) this is 14 GiB -- guard against OOM.
let total_alloc = block_size.checked_mul(total_shards).ok_or_else(|| {
io::Error::new(
io::ErrorKind::InvalidInput,
"block_size * shard count overflows usize",
)
})?;
// Large-block encoding uses 1 GiB * 14 shards = 14 GiB; allow up to 16 GiB.
const MAX_BATCH_ALLOC: usize = 16 * 1024 * 1024 * 1024; // 16 GiB safety limit
if total_alloc > MAX_BATCH_ALLOC {
return Err(io::Error::new(
io::ErrorKind::InvalidInput,
format!(
"batch allocation too large ({} bytes, limit {} bytes); block_size={} shards={}",
total_alloc, MAX_BATCH_ALLOC, block_size, total_shards,
),
));
}
// Allocate buffers for all shards
let mut buffers: Vec<Vec<u8>> = (0..total_shards).map(|_| vec![0u8; block_size]).collect();
// Read data shards from .dat file
for i in 0..data_shards {
let read_offset = offset + (i * block_size) as u64;
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
for b in buffers[i][n..].iter_mut() {
*b = 0;
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buffers[i], read_offset)?;
}
#[cfg(not(unix))]
{
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(read_offset))?;
f.read(&mut buffers[i])?;
}
}
// Encode parity shards
rs.encode(&mut *buffers).map_err(|e| {
rs.encode(&mut buffers).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon encode: {:?}", e),
@@ -823,29 +710,6 @@ fn encode_one_batch(
Ok(())
}
/// Read into `buf` at `offset` until it is full or EOF; returns bytes read.
fn read_at_most(dat_file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
let mut n = 0;
while n < buf.len() {
#[cfg(unix)]
let r = {
use std::os::unix::fs::FileExt;
dat_file.read_at(&mut buf[n..], offset + n as u64)?
};
#[cfg(not(unix))]
let r = {
let mut f = dat_file.try_clone()?;
f.seek(SeekFrom::Start(offset + n as u64))?;
f.read(&mut buf[n..])?
};
if r == 0 {
break;
}
n += r;
}
Ok(n)
}
#[cfg(test)]
mod tests {
use super::*;
@@ -882,7 +746,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.write_needle(&mut n, true).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -934,7 +798,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut needle, true, false).unwrap();
v.write_needle(&mut needle, true).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1011,7 +875,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.write_needle(&mut n, true).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1156,7 +1020,7 @@ mod tests {
// Without additional_dirs, rebuild must fail: shards 1, 3, 6 are not
// in primary and the full logical .dat content can't be reconstructed.
let res = rebuild_ecx_file(&primary, "", VolumeId(1), 10, 0, 0, &[]);
let res = rebuild_ecx_file(&primary, "", VolumeId(1), 10, &[]);
assert!(
res.is_err(),
"ecx rebuild without additional_dirs must fail when data shards are on another disk"
@@ -1167,7 +1031,7 @@ mod tests {
);
// With additional_dirs pointing at the secondary, the rebuild must succeed.
rebuild_ecx_file(&primary, "", VolumeId(1), 10, 0, 0, &[secondary.as_str()]).unwrap();
rebuild_ecx_file(&primary, "", VolumeId(1), 10, &[secondary.as_str()]).unwrap();
assert!(
std::path::Path::new(&ecx_path).exists(),
@@ -1179,118 +1043,6 @@ mod tests {
);
}
// A uniform-layout volume (block size > 1MiB) must have its .ecx rebuilt
// through the recorded geometry; the legacy 1MiB mapping would scan
// garbage past the first block boundary.
#[test]
fn test_rebuild_ecx_file_uniform_layout() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(2),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
for i in 1u64..=12 {
let data: Vec<u8> = (0..2 << 20)
.map(|b| ((b as u64).wrapping_mul(2654435761).wrapping_add(i) >> 8) as u8)
.collect();
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
let block_size = write_ec_files(&dir, &dir, "", VolumeId(2), 10, 4).unwrap();
assert!(
block_size > ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
"fixture must diverge from the legacy layout"
);
let ecx_path = format!("{}/2.ecx", dir);
let canonical = std::fs::read(&ecx_path).unwrap();
std::fs::remove_file(&ecx_path).unwrap();
rebuild_ecx_file(&dir, "", VolumeId(2), 10, block_size, 0, &[]).unwrap();
let rebuilt = std::fs::read(&ecx_path).unwrap();
assert_eq!(canonical, rebuilt, "rebuilt .ecx must match the encode-time .ecx");
}
// A truncated data shard must FAIL the .ecx rebuild, not publish the
// entries scanned so far as a successful (silently incomplete) index.
#[test]
fn test_rebuild_ecx_file_fails_on_truncated_shard() {
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::volume::Volume;
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap().to_string();
let mut v = Volume::new(
&dir,
&dir,
"",
VolumeId(3),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
for i in 1u64..=12 {
let data: Vec<u8> = (0..2 << 20)
.map(|b| ((b as u64).wrapping_mul(2654435761).wrapping_add(i) >> 8) as u8)
.collect();
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.clone(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
let block_size = write_ec_files(&dir, &dir, "", VolumeId(3), 10, 4).unwrap();
let ecx_path = format!("{}/3.ecx", dir);
std::fs::remove_file(&ecx_path).unwrap();
// Truncate shard 0 to just the superblock: the scan's very first
// needle-header read (offset SUPER_BLOCK_SIZE, shard 0 under the
// uniform layout) lands in the missing region. The pre-fix code
// broke the scan there and published an EMPTY .ecx as success.
let shard_path = format!("{}/3.ec00", dir);
let f = std::fs::OpenOptions::new()
.write(true)
.open(&shard_path)
.unwrap();
f.set_len(crate::storage::super_block::SUPER_BLOCK_SIZE as u64)
.unwrap();
drop(f);
let res = rebuild_ecx_file(&dir, "", VolumeId(3), 10, block_size, 0, &[]);
assert!(res.is_err(), "rebuild over a truncated shard must fail");
assert!(
!std::path::Path::new(&ecx_path).exists(),
"a failed rebuild must not leave a partial .ecx behind"
);
}
#[test]
fn test_reed_solomon_basic() {
let data_shards = 10;
@@ -1364,7 +1116,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.write_needle(&mut n, true).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1446,7 +1198,7 @@ mod tests {
data_size: 5,
..Needle::default()
};
v.write_needle(&mut n, true, false).unwrap();
v.write_needle(&mut n, true).unwrap();
v.sync_to_disk().unwrap();
v.close();
@@ -18,26 +18,21 @@ pub struct Interval {
}
impl Interval {
pub fn to_shard_id_and_offset(
&self,
data_shards: u32,
large_block_size: i64,
small_block_size: i64,
) -> (ShardId, i64) {
pub fn to_shard_id_and_offset(&self, data_shards: u32) -> (ShardId, i64) {
let data_shards_usize = data_shards as usize;
let shard_id = (self.block_index % data_shards_usize) as ShardId;
let row_index = self.block_index / data_shards_usize;
let block_size = if self.is_large_block {
large_block_size
ERASURE_CODING_LARGE_BLOCK_SIZE as i64
} else {
small_block_size
ERASURE_CODING_SMALL_BLOCK_SIZE as i64
};
let mut offset = row_index as i64 * block_size + self.inner_block_offset;
if !self.is_large_block {
// Small blocks come after large blocks in the shard file
offset += self.large_block_rows_count as i64 * large_block_size;
offset += self.large_block_rows_count as i64 * ERASURE_CODING_LARGE_BLOCK_SIZE as i64;
}
(shard_id, offset)
@@ -47,14 +42,7 @@ impl Interval {
/// Locate the EC shard intervals needed to read data at the given offset and size.
///
/// `shard_size` is the size of a single shard file.
pub fn locate_data(
offset: i64,
size: Size,
shard_size: i64,
data_shards: u32,
large_block_size: i64,
small_block_size: i64,
) -> Vec<Interval> {
pub fn locate_data(offset: i64, size: Size, shard_size: i64, data_shards: u32) -> Vec<Interval> {
let mut intervals = Vec::new();
let data_size = size.0 as i64;
@@ -62,14 +50,17 @@ pub fn locate_data(
return intervals;
}
let large_block_size = ERASURE_CODING_LARGE_BLOCK_SIZE as i64;
let small_block_size = ERASURE_CODING_SMALL_BLOCK_SIZE as i64;
let large_row_size = large_block_size * data_shards as i64;
let small_row_size = small_block_size * data_shards as i64;
// Number of large block rows. Mirrors Go's shardDatSize/largeBlockLength:
// the caller's ecd-size fallback already subtracts 1 to disambiguate the
// exact-multiple case, so no further -1 here — a shard size that IS an
// exact multiple (dat_file_size path) means real full large rows.
let n_large_block_rows = (shard_size / large_block_size) as usize;
// Number of large block rows
let n_large_block_rows = if shard_size > 0 {
((shard_size - 1) / large_block_size) as usize
} else {
0
};
let large_section_size = n_large_block_rows as i64 * large_row_size;
let mut remaining_offset = offset;
@@ -159,11 +150,7 @@ mod tests {
is_large_block: true,
large_block_rows_count: 1,
};
let (shard_id, offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let (shard_id, offset) = interval.to_shard_id_and_offset(data_shards);
assert_eq!(shard_id, 0);
assert_eq!(offset, 100);
@@ -175,11 +162,7 @@ mod tests {
is_large_block: true,
large_block_rows_count: 1,
};
let (shard_id, _offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let (shard_id, _offset) = interval.to_shard_id_and_offset(data_shards);
assert_eq!(shard_id, 5);
// Block index 12 (data_shards=10) → row_index 1, shard_id 2
@@ -190,11 +173,7 @@ mod tests {
is_large_block: true,
large_block_rows_count: 5,
};
let (shard_id, offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let (shard_id, offset) = interval.to_shard_id_and_offset(data_shards);
assert_eq!(shard_id, 2); // 12 % 10 = 2
assert_eq!(offset, large_block_size + 200); // row 1 offset + inner_block_offset
@@ -206,11 +185,7 @@ mod tests {
is_large_block: true,
large_block_rows_count: 2,
};
let (shard_id, offset) = interval.to_shard_id_and_offset(
data_shards,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let (shard_id, offset) = interval.to_shard_id_and_offset(data_shards);
assert_eq!(shard_id, 0);
assert_eq!(offset, ERASURE_CODING_LARGE_BLOCK_SIZE as i64); // row 1 offset
}
@@ -218,14 +193,7 @@ mod tests {
#[test]
fn test_locate_data_small_file() {
// Small file: 100 bytes at offset 50, shard size = 1MB
let intervals = locate_data(
50,
Size(100),
1024 * 1024,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let intervals = locate_data(50, Size(100), 1024 * 1024, 10);
assert!(!intervals.is_empty());
// Should be a single small block interval (no large block rows for 1MB shard)
@@ -235,14 +203,7 @@ mod tests {
#[test]
fn test_locate_data_empty() {
let intervals = locate_data(
0,
Size(0),
1024 * 1024,
10,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let intervals = locate_data(0, Size(0), 1024 * 1024, 10);
assert!(intervals.is_empty());
}
@@ -255,11 +216,7 @@ mod tests {
is_large_block: false,
large_block_rows_count: 2,
};
let (_shard_id, offset) = interval.to_shard_id_and_offset(
10,
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
);
let (_shard_id, offset) = interval.to_shard_id_and_offset(10);
// Should be after 2 large block rows
assert_eq!(offset, 2 * ERASURE_CODING_LARGE_BLOCK_SIZE as i64);
}
File diff suppressed because it is too large Load Diff
+41 -74
View File
@@ -14,10 +14,7 @@ use std::path::Path;
use std::sync::atomic::{AtomicI64, AtomicU64, Ordering};
mod compact_map;
pub mod file_pool;
pub mod sorted_file;
use compact_map::CompactMap;
use sorted_file::SortedFileNeedleMap;
use redb::{Database, Durability, ReadableDatabase, ReadableTable, TableDefinition};
@@ -753,12 +750,10 @@ impl RedbNeedleMap {
Ok(())
}
/// Look up a needle. A redb failure is an ERROR, not an absent needle:
/// answering "not found" would turn a database problem into a read miss
/// and let a delete report success without recording a tombstone.
pub fn get(&self, key: NeedleId) -> io::Result<Option<NeedleValue>> {
/// Look up a needle.
pub fn get(&self, key: NeedleId) -> Option<NeedleValue> {
let key_u64: u64 = key.into();
self.get_internal(key_u64)
self.get_internal(key_u64).ok().flatten()
}
/// Internal get that returns io::Result for error propagation.
@@ -944,31 +939,33 @@ impl RedbNeedleMap {
}
/// Collect all entries as a Vec for iteration (used by volume.rs iter patterns).
pub fn collect_entries(&self) -> io::Result<Vec<(NeedleId, NeedleValue)>> {
pub fn collect_entries(&self) -> Vec<(NeedleId, NeedleValue)> {
let mut result = Vec::new();
let txn: redb::ReadTransaction = self
.db
.begin_read()
.map_err(|e| io::Error::other(format!("redb begin_read: {e}")))?;
let table = txn
.open_table(NEEDLE_TABLE)
.map_err(|e| io::Error::other(format!("redb open_table: {e}")))?;
let iter = table
.iter()
.map_err(|e| io::Error::other(format!("redb iter: {e}")))?;
let txn: redb::ReadTransaction = match self.db.begin_read() {
Ok(t) => t,
Err(_) => return result,
};
let table = match txn.open_table(NEEDLE_TABLE) {
Ok(t) => t,
Err(_) => return result,
};
let iter = match table.iter() {
Ok(i) => i,
Err(_) => return result,
};
for entry in iter {
let (key_guard, val_guard) =
entry.map_err(|e| io::Error::other(format!("redb entry: {e}")))?;
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if bytes.len() == PACKED_NEEDLE_VALUE_SIZE {
let mut arr = [0u8; PACKED_NEEDLE_VALUE_SIZE];
arr.copy_from_slice(bytes);
let nv = unpack_needle_value(&arr);
result.push((NeedleId(key_u64), nv));
if let Ok((key_guard, val_guard)) = entry {
let key_u64: u64 = key_guard.value();
let bytes: &[u8] = val_guard.value();
if bytes.len() == PACKED_NEEDLE_VALUE_SIZE {
let mut arr = [0u8; PACKED_NEEDLE_VALUE_SIZE];
arr.copy_from_slice(bytes);
let nv = unpack_needle_value(&arr);
result.push((NeedleId(key_u64), nv));
}
}
}
Ok(result)
result
}
}
@@ -980,10 +977,6 @@ impl RedbNeedleMap {
pub enum NeedleMap {
InMemory(CompactNeedleMap),
Redb(RedbNeedleMap),
/// Read-only volumes — including every cloud-tiered one — search the sorted
/// `.sdx` on disk instead of holding an index in RAM. Mirrors Go's
/// `SortedFileNeedleMap`.
SortedFile(SortedFileNeedleMap),
}
impl NeedleMap {
@@ -992,18 +985,14 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.put(key, offset, size),
NeedleMap::Redb(nm) => nm.put(key, offset, size),
NeedleMap::SortedFile(nm) => nm.put(key, offset, size),
}
}
/// Look up a needle. Disk- and database-backed maps report their own
/// failures rather than folding them into "not found" — see the notes on
/// `RedbNeedleMap::get` and `SortedFileNeedleMap::get`.
pub fn get(&self, key: NeedleId) -> io::Result<Option<NeedleValue>> {
/// Look up a needle.
pub fn get(&self, key: NeedleId) -> Option<NeedleValue> {
match self {
NeedleMap::InMemory(nm) => Ok(nm.get(key)),
NeedleMap::InMemory(nm) => nm.get(key),
NeedleMap::Redb(nm) => nm.get(key),
NeedleMap::SortedFile(nm) => nm.get(key),
}
}
@@ -1012,7 +1001,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.delete(key, offset),
NeedleMap::Redb(nm) => nm.delete(key, offset),
NeedleMap::SortedFile(nm) => nm.delete(key, offset),
}
}
@@ -1021,9 +1009,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.set_idx_file(file, offset),
NeedleMap::Redb(nm) => nm.set_idx_file(file, offset),
// The sorted map borrows its .idx per append, so there is no
// long-lived writer to install.
NeedleMap::SortedFile(_) => {}
}
}
@@ -1032,8 +1017,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.has_idx_writer(),
NeedleMap::Redb(nm) => nm.has_idx_writer(),
// Appends open the .idx on demand, so one is always available.
NeedleMap::SortedFile(_) => true,
}
}
@@ -1042,7 +1025,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.content_size(),
NeedleMap::Redb(nm) => nm.content_size(),
NeedleMap::SortedFile(nm) => nm.content_size(),
}
}
@@ -1051,7 +1033,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.deleted_size(),
NeedleMap::Redb(nm) => nm.deleted_size(),
NeedleMap::SortedFile(nm) => nm.deleted_size(),
}
}
@@ -1060,7 +1041,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.file_count(),
NeedleMap::Redb(nm) => nm.file_count(),
NeedleMap::SortedFile(nm) => nm.file_count(),
}
}
@@ -1069,7 +1049,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.deleted_count(),
NeedleMap::Redb(nm) => nm.deleted_count(),
NeedleMap::SortedFile(nm) => nm.deleted_count(),
}
}
@@ -1078,7 +1057,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.max_file_key(),
NeedleMap::Redb(nm) => nm.max_file_key(),
NeedleMap::SortedFile(nm) => nm.max_file_key(),
}
}
@@ -1089,7 +1067,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.max_needle_end(),
NeedleMap::Redb(nm) => nm.max_needle_end(),
NeedleMap::SortedFile(nm) => nm.max_needle_end(),
}
}
@@ -1098,7 +1075,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.index_file_size(),
NeedleMap::Redb(nm) => nm.index_file_size(),
NeedleMap::SortedFile(nm) => nm.index_file_size(),
}
}
@@ -1107,7 +1083,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.sync(),
NeedleMap::Redb(nm) => nm.sync(),
NeedleMap::SortedFile(nm) => nm.sync(),
}
}
@@ -1116,7 +1091,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.close(),
NeedleMap::Redb(nm) => nm.close(),
NeedleMap::SortedFile(nm) => nm.close(),
}
}
@@ -1125,7 +1099,6 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.save_to_idx(path),
NeedleMap::Redb(nm) => nm.save_to_idx(path),
NeedleMap::SortedFile(nm) => nm.save_to_idx(path),
}
}
@@ -1137,28 +1110,22 @@ impl NeedleMap {
match self {
NeedleMap::InMemory(nm) => nm.ascending_visit(f),
NeedleMap::Redb(nm) => nm.ascending_visit(f),
NeedleMap::SortedFile(nm) => nm.ascending_visit(f),
}
}
/// Iterate all entries. Returns a Vec of (NeedleId, NeedleValue) pairs.
/// For InMemory this collects via ascending visit; the disk-backed maps read
/// it back off disk, so a truncated .sdx or a redb read fault surfaces here
/// as an error. Compaction treats the result as the complete live set, so a
/// partial scan must never be mistaken for an empty tail.
pub fn iter_entries(&self) -> io::Result<Vec<(NeedleId, NeedleValue)>> {
/// For InMemory this collects via ascending visit; for Redb it reads from disk.
pub fn iter_entries(&self) -> Vec<(NeedleId, NeedleValue)> {
match self {
NeedleMap::InMemory(nm) => {
let mut entries = Vec::new();
// The visitor never fails, so neither can this.
let _ = nm.ascending_visit(|id, nv| {
entries.push((id, *nv));
Ok(())
});
Ok(entries)
entries
}
NeedleMap::Redb(nm) => nm.collect_entries(),
NeedleMap::SortedFile(nm) => nm.iter_entries(),
}
}
}
@@ -1313,13 +1280,13 @@ mod tests {
nm.put(NeedleId(2), Offset::from_actual_offset(128), Size(200))
.unwrap();
let v1 = nm.get(NeedleId(1)).unwrap().unwrap();
let v1 = nm.get(NeedleId(1)).unwrap();
assert_eq!(v1.size, Size(100));
let v2 = nm.get(NeedleId(2)).unwrap().unwrap();
let v2 = nm.get(NeedleId(2)).unwrap();
assert_eq!(v2.size, Size(200));
assert!(nm.get(NeedleId(99)).unwrap().is_none());
assert!(nm.get(NeedleId(99)).is_none());
}
#[test]
@@ -1344,7 +1311,7 @@ mod tests {
assert_eq!(nm.deleted_size(), 100);
// Deleted entry should have negated size
let nv = nm.get(NeedleId(1)).unwrap().unwrap();
let nv = nm.get(NeedleId(1)).unwrap();
assert_eq!(nv.size, Size(-100));
}
@@ -1417,9 +1384,9 @@ mod tests {
let mut cursor = Cursor::new(idx_data);
let nm = RedbNeedleMap::load_from_idx(db_path.to_str().unwrap(), &mut cursor, Version::current()).unwrap();
assert!(nm.get(NeedleId(1)).unwrap().is_some());
assert!(nm.get(NeedleId(2)).unwrap().is_none()); // deleted and removed
assert!(nm.get(NeedleId(3)).unwrap().is_some());
assert!(nm.get(NeedleId(1)).is_some());
assert!(nm.get(NeedleId(2)).is_none()); // deleted and removed
assert!(nm.get(NeedleId(3)).is_some());
assert_eq!(nm.file_count(), 2);
}
@@ -1536,7 +1503,7 @@ mod tests {
let mut nm = NeedleMap::InMemory(CompactNeedleMap::new());
nm.put(NeedleId(1), Offset::from_actual_offset(0), Size(100))
.unwrap();
assert_eq!(nm.get(NeedleId(1)).unwrap().unwrap().size, Size(100));
assert_eq!(nm.get(NeedleId(1)).unwrap().size, Size(100));
assert_eq!(nm.file_count(), 1);
}
@@ -1547,7 +1514,7 @@ mod tests {
let mut nm = NeedleMap::Redb(RedbNeedleMap::new(db_path.to_str().unwrap()).unwrap());
nm.put(NeedleId(1), Offset::from_actual_offset(0), Size(100))
.unwrap();
assert_eq!(nm.get(NeedleId(1)).unwrap().unwrap().size, Size(100));
assert_eq!(nm.get(NeedleId(1)).unwrap().size, Size(100));
assert_eq!(nm.file_count(), 1);
}
}
@@ -1,247 +0,0 @@
//! Bounded pool of open index-file descriptors.
//!
//! Read-only volumes — cloud-tiered ones above all — outnumber writable ones by
//! orders of magnitude on a large server, and a volume that pins its `.idx` and
//! `.sdx` for the life of the process costs two descriptors whether or not
//! anybody reads it. At ~600K volumes per server that alone exhausts any fd
//! limit. Neither file is needed except while a lookup is in flight, so
//! [`SortedFileNeedleMap`](super::sorted_file::SortedFileNeedleMap) borrows them
//! from this pool: an idle volume holds nothing, a busy one keeps its handles
//! hot rather than paying an `open()` per needle.
//!
//! Mirrors Go's `weed/storage/needle_map_file_pool.go`. Handles are handed out
//! as `Arc<File>`, so an eviction cannot close a descriptor a reader still
//! holds — the file closes when the last borrower drops its `Arc`.
use std::collections::{BTreeMap, HashMap};
use std::fs::{File, OpenOptions};
use std::io;
use std::sync::{Arc, Mutex, OnceLock};
/// Descriptors the pool keeps open. Matches Go's `maxPooledIndexFiles`.
pub const MAX_POOLED_INDEX_FILES: usize = 1024;
struct Entry {
file: Arc<File>,
tick: u64,
}
#[derive(Default)]
struct Inner {
entries: HashMap<String, Entry>,
/// Recency order, oldest tick first, so eviction is a `pop_first`.
order: BTreeMap<u64, String>,
next_tick: u64,
}
pub struct IndexFilePool {
capacity: usize,
inner: Mutex<Inner>,
}
/// Writable and read-only handles for the same path are pooled separately so a
/// read never depends on the file being openable for write — a volume served
/// off a read-only mount still answers lookups.
fn pool_key(path: &str, writable: bool) -> String {
if writable {
format!("{path}\0rw")
} else {
path.to_string()
}
}
impl IndexFilePool {
pub fn new(capacity: usize) -> Self {
IndexFilePool {
capacity: capacity.max(1),
inner: Mutex::new(Inner::default()),
}
}
/// Hand out an open handle for `path`, reusing the pooled one when there is
/// one. The descriptor lives as long as the returned `Arc`.
pub fn borrow(&self, path: &str, writable: bool) -> io::Result<Arc<File>> {
let key = pool_key(path, writable);
if let Some(file) = self.touch(&key) {
return Ok(file);
}
// Opened outside the lock: a cold open blocks on disk, and holding a
// process-wide mutex across it would serialize every volume's lookups.
let file = Arc::new(OpenOptions::new().read(true).write(writable).open(path)?);
Ok(self.insert(key, file))
}
/// Forget the pooled handles for `path`, so a later rename or delete of that
/// path cannot be served from a descriptor on the old inode.
pub fn discard(&self, path: &str) {
let mut inner = self.inner.lock().unwrap();
for key in [pool_key(path, false), pool_key(path, true)] {
if let Some(entry) = inner.entries.remove(&key) {
inner.order.remove(&entry.tick);
}
}
}
/// Descriptors currently pooled. Test-only visibility into the bound.
#[cfg(test)]
pub fn pooled_count(&self) -> usize {
self.inner.lock().unwrap().entries.len()
}
fn touch(&self, key: &str) -> Option<Arc<File>> {
let mut inner = self.inner.lock().unwrap();
let tick = inner.next_tick;
let entry = inner.entries.get_mut(key)?;
let file = entry.file.clone();
let old_tick = std::mem::replace(&mut entry.tick, tick);
inner.order.remove(&old_tick);
inner.order.insert(tick, key.to_string());
inner.next_tick += 1;
Some(file)
}
fn insert(&self, key: String, file: Arc<File>) -> Arc<File> {
let mut inner = self.inner.lock().unwrap();
if let Some(entry) = inner.entries.get(&key) {
// Another borrower opened the same path first; keep one descriptor.
return entry.file.clone();
}
let tick = inner.next_tick;
inner.next_tick += 1;
inner.order.insert(tick, key.clone());
inner.entries.insert(
key,
Entry {
file: file.clone(),
tick,
},
);
while inner.entries.len() > self.capacity {
let Some((_, oldest)) = inner.order.pop_first() else {
break;
};
inner.entries.remove(&oldest);
}
file
}
}
/// Process-wide pool shared by every read-only volume on this server.
pub fn pooled_index_files() -> &'static IndexFilePool {
static POOL: OnceLock<IndexFilePool> = OnceLock::new();
POOL.get_or_init(|| IndexFilePool::new(MAX_POOLED_INDEX_FILES))
}
/// Descriptors this process holds on `.idx`/`.sdx` files under `dir`, read from
/// `/proc/self/fd` where it exists and from `lsof` otherwise. `None` when
/// neither is available, so a caller can skip rather than assert vacuously.
#[cfg(test)]
pub(crate) fn open_index_fds(dir: &std::path::Path) -> Option<usize> {
let prefix = std::fs::canonicalize(dir).unwrap_or_else(|_| dir.to_path_buf());
let is_index = |target: &std::path::Path| {
target.starts_with(&prefix)
&& matches!(
target.extension().and_then(|e| e.to_str()),
Some("idx") | Some("sdx")
)
};
if let Ok(entries) = std::fs::read_dir("/proc/self/fd") {
return Some(
entries
.filter_map(|e| std::fs::read_link(e.ok()?.path()).ok())
.filter(|target| is_index(target))
.count(),
);
}
let out = std::process::Command::new("lsof")
.args(["-p", &std::process::id().to_string(), "-F", "n"])
.output()
.ok()?;
Some(
String::from_utf8_lossy(&out.stdout)
.lines()
.filter_map(|line| line.strip_prefix('n'))
.filter(|line| is_index(std::path::Path::new(line)))
.count(),
)
}
#[cfg(test)]
mod tests {
use super::*;
use std::io::Write;
fn write_file(dir: &std::path::Path, name: &str, contents: &[u8]) -> String {
let path = dir.join(name);
let mut f = File::create(&path).unwrap();
f.write_all(contents).unwrap();
path.to_str().unwrap().to_string()
}
#[test]
fn evicted_handle_stays_usable_for_its_borrower() {
let dir = tempfile::tempdir().unwrap();
let first = write_file(dir.path(), "first", b"first");
let second = write_file(dir.path(), "second", b"second");
let pool = IndexFilePool::new(1);
let borrowed = pool.borrow(&first, false).unwrap();
// Pushes the single slot over, evicting the entry still in use.
let _other = pool.borrow(&second, false).unwrap();
assert_eq!(pool.pooled_count(), 1);
let mut buf = [0u8; 5];
#[cfg(unix)]
{
use std::os::unix::fs::FileExt;
borrowed.read_exact_at(&mut buf, 0).unwrap();
}
assert_eq!(&buf, b"first");
}
#[test]
fn borrow_reuses_the_pooled_handle() {
let dir = tempfile::tempdir().unwrap();
let path = write_file(dir.path(), "idx", b"x");
let pool = IndexFilePool::new(4);
let a = pool.borrow(&path, false).unwrap();
let b = pool.borrow(&path, false).unwrap();
assert!(Arc::ptr_eq(&a, &b));
assert_eq!(pool.pooled_count(), 1);
// A writable handle is pooled separately from the read-only one.
let w = pool.borrow(&path, true).unwrap();
assert!(!Arc::ptr_eq(&a, &w));
assert_eq!(pool.pooled_count(), 2);
}
#[test]
fn discard_drops_both_handles() {
let dir = tempfile::tempdir().unwrap();
let path = write_file(dir.path(), "idx", b"x");
let pool = IndexFilePool::new(4);
let _r = pool.borrow(&path, false).unwrap();
let _w = pool.borrow(&path, true).unwrap();
assert_eq!(pool.pooled_count(), 2);
pool.discard(&path);
assert_eq!(pool.pooled_count(), 0);
}
#[test]
fn pool_stays_within_capacity() {
let dir = tempfile::tempdir().unwrap();
let pool = IndexFilePool::new(3);
for i in 0..10 {
let path = write_file(dir.path(), &format!("f{i}"), b"x");
let _ = pool.borrow(&path, false).unwrap();
}
assert_eq!(pool.pooled_count(), 3);
}
}
File diff suppressed because it is too large Load Diff

Some files were not shown because too many files have changed in this diff Show More