Compare commits

...
150 Commits
Author SHA1 Message Date
github-actions[bot] 3563738699 4.44 2026-08-22 07:41:01 +00:00
Chris LuandGitHub c1a993bc3b filer: keep the TUS sub-chunks that already landed when a write fails (#10876)
* filer: keep the TUS sub-chunks that already landed when a write fails

A PATCH is split into 4MB sub-chunks, and each one is recorded in the
session as soon as it is stored. The session listing is what HEAD reports
as Upload-Offset and what the final entry is assembled from, so a record
is a promise that the data behind it exists.

When a later sub-chunk failed - a read-only volume, or a client that hung
up mid-body - the error path deleted the needles of every sub-chunk the
same PATCH had written but left their records in place. The resuming
client was then told to continue past bytes the filer had just queued for
deletion, and the upload completed into a gapless manifest pointing at
needles that were gone: HEAD returned the right size, GET died mid-body
once a vacuum reclaimed them.

Recorded sub-chunks now stay, which is what resumption expects: the
client picks up at the offset the session reports, and an upload that is
abandoned frees its chunks with the session.

* filer: drop a TUS chunk's record before freeing its data

filer.CreateEntry can return an error with the entry already inserted -
the parent-directory pass runs after the insert and keeps the entry when
it fails. A failed saveTusChunk therefore does not mean the record is
absent, and deleting the needle outright left the same corruption the
resume path used to cause: a session record pointing at data that is gone.

Remove the record first and only free the needle once it is gone. A
record lost with its data still stored merely leaks, which the vacuum and
fsck paths already account for.

* test: cover a TUS PATCH that is cut off mid-body

Resets the connection after one 4MB sub-chunk has landed, resumes from the
offset the session reports, and vacuums before reading the file back, so
anything the filer deleted behind a kept record shows up as a short read.
2026-08-22 00:30:14 -07:00
Chris LuandGitHub 34bb444f33 test: drive the Lance namespace with Spark (#10864)
* test: drive the Lance namespace with Spark

The counterpart of catalog_spark, which does this for the Iceberg REST
catalog. Spark is the engine most likely to be pointed at a lakehouse,
and it reaches the Lance catalog through the connector's DSV2 catalog -
org.lance.spark.LanceNamespaceSparkCatalog with impl=rest - over the same
routes every other client uses.

    SHOW NAMESPACES -> ['`sparklance-lcephd80`.ml']
    SHOW TABLES -> ['sparklance-lcephd80$ml$embeddings']
    count -> 3
    filtered -> [(2, 'two'), (3, 'three')]
    count after a second commit -> 4

The second insert is there on purpose: a store that cannot order commits
fails on the second one, not the first.

Two things the run settled that were guesses beforehand. CREATE TABLE
works, because the connector declares through the namespace and writes the
data itself rather than pushing Arrow at the server. And SHOW TABLES
returns the namespace's own identifiers - bucket, namespace and name
joined by the delimiter - not bare Spark table names.

Credentials go under the catalog's storage.* prefix, which is handed to
lance as object_store options; a gateway without STS vends none, the same
trap the LanceDB suite documents.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: verify the Lance table bucket was actually created

weed shell prints a command's own failure and still exits 0, so the harness
would go on to blame Spark for a bucket that was never made.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: bound the Docker probe

An unhealthy daemon makes docker version hang, and the probe runs before the
test has a timeout of its own.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: keep the ivy cache under the user's cache directory

It is mounted into a container running as root, so a shared temp path lets
another local user pre-create it and choose what Spark loads.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: assert the vector column's type, not only its name

A column that came back as array<double> or array<string> would still be
called vector and still pass.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: read the dataset off its location for real

The catalog being optional is the property that lets duckdb and pandas read
these tables; it was asserted in a comment and printed, never exercised.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: do not persist credentials in the Spark Lance checkout

The job only uploads a log on failure; nothing in it pushes.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: say the hosts in the README are placeholders

The suite passes dynamically allocated host.docker.internal ports.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-21 23:49:09 -07:00
Chris LuandGitHub 4af6798639 helm: render the mysql secret and env only for the mysql filer store (#10872)
The db credential secret and the filer's WEED_MYSQL_* env were gated on
filer.enabled alone, so a filer on mongodb, redis, postgres or leveldb2 got a
generated mysql secret it never reads - kept forever by resource-policy: keep -
plus a mysql-db-host pointing nowhere.

Gate all of it on WEED_MYSQL_ENABLED, which is how the store is selected, plain
keys and secret-backed ones alike. An enable flag the chart cannot read - a
valueFrom, or one in secretExtraEnvironmentVars - counts as selected, so nothing
is dropped from a filer that is actually on mysql.
2026-08-21 23:43:56 -07:00
df93d01c06 admin: add bucket lifecycle rule editing (#10860)
* admin: add bucket lifecycle rule editing

* address greptile's comments

* more small fixes

* coderabbit's comments

* more comment fixes

* more fixes

* more

* maybe last

* last ?

* 14850

* 14851

* filer: stamp the content MD5 on every SaveInsideFiler write

An entry's ETag falls back to Attributes.Md5, so conditional writers key
IF_ETAG_MATCH off it. SaveInsideFiler carried the looked-up attributes
forward without refreshing the hash, leaving it describing whatever the
previous writer stored: a later conditional write matched the stale hash
and overwrote content that had already changed.

* s3api: give the bucket lifecycle constants and the write route key one definition each

The extended-attribute keys, the XML size cap and the object-write ring key
prefix were each spelled out in two places, so the admin dashboard's copies
could drift from the gateway's. Move them to the packages both sides already
import and alias them where the short local name reads better.

* admin: patch the bucket entry's lifecycle keys instead of rewriting the entry

The save read the bucket entry, edited its extended map and wrote the whole
entry back, guarded by IF_UNMODIFIED_SINCE. Nothing that writes a bucket
entry advances its mtime - not the S3 gateway's patchBucketEntry, not
SetBucketOwner, not SetBucketQuota - so the guard never fired and the stale
snapshot reverted whatever else had changed since the lookup.

Send the PATCH_EXTENDED mutation the S3 gateway already uses for these keys:
the filer re-reads and merges under the bucket path lock, so only the two
lifecycle keys move. That removes the reason for the mtime snapshot, the
verification retry loop and the compensating restore of the cleared day-TTL
rules, which the migration now logs instead.

* s3api: run the delete-lifecycle day-TTL migration through the shared helper

DeleteBucketLifecycleHandler kept its own copy of the read-strip-write
sequence the put handler now shares, including a missing return that let a
ToText failure persist a truncated filer.conf and write a second response.
It also wrote the whole file back unconditionally, reverting any concurrent
edit; the shared helper writes conditionally.

* admin: answer 404 when a lifecycle request names a bucket that does not exist

Every SetBucketLifecycle failure came back as 500, including the lookup miss
for an unknown bucket, so a client or monitor read a caller error as a server
fault and retried it.

* s3api: emit lifecycle XML a client would recognize

Two changes to what MarshalCanonical writes, both visible through
GetBucketLifecycleConfiguration, which replays the stored bytes verbatim:
stamp the S3 namespace on the root, and put a size range under <And>. A
<Filter> carries one predicate, so two size bounds side by side is a shape
AWS does not document. Parsing still accepts either.

* admin: fix the lifecycle editor's handling of stored status, deletes and empty saves

Four things the editor got wrong:

A stored <Status> the S3 API never validated, say 'enabled', left both radio
buttons unchecked, so reading the form threw on a null querySelector result
and Save did nothing. Collapse anything but an exact 'Enabled' to 'Disabled',
which is what the engine already does with it.

Deleting a rule re-rendered an open edit form from the snapshot taken when
editing began, discarding what had been typed; every other transition folds
the form in first.

The Transition warning only matched a bare <Transition>, missing the form
with attributes, self-closed or namespace-prefixed.

Saving an emptied rule list clears the configuration through a path with no
prompt, next to a Delete-all-rules button that asks.

Also collapses the three divergent copies of formatBytes on this page to one.

* filer: stop the day-TTL migration from deleting an operator's path rule

The migration removed every rule under the bucket's path that carried a day
TTL in the bucket's collection. The add path it is retiring used
AddLocationConf, which merged its TTL onto whatever already sat at the
prefix, so a rule can hold operator settings the lifecycle path never wrote -
a disk type, WORM retention, a read-only flag, a placement pin. Deleting the
whole rule to retire its TTL took those with it, leaving objects under that
prefix on defaults nobody asked for.

Delete only rules shaped like ones the add path created from scratch;
anything else keeps its settings and loses just the TTL.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-08-21 23:42:26 -07:00
Chris LuandGitHub 301d83cc7a test: wait for the master to register the volume servers before failover tests run (#10871)
The failover harness treated an open volume server port as readiness, but the
master only learns of a volume server from its heartbeat. A lone master refuses
heartbeats until its bootstrap check elects it, and the servers back off and
retry, so registration lands seconds after the ports answer. Tests that started
writing in that window assigned against an empty topology, which fails with
"no free volumes left" and reaches the mount as ENOSPC.
2026-08-21 23:05:14 -07:00
Chris LuandGitHub c0a9b110dd volume: stop reporting read-only volumes that are no longer here (#10867)
* volume: clear per-collection metrics when a collection leaves a server

The read-only and disk size gauges are only ever set for collections the
heartbeat still finds here, and nothing zeroes the rest. volume.balance marks a
volume read-only to move it, so the last heartbeat that saw it counts it
read-only - and if it was the collection's last volume on that server, that
count stands until the process restarts. The dashboard then shows read-only
volumes that volume.list -readonly cannot find anywhere.

Remember what each heartbeat set, and drop what is gone on the next one.

* volume: stop the read-only volume count from wrapping at 256

The per-collection counters were uint8, so a server holding 256 read-only
volumes of one collection reported zero of them.

* volume: read the read-only flags once when counting them

The heartbeat asked IsReadOnly for the verdict and then read noWriteOrDelete
and noWriteCanDelete straight off the volume, unlocked, so the reasons could
disagree with the verdict they were explaining. Take them together, under one
lock. The location is now nil-checked rather than skipped by short-circuit
evaluation, so a volume that has not joined a disk location yet stays safe.

* volume: let only a surviving volume keep its collection reported

A volume being deleted for expiry still made an entry in the read-only counts,
which is what the cleanup reads as "this collection is still here". The
collection's last volume could go and its series would stand for one more
heartbeat. Count the survivors only.

* volume: size a collection from the volumes it still has

The size totals are rebuilt from scratch every heartbeat, so subtracting a
volume that is about to be deleted took the surviving volumes' sizes down with
it: a collection keeping a small volume and losing a larger one reported the
difference, or lost its entry and kept the previous heartbeat's number.

* volume: cover the deleted bytes total in the surviving volume test

Deleted bytes are totalled the same way as sizes and were going unchecked, so
the test now leaves deleted needles on both volumes and pins that gauge too.
2026-08-21 22:33:01 -07:00
Chris LuandGitHub 96304b6870 S3: source config credentials from the environment, and let the chart point at an existing secret (#10868)
* s3: resolve ${VAR} in static config credentials from the environment

A deployment that keeps its S3 keys in a secret store had no way to hand
them to the gateway: -config takes a file, so the keys had to be written
into that file. Let a key in the static config name an environment
variable instead, and drop any credential whose reference stays unset so
the placeholder never becomes a usable key.

* helm: source the generated s3 identities from an existing secret

The only way to reuse credentials that already live in a Secret was to
hand-author the whole seaweedfs_s3_config JSON, since the literal keys in
values.yaml end up in git and a lookup-based keyRef renders empty under
helm template and Argo CD. Let s3.credentials.admin/read name a Secret and
its keys instead: the generated config references them as ${VAR} and the
gateway resolves them from the environment, so nothing is read from the
cluster at render time.

* s3: treat an empty environment value as an unresolved credential reference

A secret store can hand over a key that exists but is blank. Resolving it
would leave an access key whose signing secret is empty, so count it as
unresolved and drop the credential.

* helm: render the s3 secret when only the all-in-one auth flag is set

The all-in-one deployment mounts the s3 secret whenever any of the three
enableAuth flags is set, but the secret itself only rendered for the s3 and
filer flags, so allInOne.s3.enableAuth on its own left the pod waiting on a
secret nothing creates.

* helm ci: check the credential wiring on every workload that mounts it

The render check only looked at the standalone s3 deployment and only at
one of the four variables, so a helper that bound a variable to the wrong
secret key would still pass.

* helm: create the all-in-one s3 secret for every flag that mounts it

The all-in-one pod mounts the secret on any of the three enableAuth flags,
so keying its creation off allInOne.s3.enableAuth alone still left
filer.s3.enableAuth without filer.s3.enabled pointing at a secret nothing
creates. Mirror the deployment's own condition instead, and check each
flag renders both the mount and the secret.

* s3: reject a malformed credential reference instead of keying on it

A typo such as ${MY-VAR} matches no substitution, so it survived expansion
and the placeholder itself became the access key the gateway accepted.
Require every ${ in a static credential to open a well-formed reference.
2026-08-21 22:32:47 -07:00
Chris LuandGitHub 480795d40d release: cut the whole release from the version bump workflow (#10870)
* release: cut the whole release from the version bump workflow

The bump workflow stopped after pushing the version commit, and the rest was
manual: create the release, then run "Prepare release" in the csi-driver and
the operator. It now pushes the tag itself, which is what starts the binary,
container and helm workflows, creates the release with generated notes, and
dispatches the other two repositories, waiting for both.

Pushing the tag and reaching the other repositories both need RELEASE_PAT;
GITHUB_TOKEN raises no events that start workflows.

* release: tighten the release workflow after review

Check out master explicitly: a dispatch can select any branch, and the tag,
the commit and the release would then come off that branch while the
downstream job dispatches master.

Scope contents:write to the job that pushes; the downstream job talks to the
other repositories with RELEASE_PAT and needs nothing here.

Wait for the module proxy to serve the release commit as the tip before
dispatching, instead of priming it and hoping. The dispatched workflows pin
seaweedfs with `go get -u ...@latest`, so a stale tip means they release
against a pre-release commit, silently.

Identify the dispatched run by diffing the run list against the snapshot
taken before dispatching, rather than assuming the newest run is ours.

* release: wait on the downstream release, not on the run that makes it

A dispatched run cannot be told apart from a concurrent one: the API does not
report the inputs a run was dispatched with, so watching "the run that appeared
after mine" can watch someone else's and report their result as ours.

Wait for a release to appear in the downstream repository instead. That is the
thing being waited for, and it holds however many runs are in flight.
2026-08-21 22:19:58 -07:00
Chris LuandGitHub 5e7ab43ddd test: read Lance tables from DuckDB (#10866)
* test: read Lance tables from DuckDB

The LanceDB and Spark suites go through the catalog. DuckDB does not: its
lance extension reaches the data over S3 with no namespace involved, which
exercises the other half of the design - a table bucket's layout is a
valid Lance dataset directory, so a table stays readable when the catalog
is not in the path.

    scan_rows=128
    scan_columns=id,title,vector
    filtered_rows=5
    nearest=1,0,2

It also pins the one place the layout costs us. DuckDB's replacement scan
recognises a dataset by a .lance path suffix, and tables created through
this catalog deliberately have none: the catalog entry is the dataset
directory, a table name may not contain a dot, and a suffix would leak
into ARNs and policies. So __lance_scan is the way in, and the bare
SELECT ... FROM 's3://...' form does not see these tables.

The test asserts both halves - a suffixed path is read, a suffix-less one
is not - so if the extension ever recognises a bare directory, it fails
and says to update the documentation rather than leaving it wrong.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: require the catalog error from the suffix-less read

Any failure satisfied the old check - a missing extension, bad credentials,
an unreachable endpoint - so the assertion could pass without the
replacement scan ever classifying the path.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: verify the Lance table bucket was actually created

weed shell prints a command's own failure and still exits 0, so the harness
would go on to blame DuckDB for a bucket that was never made.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: bound the Docker probe

An unhealthy daemon makes docker version hang, and the probe runs before the
test has a timeout of its own.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: order the aggregates the assertions read

string_agg over an unordered relation may return the names, and the vector
search's ids, in any order, so the expectations could fail on a run where
nothing changed.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: do not persist credentials in the DuckDB Lance checkout

The job only uploads a log on failure; nothing in it pushes.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-21 15:22:38 -07:00
Chris LuandGitHub 35d53a20f6 master: let the leader admit a master that starts with no raft state (#10865)
* master: answer with the leader raft already knows

Topo.Leader() backs off for up to 20 seconds waiting for an election.
Callers that a health probe or a client is blocked on cannot afford that:
/cluster/status, /cluster/healthz and /readyz all sit past the probe
timeout of both the helm chart and the operator, so a master that is
still joining looks dead rather than joining, and the kubelet restarts
it. informNewLeader and SendHeartbeat hold the client on a master that
cannot serve it, exactly when it should move on to find the one that can.

Answer these from MaybeLeader instead, which reports what raft knows
right now. MaybeLeader takes over the "am I the leader myself" fallback
that Leader() used to apply on top of it, so one non-blocking call is
still correct; Leader() keeps the backoff for callers that must wait.

* master: let the leader admit a master that starts with no raft state

Neither raft implementation lets a server outside the configuration
campaign: goraft's promotable() requires a non-empty log, and hashicorp
rejects vote requests from a candidate that is not in its configuration.
A master that comes up with fresh state therefore cannot elect itself in
— the leader has to pull it in. Nothing did.

The peer list is static, rendered from the replica count, so scaling it
up leaves the sitting leader running the old list with no idea the new
masters exist. Under goraft they wait forever. Under hashicorp they are
worse off: each bootstraps a cluster of its own from the new list, and
two of them form a quorum next to the live leader, with their own
TopologyId. That is the split brain SetTopologyId kills a master over.

Admit the peer where it registers instead. Only the leader gets past the
IsLeader check in KeepConnected, and a joining master's client lands
there, so that is the moment it joins. The broadcast OnPeerUpdate rides
on is not enough on its own: it only reaches masters already connected,
which is why a leader that came up first missed both newcomers.

RaftAddServer grew a goraft branch on the way, so cluster.raft.add stops
silently doing nothing on the default raft, and RaftRemoveServer with it.
Bootstrapping is now one call for both implementations, made only after
the peers confirm nobody has a leader, and retried until this master is
in rather than checked once and dropped.

* master: do not evict a peer that is still in -peers

The hashicorp leader drops a master from the raft configuration as soon
as it stops answering pings. A master that is merely restarting answers
nothing, so an ordinary bounce shrinks the quorum behind the operator's
back — and then races its own return: the master comes back, registers,
gets re-admitted, and the eviction lands after it.

A randomized start/stop walk lands on it. Two of three masters running,
the leader evicts the one that just went down, the restart re-adds it,
the removal commits late and takes the leader's own leadership with it.
What is left is a two-server configuration whose other half is down, and
a running master that nobody will ask for a vote — no quorum, no way
back until the third master returns.

-peers is what declares membership. updatePeers already reconciles the
configuration against it on every leadership change, and an operator who
really means to drop a master can say so with cluster.raft.remove, so
keep the eviction for masters that are no longer listed at all.

* test: bounce masters at random and hold the election to it

Twelve rounds of stopping or starting a random master, on both raft
implementations, checking the two things an election must never get
wrong: two masters claiming leadership at once, and a quorum that comes
back without agreeing on one. The cluster's identity has to survive the
whole walk, since a master that re-mints a TopologyId is the split brain
SetTopologyId kills its peers over. The seed is random and logged, so a
failure names the walk that reproduces it.

Below a quorum the walk moves straight on. A master that has lost its
quorum cannot commit anything, and goraft only checks whether it still
has one on an election-timeout ticker, after its peers have been quiet
for a full timeout — measured taking over 30 seconds to step down. That
direction belongs to TestTwoMastersDownAndRestart, which was giving it
ten seconds and would have started failing on a slower machine; it now
waits on that behaviour explicitly rather than sleeping twice and hoping.

WaitForTopologyId returns the id it waited for. Reading it separately
raced the leader applying the raft entry that carries it, which shows up
as an empty id right after an election rather than as a wrong one.
2026-08-21 15:22:22 -07:00
Chris LuandGitHub 0c95137528 filer: stop aggregated metadata subscribers from spinning on a peer watermark hold (#10863)
* fix(filer): stop logging a held aggregated read as an error

An aggregated subscriber may not read past the peers' low-watermark, and
it stops at the first entry beyond it by returning a sentinel from the
read callback. LoopProcessLogData logs every callback error, so on a
cluster that keeps writing - where there is almost always an entry newer
than the watermark - every read wrote an ERROR line naming the entry it
stopped at, thousands per minute per filer.

Mark the stop as control flow: an error wrapping StopReadingError is
handed back to the caller unlogged, and the held-read sentinel wraps it.

* fix(filer): release an aggregated watermark hold on peer progress

A held read waited on the aggregated buffer's data channel, which the
next write signalled - but a write cannot release a hold, only a peer
reporting further progress can. On a cluster that keeps writing the loop
therefore re-ran a whole pass per arriving event, log file listing and
all, and held again on the same entry every time.

Signal held readers from the meta aggregator instead, whenever a
low-watermark rises: a peer reporting, or one dropped past its removal
grace. The retry interval stays as the backstop for what no watermark
covers. Count the holds so a parked subscriber stays visible.

* fix(filer): floor how often an aggregated watermark hold releases

Peers advance their delivery watermark on every event they stream, so
releasing a hold on every advance is the same pass-per-event storm as
releasing on every write, just without the log lines - and each pass
lists a day of log files.

Floor the release at 20ms. Advances inside the floor collapse into one
release, which then delivers everything they covered.

* fix(filer): pace a peer's delivery claim by what its subscribers hold at

A filer's local metadata stream carries an idle heartbeat to its peer
aggregators, and each peer turns it into that filer's delivery
low-watermark. Aggregated subscribers hold at the minimum across peers,
so a filer quiet enough to fall back on the heartbeat parked every
subscriber in the cluster up to a keepalive interval - 5 seconds -
behind live writes. With nine filers, most of them quiet at any moment,
the minimum sat there permanently.

Pace that heartbeat at 200ms once the filer has peers. It stays a
keepalive, at the keepalive interval, for a filer with none.

* fix(filer): wake each aggregated hold on its own watermark

A persisted-log read is held by what the peers have flushed, an
in-memory read by what they have delivered, but both parked on one
channel closed whenever either minimum rose. Peers advance their
delivery watermark on every event they stream, so a flush-held reader
woke at the coalescing floor to re-list a day of log files and park
again on the same entry - the storm this set out to fix, in the one
place asymmetric peer progress still reached.

Signal the two separately and park each read on the one that bounds it.
2026-08-21 15:22:05 -07:00
Chris LuandGitHub 0dfaa103d0 test: take a table through its whole life, for Iceberg and Lance (#10862)
* lance worker: share the integration tests' scaffolding

The recorder that keeps what a handler sent, the config builder and the
storage-option fallback all lived inside compaction.rs, so a second test
binary would have had to copy them. They move to tests/common.

The fallback now reads AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY and
AWS_ENDPOINT_URL from the environment, defaulting to what it used before.
A harness can then point these tests at a gateway that checks what it is
given rather than one that accepts anything.

* lance worker: maintain one named table, for a harness to drive

Compacts and cleans up whatever WEED_LANCE_TABLE names, through the
handlers' own detect-then-execute path: a proposal the worker would not
have made is not one worth running.

The existing tests seed the tables they check. This one deliberately does
not, so a harness that has already written a table and knows what is in it
can have the real handlers maintain it and then read it back.

* test: take a table through its whole life, for Iceberg and Lance

Created in the catalog, filled by a real client, maintained by the worker,
read again, dropped. The step nothing was checking is the read after
maintenance: compaction once rewrote every dictionary-encoded column onto
a single value and shipped, because the maintenance tests were thorough
about sequence numbers, manifest entries and metadata versions and none of
them opened the parquet file the worker had just written.

So the assertion is a tally - row count, the cardinality of each
dictionary-encoded column, and an md5 over whole rows - taken before
maintenance and again after, required to be equal. The cardinalities name
the failure that happened; the digest catches a rewrite that keeps every
column's cardinality and hands the values to the wrong rows. A compaction
that merged nothing fails rather than passes, or the read afterwards is
checking a file the worker never wrote.

The Iceberg half runs two clients. DuckDB is the one the corruption was
reported against and the only one here that writes the deprecated
PLAIN_DICTIONARY encoding, which parquet-go normalizes away on write, so a
Go writer cannot produce it. PyIceberg writes the modern spelling. Pinning
parquet-go back to v0.30.1 fails the DuckDB half and passes the PyIceberg
one, which is why both are here.

Lance maintenance lives in the Rust worker, so it runs there where cargo
is installed and through the two lance calls those handlers wrap where it
is not. WEED_LANCE_MAINTENANCE picks one instead of letting the test guess.

* ci: run the table lifecycle tests

CI maintains the Lance table through the lance library rather than the
worker: a cold build of the lance crate costs more than the glue it would
be checking, and the worker's own tests cover its handlers.

The suite drives the Iceberg maintenance worker, so a change to it now
triggers this workflow too.

* test: let the lifecycle harness fail instead of skipping

Setup failures all exited zero, so a cluster that would not come up, or a
port allocation that lost, reported a green run for code nothing had
executed. That is the failure mode this whole directory exists to close,
and it was in the harness itself.

Only a checkout without a weed binary skips now, and it runs the tests so
each one says so rather than the package quietly passing. Everything else
fails.

The filer existence probe gets a deadline while I am here: it ran without
one, so an unresponsive filer would hang the suite past every timeout the
clients have.

* test: make the lifecycle checks check what they claim to

Three of them could pass without having looked.

The DuckDB skip matched "syntax error", "not implemented" and "Failed to
load" anywhere in the output, in any phase. A parse error in the SQL this
test generates, or a refusal from our own catalog, would have taken the
only coverage of the PLAIN_DICTIONARY encoding out of CI and left it
green. It now matches the extension failing to install, and only in the
phase that installs it. Everything past LOAD is ours and fails.

The digests covered id, category and value. Compaction rewrites the whole
row, so a defect confined to ts, or to a Lance vector, changed nothing
either side of maintenance. Every persisted column goes in now, ts as
microseconds so no timezone sits between the two runs.

The Lance drop check caught every exception as proof the dataset was
gone. pylance turns credential and transport failures into the same
ValueError, so it only accepts the message that means not found.

* docs: say up front which maintenance path the Lance half takes

The opening summary said the worker maintains both tables. It maintains
the Iceberg one always and the Lance one only where cargo is installed,
which is not what CI does.
2026-08-21 15:16:11 -07:00
Chris LuandGitHub 3bd218e030 volume: cut idle memory at high volume counts (#10861)
* volume: start a volume's batch write worker on first use

Mounting a volume started a goroutine parked on a 128-slot channel, plus
the 128-entry batch slice it had already allocated. That is around 6.7KB
per volume the server pays whether or not the volume ever takes a write:
7231 bytes per mounted volume, of which 4101 is goroutine stack.

Only a write that asks for fsync ever reaches the worker, and a
remote-tiered or read-only volume never can. Create the channel and its
goroutine on the first such request instead, and let a write arriving
after Destroy fall back to the inline path rather than queue onto a
worker that has gone.

Measured over 20000 mounted volumes: 7231 -> 1269 bytes each.

* volume: update the heartbeat report state in place

Every heartbeat built a second map of what it was about to tell the
master, holding a freshly allocated short information message per volume,
then swapped it in over the old one -- and computed departures through a
third map of the live volume ids. A server holding 2M volumes rebuilt all
three every VolumePulsePeriod for a report that usually says nothing.

Number the heartbeats instead and mark the entry already held with the
pass that found the copy, so a quiet volume costs a map lookup and no
allocation. Departures are the entries a pass did not mark; the live-id
map is now built only when there are some, sized to them.

Measured over 10000 mounted volumes: 436 -> 196 bytes allocated per
volume per heartbeat.

* volume: fill one volume information message per heartbeat, not per volume

The heartbeat built a message for every volume held so it could hash it,
then dropped all but the few it had something to say about. At 2M volumes
that is 2M messages allocated every VolumePulsePeriod to send almost none
of them.

Fill a message the caller supplies instead, and replace it only when the
heartbeat keeps it, so a server with nothing to report fills the same one
all the way through.

Measured over 10000 mounted volumes: 196 -> 4 bytes allocated per volume
per heartbeat, and a heartbeat runs a third faster.

* volume: drop the per-volume trace from the heartbeat's status read

glog.V(4).Infof evaluates its arguments whether or not the verbosity is
on, so every volume boxed its id into a fresh interface slice on every
heartbeat: 759 of the 773 allocations a 1000-volume heartbeat made, for a
line that at this scale would print millions of unreadable rows.

Measured over 1000 mounted volumes: 4776 -> 1792 bytes and 759 -> 14
allocations per heartbeat, which no longer grows with the volume count.

* seaweed-volume: mirror the in-place heartbeat report state

Same change as the Go volume server: number the heartbeats and mark the
entry already held with the pass that found the copy, instead of building
a second map of hashes and swapping it in.

The volume snapshot must leave the reporting state as it found it, so it
keeps asking through changed() while a real heartbeat marks through
record().

* volume: refuse writes to a closed volume instead of dereferencing nil

Close and Destroy leave the needle map and data backend nil, but a caller
that already holds the volume can still reach the write path, where both
are used unguarded: a write racing a volume deletion took the server down.
syncDelete has always checked; syncWrite and the batch worker had not.

Reachable before this series and now also from the inline fallback a
durable write takes when the worker has gone.

* seaweed-volume: guard the report state with one mutex, as Go does

The full-list flag and the generation that answers it have to move
together. Split across separate atomics they cannot: a request landing
between begin's two reads returns full == false with the generation it
just raised, and one landing between commit's read and its clear is
marked answered by a heartbeat that carried no list. Either way the
resend is dropped.

Neither is reachable today -- every caller reaches this through the
store's RwLock, the flag setters under a read lock and the heartbeat
build under a write lock, so they cannot interleave. The type should not
depend on that being true two files away, and Go holds a single mutex
over exactly these fields.

* test: build the servers under test to match the harness's offset size

The mixed Go/Rust suites run both servers against one dataset, so both
have to agree on the offset width. They did not: the harness built Go
with no tags, 4-byte offsets, while the Rust crate defaults to its 5bytes
feature, and the Rust server then refused the .vif the Go server had just
written -- "bytes_offset mismatch: found 4, expected 5".

Build each side to match the offset size the test binary itself was
compiled with, so a plain `go test` and one with -tags 5BytesOffset both
get a matched pair.
2026-08-21 13:04:56 -07:00
930603eb74 S3: optionally serve remote-mounted objects from remote when the local read fails (#10837)
* feat(s3): serve from remote on local read failure

When a locally-cached chunk of a remote-mounted object becomes unreadable
(volume server down/restarting, or an evicted needle 404ing under
retry-backoff), fall back to serving the object from its mounted remote
instead of erroring. A bounded pre-flight probe makes a stuck volume trip
the timeout rather than stalling the request.

Gated by -localReadFallbackToRemote (default off) with
-localReadFallbackTimeout (2s default), so existing deployments are
unaffected until they opt in.

* fix(s3): register local-read-fallback flags for mini/server/filer

The mini, server and filer launchers build S3Options directly and only
populate the flag pointers they register. Without registering the two new
flags there, startS3Server dereferenced nil pointers and crashed at boot,
failing every integration suite that runs `weed mini`.

* fix(s3): treat a zero-byte probe read as unreadable

A read that returns no byte -- whether it reports io.EOF or no error at all
-- means the offset is not locally readable, so the probe must fall back to
the remote rather than proceeding to stream a truncated response. Only a
returned byte (including the object's final byte with a trailing io.EOF)
counts as readable.

* s3: finish a mid-stream local read failure from the remote mount

The pre-flight probe only proves the byte at the requested offset readable.
A multi-chunk object can still lose a later chunk after the 200/206 and its
Content-Length are committed, which truncated the body with no fallback.
Resume from the mounted remote at the byte the local copy stopped at, so the
response still carries the declared length. A short local read that surfaces
as a clean EOF is treated the same way instead of silently truncating.

* s3: fall back to the remote mount without a CLI switch

Serving a remote-mounted object from its authoritative remote is what the
read should have done all along -- the alternative is a 500 on an object the
cluster can still reach -- so make it the behavior instead of two new flags,
with the probe bounded by a constant.

* s3: trim the comments on the fallback path

* filer: report only the contiguous prefix when a parallel chunk read fails

The parallel branch of doReadAt fans the chunk reads straight into their own
windows of the output buffer, then sums every task's bytesRead. A middle chunk
failing while a later one succeeds therefore returned a length covering a hole
the reader never filled, handing the caller zeros in the middle of otherwise
valid data.

* s3: only splice the remote onto a local prefix while it is the cached generation

Eligibility establishes a size match, not byte identity: a remote key
overwritten with same-size content between the cache fill and the fallback
would have finished the response with bytes from a second generation, under
the first one's ETag. Stat the remote before resuming and keep the local
error when it no longer matches -- a truncated body is a visible failure,
a spliced one is not.

---------

Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-21 11:47:53 -07:00
Chris LuandGitHub 6faa9d20e8 iceberg: stop compaction from corrupting dictionary-encoded columns (#10857)
* deps: upgrade parquet-go to v0.32.0

Iceberg compaction writes the merged file with the schema of its first
input, encodings included. parquet-go before v0.31.0 took the deprecated
PLAIN_DICTIONARY encoding that DuckDB writes at face value and encoded
those pages as plain int32 indices, but the spec gives PLAIN_DICTIONARY
the same bit-width-prefixed RLE layout as RLE_DICTIONARY. Every
dictionary-encoded column in a compacted file then decoded onto a single
dictionary entry, and anything past one page failed to decode at all.

* iceberg: cover compaction of dictionary-encoded input

The fixture is a DuckDB-written file, so it carries the PLAIN_DICTIONARY
encoding a Go writer will not produce.

* iceberg: tally whole rows in the dictionary merge test

Counting each column on its own passes a merge that remaps names while
leaving their cardinality intact.
2026-08-21 10:12:01 -07:00
Chris LuandGitHub 813c439af6 admin: regenerate the gzipped static mirror (#10859)
The toast and modal changes edited static/js/ without rerunning
gen_static_gz.go, so the embedded assets served to browsers still carry
the old scripts and TestStaticGzMirror fails on master.
2026-08-21 10:10:10 -07:00
Chris LuandGitHub bd34565e56 admin: keep the copy confirmation in front of the access key modal (#10856)
* admin: raise nested modals above the ones already open

Bootstrap gives every modal and every backdrop the same z-index, so a
modal opened while another is showing paints behind it and its buttons
cannot be clicked. Viewing an access key secret and then copying a field
left the confirmation stuck behind the details modal with no way to
dismiss it.

Give each nested modal, and the backdrop Bootstrap creates for it, a
z-index above what is already on screen, and put back the scroll lock
that Bootstrap drops as soon as any one of them closes.

* admin: confirm clipboard copies with a toast

The access key details modal offers three copy buttons, and each one
raised a modal that had to be dismissed before the next copy. Confirm
with a toast instead, so the credentials stay in view and nothing has to
be clicked away.
2026-08-21 09:20:27 -07:00
github-actions[bot] 6c7f184381 4.43 2026-08-21 07:08:16 +00:00
Chris LuandGitHub 8a532cc0cf mini: state the format of a -tableBucket, do not infer it (#10851)
* mini: state the format of a -tableBucket, do not infer it

A table bucket holds one format and that format decides which catalog can
serve it, but the flag only took names. The format came from
miniTableBucketFormat(): Iceberg whenever its port was up, Lance only when
it was not. So -tableBucket=vectors on a default mini quietly made an
ICEBERG bucket that the Lance namespace then refused every table in, and
the only way to get a Lance one was -s3.port.iceberg=0, which buys it by
deleting the other catalog. One flag, two meanings, decided by an unrelated
port.

Each entry is now name[:FORMAT], unsuffixed meaning ICEBERG as before:

    weed mini -tableBucket=warehouse,vectors:LANCE

Both catalogs stay up and both buckets are reachable. A name whose format
has no endpoint here is skipped with a warning rather than created out of
reach, and the Iceberg-only S3_TABLE_BUCKET default-routing hint gets the
Iceberg names alone, without their suffixes.

* mini: do not reuse a table bucket that holds another format

CreateTableBucket answers BucketAlreadyExists on the name alone, so
-tableBucket=vectors against a bucket created as LANCE logged "already
exists" and moved on, and the Iceberg default-warehouse hint then pointed
at it. Every table create against that catalog fails with "table bucket
vectors holds LANCE tables", far from the flag that chose it.

ensureMiniTableBuckets now reads the format of a bucket it did not create,
warns when it is not the one asked for, and returns only the buckets that
hold what was requested. S3_TABLE_BUCKET is seeded from that list, so an
unprefixed Iceberg request falls back to its own default rather than
committing into a Lance bucket. A bucket predating declared formats reports
an empty one and still accepts either.

* mini: normalize S3_TABLE_BUCKET whichever way the spec arrived

The rewrite that keeps Lance names out of the Iceberg default warehouse only
ran when the flag supplied the spec. Set the variable directly, as the docker
quickstart does, and it reached the catalog untouched: S3_TABLE_BUCKET=
vectors:LANCE,warehouse made the unprefixed default the literal string
"vectors:LANCE", a bucket no lookup finds, while warehouse sat behind it.

The variable is both mini's input and the catalog's routing hint, so it is
now always rewritten from the buckets that came back holding Iceberg tables,
and unset when there are none rather than left pointing somewhere stale.

* mini: reuse a table bucket only when its format reads back

An ordinary S3 bucket wearing the name answers CreateTableBucket with the
same BucketAlreadyExists as a table bucket does, and the format lookup that
follows returned "" for a failed read exactly as it does for a bucket
predating declared formats. So -bucket=data -tableBucket=data reported
nothing and published data as the Iceberg default warehouse, where every
unprefixed request 404s on a bucket that is not a catalog.

The lookup now returns its error, and only a bucket that reads back as the
format asked for is reused. Anything else is left alone with a warning
naming why, rather than routed to and discovered later.
2026-08-20 23:46:32 -07:00
Chris LuandGitHub 83753ccdad test: drive the Lance namespace with LanceDB (#10850)
* test: drive the Lance namespace with LanceDB

The Iceberg catalog is checked against Spark, Trino, ClickHouse, Doris,
Dremio and RisingWave. The Lance one had only its own reference client,
which is the same thing as checking it against ourselves.

LanceDB connects with connect_namespace("rest", ...), which speaks the
routes this catalog implements, so the suite exercises the protocol rather
than our idea of it: list the catalog, open a table through it, read the
schema, run a vector search and a filtered scan, create a table, and read
the same dataset straight off its URI with no catalog at all.

    table_names -> ['lancedb-p0guidmm$ml$embeddings']
    open_table  -> 64 rows
    search      -> [1, 0, 2]
    create_table -> 4 rows, listed by the catalog
    direct read without the catalog -> 64 rows

Seeding is pylance, because the namespace records where a table lives and
does not carry its data. That split is the design rather than a limit of
the test.

One interop note the test encodes: a gateway without STS vends
storage_options carrying an endpoint and a region but no credentials, and
LanceDB uses what the namespace vends on some paths. The container gets
credentials in its environment as well, which is what a deployment without
STS would do.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: pin the LanceDB client, and index before searching

Three from review.

The client's dependencies were unpinned, so an unrelated upstream release
could change what an old commit reproduces. Pinned to the versions this
suite was verified against; the client is as much the thing under test as
the server.

The search was called ANN and was not: without an index LanceDB scans.
The test now builds an IVF_PQ index over 1024 rows first, which is worth
more than the wording fix - an index writes into a directory of the table
that the S3 door has to admit, and that guard has refused a Lance
directory before. It builds, covers all 1024 rows, and searches.

The assertion moved with it. Demanding the exact nearest neighbour was
right for a brute-force scan and wrong for a quantized index, which
answered 0 as readily as 1; both are correct, so the check is now the
neighbourhood.

And the pushdown check accepted any failure. It now requires the refusal
to be the catalog's Unsupported and requires that nothing was left behind,
or, when the client falls back, that the table is complete.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-20 20:11:48 -07:00
Chris LuandGitHub 6f3b5a4f4d metrics: graph the plugin workers (#10849)
Nothing displayed the worker metrics, and the panels that look like they
did are about something else: Workers Connected, Worker Slots and Worker
Events in the Admin / Maintenance row read SeaweedFS_admin_*, which the
older maintenance queue feeds. A cluster running plugin workers - Go or
Rust - reads zero there while they are connected and busy.

A Plugin Workers row graphs what the workers themselves publish: how many
are connected, jobs and their failures, detection and proposal rates, job
duration, slot usage, stream events, and what the Lance jobs reclaimed.
The panel worth having is Objects Seen vs Skipped, since a sweep with
nothing to do and a sweep that could read nothing report the same number
of proposals.

Also a commented scrape target in the sample Prometheus config. It is 9328
rather than 9327: the sample compose already gives 9327 to the S3 gateway,
so the port the worker's own usage text suggests collides with it on a
single host.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-20 11:49:30 -07:00
Chris LuandGitHub f56a7a1557 seaweed-worker: serve health, readiness and metrics (#10848)
* seaweed-worker: serve health, readiness and metrics

A Rust worker had no surface of its own. If it wedged, the only signals
were its stdout and whatever admin could infer from a stream that had gone
quiet; nothing could be scraped and nothing could be alerted on.

--metrics-port serves /health, /ready and /metrics, the same three the Go
worker serves under -metricsPort, so one scrape config covers workers in
either language. Off by default, loopback unless --metrics-ip says
otherwise, since the endpoint is unauthenticated. Names follow the Go
convention, SeaweedFS_worker_*.

The counters live in core and are raised where the stream already knows
what happened - connect, close, detection, execution, preview - so a
worker for another format gets them without writing any of this. Slots are
published from the heartbeat that already computes them, so a scrape and
the admin UI cannot disagree.

The pair worth having is objects_seen_total and objects_skipped_total. A
sweep that proposed nothing because there was nothing to do and a sweep
that proposed nothing because it could not read anything are the same
number of proposals; they are not the same event, and until now only a log
line told them apart.

The Lance jobs add what they reclaimed - fragments, rows brought under an
index, versions, bytes - on the same registry, so one endpoint serves both.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* seaweed-worker: fix the metrics address, the count, and a dead field

Three from review.

--metrics-ip ::1 failed at startup: the address was built by joining host
and port with a colon, and "::1:9327" is not an address. It is parsed as a
host and combined with SocketAddr::new now, so an IPv6 literal works, with
or without the brackets an operator will reasonably type after seeing one
in a URL.

proposals_total counted before the send rather than after, so a stream
that closed mid-sweep left the counter claiming proposals admin never
received.

And MeteredSender carried a Metrics clone and a job type it never read,
kept alive by two statements that existed only to silence the warning
about them. Everything is recorded by the caller, so both are gone.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-20 10:11:58 -07:00
Chris LuandGitHub fc97f8ea8f mount: index directory state by path (#10827)
Every directory-state lookup went through path2inode, the map that holds one
full path per inode in the table, and then through dirStates. Directories now
carry their own path and are indexed by it directly.

There are orders of magnitude fewer directories than files, so this map stays
small whatever the mount holds, and it is what a file needs before it can stop
carrying a full path of its own: a child's path is its parent's plus its name.
No behavior change - the two indexes are asserted to agree.
2026-08-20 10:09:42 -07:00
f7c4636d22 topology: refresh oversized mark on every heartbeat (#10829)
* topology: refresh oversized mark on every heartbeat

The oversized flag on a volume location was only set when the volume was
registered (RegisterVolume). A volume that later grew past the size limit
kept its stale "not oversized" mark, so the heartbeat path
(ensureCorrectWritables) kept re-adding it to the writable list while
RecordAssign removed it on every assign - a writable/unwritable flip loop
that let writes continue past the limit and made vacuum race in-flight
writes.

Refresh the mark from each heartbeat's reported size in both heartbeat
paths (ApplyVolumeChanges and SyncDataNodeRegistration), mirroring what
RegisterVolume already did at registration time. A volume that grew past
the limit now stays unwritable, and one that shrank back clears the mark
and can recover.

* topology: order heartbeat writable correction after decay and honor cooldown

Review feedback (Greptile, CodeRabbit) on the oversized-mark refresh:

1. Greptile: clearing the oversized mark before EnsureCorrectWritables let
   the delay-unaware helper re-add a just-compacted volume to writables,
   bypassing capacityRecoveryDelay. ensureCorrectWritables now checks
   fullSince and skips the re-add while the cooldown is pending, so a
   volume removed for capacity only recovers through UpdateVolumeSize's
   heartbeat recovery path.

2. CodeRabbit: in the full-heartbeat path the mark was refreshed after
   the writable correction, so a newly oversized volume stayed writable
   for an extra heartbeat cycle. The standalone changedVolumes loop is
   merged into the volumeInfos loop and EnsureCorrectWritables now runs
   after UpdateOversizedState + UpdateVolumeSize in both heartbeat paths,
   using the freshly refreshed mark.

3. TestHandlingVolumeServerHeartbeat used a size (254320) that is past
   the test's volumeSizeLimit (32768); it only passed because the stale
   mark hid the oversized state. Sized down to 30000 to keep testing the
   add/remove flow, and added TestEnsureCorrectWritablesHonorsRecoveryCooldown
   covering the cooldown window and the recovery after it.

* topology: do not restore a still-crowded volume after the cooldown

Greptile review: after capacityRecoveryDelay elapses, ensureCorrectWritables
could restore a volume whose effective size is still past the crowded
threshold. UpdateVolumeSize refuses the recovery (effectiveSize > crowded
threshold -> setVolumeCrowded + return false), but the cooldown check in
ensureCorrectWritables only looked at fullSince, so once the delay passed
it re-added the volume even though capacity tracking still considers it
crowded.

Check the crowded mark before re-adding: a volume UpdateVolumeSize just
marked crowded must not be restored here, otherwise assignments resume
while the volume is still flagged for growth.

Adds TestEnsureCorrectWritablesDoesNotRestoreCrowdedVolume: effectiveSize
decays to 10500 (past the 9000 crowded threshold) after a report of 8000,
and ensureCorrectWritables keeps the volume unwritable past the cooldown.

* ci: trigger re-run of flaky FUSE jobs

* topology: gate the writable restore on the limit, not on crowded

A crowded volume is above the growth threshold, not full, and is normally
writable. Refusing to restore one locks it out for good: nothing writes to
a volume that is not writable, so its size can never fall back under the
threshold. Gate on the same size the assign path uses to remove it.

* topology: let only the heartbeat refresh set the oversized mark

Registration also set it, from whatever VolumeInfo it was handed. The
incremental path builds that from a short heartbeat message, which carries
no size, so every arrival announcement cleared the mark and handed the
volume back to the writable list until the next full report.

* topology: use the re-resolved layout after a dropped one is replaced

A layout dropped with its collection makes RegisterVolume refuse, and the
full heartbeat then re-registered against a fresh layout but kept applying
the size, oversized and writable updates to the dropped one.

---------

Co-authored-by: hzsunchao <hzsunchao@corp.netease.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-20 09:56:53 -07:00
Chris LuandGitHub 8c7d714d5e Lance catalog, and a Rust plugin worker to maintain it (#10841)
* iceberg: skip tables the maintenance worker does not own

A Lance dataset registered through the Lance namespace's Iceberg REST adapter
arrives as an Iceberg table with a placeholder schema and table_type=lance, and
keeps its fragments under data/ - the same subdirectory the orphan cleaner
walks. Every fragment is unreferenced by the Iceberg metadata, so a maintenance
pass deletes the dataset. Views share the entry shape and were only skipped
because parsing their metadata happened to fail first.

Gate the scan and the execution path on the entry actually being an Iceberg
table. Maintenance is off by default, so this was latent rather than live.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: let a table declare a format the catalog does not interpret

CreateTable accepted ICEBERG and nothing else. A Lance table has no metadata
file for the catalog to maintain - the entry records a name and the dataset
root, and the client owns everything under it - so accept LANCE, and carry the
declared format on the entry instead of hardcoding it back on the way out.

ListTables now reports format and metadataLocation, so listing a catalog that
holds both kinds takes one pass rather than a GetTable per row. AWS omits both
fields; adding them is additive.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: move the in-memory filer into its own package

The Lance namespace tests need the same harness, and copying it would leave two
of them to keep in step. Extracted as it was, plus the two fidelity gaps that
only surface once a paginating caller uses it: ListEntries ignored
startFromFileName and limit, so a caller that paginates re-read the first page
until it hit its own cap and reported the same entry over and over, and
GetFilerConfiguration was missing, which CreateTableBucket needs to resolve the
buckets directory.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: serve the Lance Namespace REST spec

A second catalog surface beside the Iceberg one, over the same table buckets:
the namespace and table metadata operations, the $-delimited identifier codec,
the spec's numeric error model, the directory-catalog marker files, and
storage_options vending through the STS path the Iceberg catalog already uses.
Listens on -port.lance, 9101 by default, and inherits ARNs, policies and tags
from the storage layer, so a Lance table needs no second permission model.

Identifiers map bucket / namespace / table onto the three levels Lance clients
already use, which is why there is no warehouse selector to invent. The data
plane needs Lance format support that does not exist in Go and answers with the
spec's Unsupported code rather than a bare 404.

Two things it deliberately will not do: create a table bucket as a side effect
of creating a namespace inside one, since a bucket carries its own policy and
lifecycle, and resolve an Iceberg table's location for a Lance client, which
would hand it a table another engine owns.

The design note this follows is in design-lance-catalog.md, including the
.lance directory suffix it proposed and this does not implement.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* mini: give the Lance port the same treatment as the Iceberg one

The flag was registered but nothing else knew about it, so mini would start the
server without reserving its port, waiting for it, or saying where it is. Adds
it to the startup service list, the conflict resolver, the gRPC allocator's
reserved set, the readiness wait, the stop reporting and the banner.

The admin server still takes only the Iceberg port, because there is no Lance
page for it to link to.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: stop deregister and repoint from deleting the dataset

Deregistering preserves data by definition, and this did the opposite: the
catalog entry is the dataset directory, so DeleteTable took the files with it.
Registering over an existing name had the same shape, destroying the dataset
the name used to hold. Found by driving the running server rather than the
in-memory filer, where both looked like success because the table did stop
being listed.

Deregistering is now a state on the entry - the marker file hides it, and
declaring or registering the name again brings it back. Repointing a name at
another dataset is an UpdateTable against the version token, so neither dataset
loses files. Drop is left alone; it is the operation that does remove data.

The storage endpoint now falls back to the advertised -ip where the Iceberg
derivation gives up. An Iceberg client brings its own s3.endpoint and
advertising the wrong one hijacks it, but storage_options is the only place a
Lance client learns where the store is, and without it object_store quietly
talks to real AWS.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: refuse to create a table over one of another format

Creating a table that already exists is idempotent, and that path returned the
existing table without looking at its format. A Lance declare over an Iceberg
table answered 200 and handed back a directory Iceberg owns, so the client
would write its dataset on top. The view check immediately above it already
guards the same class of collision.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: let a table bucket hold a format other than Iceberg

The S3 door validated every object written into a table bucket against
Iceberg's file layout, so a Lance client could not write its dataset at all: it
got 403 on data/*.lance, on _versions/, and on the _transactions/ directory it
turned out to write as well. Table buckets were only neutral containers by
intention; in practice they were Iceberg-shaped and enforced as such.

The allowed set is now the union of what the supported formats write, because
the validator runs where the table's format is not in hand. Underscore-prefixed
directories are treated as belonging to the format, since enumerating them
means guessing at the next one - _transactions is exactly the one this missed -
and their contents are checked only for traversal. Iceberg writes none of them,
so it loses nothing. Marker files at the table root are admitted too, which the
namespace/table/dir/file shape had rejected as too shallow.

Describe also honours the request-body spellings of with_table_uri,
load_detailed_metadata and check_declared. The spec puts them in the query
string, but real clients send both.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record what the implementation found

The table bucket being an Iceberg-shaped container, enforced at the S3 door,
was the premise this design never questioned and the one that had to change
before anything worked end to end.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* iceberg: prove the data loss the foreign-format guard prevents

The guard landed with a unit test for the predicate and nothing showing what it
saves. These seed what the Lance namespace's Iceberg REST adapter actually
leaves behind - an Iceberg table with a placeholder schema and table_type=lance
whose directory holds a Lance dataset - and assert both halves: orphan
collection does flag the dataset's fragments, because the Iceberg metadata
beside them references nothing, and the scan never reaches the table. An
ordinary Iceberg table in the same shape is still scanned, so the guard is not
just skipping everything.

Confirmed against a running gateway first: our Iceberg catalog accepts the
adapter's registration, and a real Lance client then writes a dataset into that
table's location.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tablestest: make the in-memory filer safe to race against

Two gaps that only matter once a test drives concurrent writers, which is what
an exclusive create has to be tested with: the entry map had no lock, and
CreateEntry ignored O_EXCL entirely, so both writers of the same name would
have won and the test would have passed while proving nothing.

The BeforeUpdate hook runs before the lock is taken. Its whole purpose is to
land a competing write in a handler's read-to-write window, and that write
needs the lock the hook would otherwise be holding.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: make the namespace an external manifest store

Lance commits a version by writing _versions/{v}.manifest with
put-if-not-exists. The S3 layer in front of this same filer evaluates
If-None-Match by looking the entry up and then writing without a precondition,
so two writers can both pass the check and one commit is lost. The filer itself
has the primitive: CreateEntry with o_excl.

Adds the four version operations a Lance client actually calls - create, list,
describe and batch-delete - recording one entry per version under
_lance_versions/, and advertises managed_versioning so the client routes its
commits here. Reserving a version is the exclusive create, so exactly one of
several racing writers wins and the rest rebase.

Off by default, behind -lance.managedVersioning. Turning it on moves where a
table's version history lives, and a reader that does not come through this
namespace no longer sees all of it; that is the operator's call, not a default.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record what managed versioning does and does not reach

The first commit through a namespace-backed store works and is recorded the way
the protocol specifies. Later commits do not, because lance 4.0.0 refuses
put_if_exists on that path in its own code, so the feature is capped upstream
rather than here.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: integration tests for the Lance namespace

Everything this surface got wrong so far - a deregister that deleted the
dataset, an S3 door that refused every Lance file, a version reservation that
could not actually be exclusive - passed against an in-memory filer first. So
these run against a live gateway, and where the claim is about data they check
storage rather than visibility.

Five Go tests on the shared harness: namespace and table lifecycle including
that deregister keeps the bytes and drop removes them, that a Lance client
cannot resolve or declare over an Iceberg table, that a Lance dataset's files
get past the table-bucket layout guard while junk still does not, and that
eight writers racing for one version produce exactly one winner.

One Docker-gated test drives the real Lance client, which is the only way to
check that the location and storage_options the namespace vends are between
them enough to write and read a dataset. It overrides the endpoint with the
container's view of the same gateway, because the shared harness binds a
wildcard address and so vends none.

The harness gains a Lance port and turns managed versioning on; the flag
touches nothing outside that surface.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: a directory with no namespace metadata is a missing namespace

Three callers resolved a namespace by reading its metadata attribute and each
tested only for a missing entry, so a directory that carried no metadata came
back as an internal error saying "attribute not found". Creating a table under
a namespace that does not exist answered 500.

Collapses the three copies into one helper that reports both conditions as
absent, which is what they are: a directory without namespace metadata is not a
namespace.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* iceberg: stop reporting storage-layer refusals as server faults

writeManagerError recognised a missing table bucket and sent everything else to
500, so a missing namespace, a duplicate name and a commit conflict all reached
the client as InternalServerError with nothing to act on. Creating a table in a
namespace that does not exist is the case that turned up: 500 where the spec
wants 404 NoSuchNamespaceException.

Maps the storage error types onto the exception names this package already
uses, and keeps the existing bucket message, which explains how to select a
table bucket.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* iceberg: skip a foreign-format table by name, not by failing to parse it

A table the namespace created as LANCE carries no Iceberg metadata, so the
worker skipped it only because the parse failed, and logged that as damaged
metadata. The catalog records the format on the entry and this never read it.

Reading it turns an accident into a decision, and separates a mixed catalog
from a corrupt one in the logs. The property check beside it still covers the
other shape: a real Iceberg table wearing table_type=lance, which is what the
Lance namespace's Iceberg REST adapter writes.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: answer whether a Lance table needs maintenance

It does, and index optimization has no Iceberg equivalent: rows written after
an index was built are not covered by it, so a vector search quietly misses
them. None of the three jobs can run in the Go worker, and there is no useful
subset, because deciding what an old version still references means parsing
Lance manifests. Version cleanup at least has an answer that needs nothing from
us - Lance can enable it on the dataset itself.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: the Lance maintenance worker is a plugin worker, in Rust

Framing it as a sidecar was wrong. plugin.proto already defines a
language-agnostic gRPC contract for external maintenance workers, and
"weed worker -admin=..." is the Go reference implementation of it from outside
the admin process. seaweed-volume already compiles protos out of weed/pb with
tonic_build, so a Lance worker is that build plus plugin.proto and the lance
crate.

Scheduling, retries, dedupe, progress and the admin settings page all come from
the protocol: a worker that answers RequestConfigSchema with a descriptor gets
its configuration form rendered without a line of Go.

The data plane is the part that genuinely does need a process answering HTTP,
and this had the two conflated.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* seaweed-worker: Rust plugin worker workspace, with Lance as the first one

plugin.proto is language-agnostic and the Rust toolchain was already in the
tree, so a Lance maintenance worker needs no new integration surface: core is
the contract and nothing else, and a worker crate beside it supplies handlers
and a binary. A second worker is a new member here rather than a fork of the
protocol, which is why this is seaweed-worker and not seaweed-lance-worker.

Verified against a running admin: it connects, is accepted, and admin prefetches
descriptors for lance_compact, lance_optimize_indices and lance_cleanup_versions,
so their settings pages render from the Rust side without a line of Go. The
stream stays up across heartbeats.

The job bodies are stubs that report failure. Doing the work means adding the
lance crate and opening the dataset, and claiming success before that would be
worse than saying so.

Two things running it caught that reading the proto did not: the admin address
has to be converted to the gRPC port the way pb.ServerToGrpcAddress does, or the
dial fails as an h2 frame error; and the generated field names differ from the
Go ones in several places, so JobCompleted carries success rather than a state
enum.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: implement compaction

Detection lists tables from the namespace, opens each one, and proposes a job
for any with more fragments than the policy allows; opening a dataset reads its
manifest and not its data, so a sweep stays cheap. Execution re-resolves the
table rather than trusting what detection saw - it may have been repointed, and
the vended credentials expire - then compacts and reports the fragment counts
either side.

Verified against a live gateway: a twelve-fragment dataset became one fragment
with all twelve rows intact. The test drives the handler directly and skips
unless WEED_LANCE_NAMESPACE names a namespace, the way the Go integration tests
skip without Docker.

Running it turned up a gap the design had not: a gateway without STS vends no
credentials at all, so the worker could not open anything and detection quietly
proposed nothing. --access-key/--secret-key are the fallback, and whatever the
namespace vends still wins over them.

Two API assumptions did not survive contact either. Datasets open through
DatasetBuilder::with_storage_options, not ReadParams, and lance 10's
ObjectStoreParams has no storage_options field at all.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: implement index optimization and version cleanup

Index optimization is the job with no Iceberg equivalent: rows appended after
an index was built are invisible to a search of it until this runs. Detection
reads num_unindexed_rows from each index's statistics and proposes a table once
more rows sit outside its indices than the budget allows; a table with no
indices is skipped, which is different from one whose indices have fallen
behind.

Cleanup applies a retention window, refusing rather than silently dropping a
tagged version, and leaving unverified files alone because they may belong to a
commit still in flight.

Both verified against a live gateway: 512 uncovered rows became 0, and a
fourteen-version table lost its old ones. Each test now seeds what it needs,
including building an IVF_PQ index and appending rows outside it. The first
version of these depended on state a script had left, so the second run found
the work already done and asserted nothing - a test that passes by doing
nothing is worse than no test.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: answer an empty catalog with an empty list, not null

ListAllTables built its result from a nil slice, so a namespace holding no
tables answered {"tables":null} on a field the spec marks required. A generated
client may decode that differently from an empty list. Found running the
namespace on a dev box, where the catalog was empty.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: give Lance maintenance its own scheduler lane

Lane assignment is a hardcoded map, so the three lance_* job types fell through
to the default lane. That lane serialises its work under the cluster admin lock
because volume management shares global state, which would queue a table's
compaction behind volume balancing for no reason - Iceberg has its own
lock-free lane for exactly this.

Adds the lane, maps the three job types to it, and puts it in the sidebar
beside Iceberg and Lifecycle. The lane routes were already generic, so only the
nav was hand-written.

The lane-coverage test spelled out the three known lanes, so a fourth failed
it. It now checks against AllLanes(), which is the property it was reaching for
and does not need editing next time.

Found by connecting the Rust worker to a real admin: it registered fine and its
job types were known, but they were filed under "default" and had no page.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: log what detection saw

"Detection proposed nothing" and "the worker could not read the table" look
identical from the admin side, and the second is what a missing credential
produces. One line per table separates them.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: fix a leaked heartbeat and a silent reconnect loop

spawn_heartbeat returned a handle to an empty task rather than the ticker it
had just spawned, so aborting it aborted nothing and every reconnect left
another heartbeat running against a dead channel.

A stream that admin closes cleanly is not an error, but reconnecting in silence
hides why. Two workers sharing an id evict each other forever and the log shows
nothing but a login every five seconds - which is exactly how this presented on
a dev box, and it took a look at the admin's own log to see it. The message now
names the id to check.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: a namespace cannot be created without its parent

Storage keeps a namespace's parts flattened, so creating "a.b" with no "a"
was accepted and left an intermediate that only existed inside a name. Listing
derives child names by slicing those parts, so it reported "a", while describe
and exists on "a" both answered 404 - a client walking the tree got a 404 on
something the listing had just handed it.

The spec asks for NamespaceNotFound when the parent is missing, which is also
what keeps listing and describe telling the same story.

Namespaces created through the S3 Tables API still bypass this, so listing
keeps deriving intermediates rather than hiding whatever is already there.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: say why a non-Iceberg table shows no schema

The table pages read Iceberg metadata for schema and snapshots, and a Lance
table has none, so both panels rendered "No schema available" - which reads as
an empty table rather than a table this page cannot describe. The dataset
behind the one that prompted this holds 1024 rows.

The format is already on the entry and shown two rows above, so the empty
states now use it: the catalog records where a LANCE table lives, not what is
in it.

Reading the schema for real needs Lance format code, which is the same wall as
the data plane.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* seaweed-worker: run rustfmt over the workspace

Committed the crates unformatted, so `cargo fmt --all --check` failed on
files nothing had touched since.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* plugin: let a worker report what it saw about an object

Admin cannot read a Lance table: it knows where the dataset lives and
nothing else, so the details page had a location and two empty panels.
The worker already opens every dataset during detection to decide whether
it needs compacting, so it knows the schema, the row count and the
fragment count at that moment. It just had no way to say so.

Add a WorkerObservations body to the worker stream. Admin caches the last
observation per object and serves it back, timestamped, for display;
nothing schedules from it. The Lance compaction sweep reports what it
opened, and the S3 Tables details page fills its schema panel from the
cache when it has no metadata of its own, badged with when the worker
looked and which worker it was.

Nothing about this is Lance-specific past the reporting side, which is
the point: any format admin cannot parse can describe itself the same way.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record the observation channel

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* plugin: ask a worker for sample rows of a table admin cannot read

Browse Data reads an Iceberg table's Parquet files directly, so it shows
real rows. For a Lance table it showed "Table has no Iceberg metadata"
and an empty grid, because there is no Go Lance reader and never will be
one worth maintaining.

The worker has the reader. Add RequestObjectPreview / ObjectPreviewResponse
to the stream, mirroring the config-schema round trip that already exists,
and give the Rust worker a PreviewProvider that scans the dataset and
formats the rows with Arrow's own formatter, so a vector column reads as a
vector. Admin picks the worker from the observation store: whichever one
last described this table is the one that can read it.

Unlike an observation the rows are not cached. They are the table's data
rather than a description of it, and a copy sitting in admin would be both
stale and nobody's business. The page fetches on load, bounded at 200 rows
and a 15 second round trip, and drops the snapshot and data-file panels
that only mean something for Iceberg.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record the preview channel

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: disable the lance listener when two gateways share a host

* test: keep AllocatePorts away from the lance default port

* s3tables: let a table bucket declare the format it holds

A bucket is a catalog, and a catalog serves one protocol. Format was
recorded per table, so nothing could answer "where do I point a client at
this bucket" without opening a table first, and an empty bucket had no
answer at all.

CreateTableBucket takes an optional format, stored with the rest of the
bucket metadata and returned by Get and List. Empty means ICEBERG, which
is what AWS S3 Tables serves and therefore what an SDK that has never
heard of the field means. CreateTable refuses a table of another format,
and CreateView refuses outright in a bucket that is not Iceberg, since a
view is Iceberg metadata.

Buckets that already exist carry no declaration and keep accepting
anything, so nothing is migrated and nothing that worked stops working.
The Lance namespace declares LANCE for the buckets it creates, which is
what stops one of them being described to a client as an Iceberg catalog.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: take the Lance port the way it takes the Iceberg one

The UI cannot name the endpoint that serves a Lance bucket without it,
and every format-aware page below needs to.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: show which format a table bucket holds

The bucket list printed an Iceberg endpoint for every bucket, including
ones holding Lance datasets, where that endpoint serves nothing. It was
the most visible place the UI assumed one format.

The list gains a Format column and its endpoint column follows the
bucket's declaration. The banner names both endpoints rather than
asserting everything is Iceberg, and says so only for the servers that
are actually running. Create Bucket picks a format with two cards rather
than a dropdown, since what matters is not the name but which clients can
read the result, and the endpoint under them updates as you choose so the
operator leaves the modal knowing where to point one.

A bucket from before the declaration existed shows "unset" in an outline
badge, explained on hover. It is a fact about the bucket's age, not a
fault, so nothing nags about it.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: carry the bucket's format into the pages inside it

Namespaces and tables are reached through a bucket, so both now say which
catalog they belong to rather than making you go back up to find out. The
tables list gains a Format column and a Rows column filled from what a
worker last observed, since for a format admin cannot read that is the
only row count there is; a table nothing has looked at shows a dash, not
a zero.

Create Table stops offering a choice the bucket has already made: in a
declared bucket the format is fixed and says why, and only an undeclared
one still offers both. Before this the select had exactly one option,
hardcoded, which made a Lance table impossible to create from the UI at
all.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: let the table page speak the table's own format

Partitions and Snapshot History are Iceberg's shape. Rendering them empty
for a Lance table reads as a fault; a Lance table has neither, and says
so by not showing them. In their place is a Versions panel, which is what
that format calls its history, carrying the worker's timestamp so it is
clear the numbers are a cached look rather than something read live.

The breadcrumb carries the format badge, so the page names what it is
looking at before you read a panel and wonder why it is empty.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: show how to connect to either catalog, and group the two format workers

The client examples on the buckets page were Iceberg's alone, so the one
thing an operator wants after creating a Lance bucket - what to type to
reach it - was not written down anywhere in the UI. Both formats now get
a pair of snippets, and only for a server that is running.

In the Workers menu, Iceberg moves below Lifecycle so it sits next to
Lance: the two table-format workers together, the two cluster-wide ones
above them.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* shell: create a table bucket of either format

s3tables.bucket -create takes -format, so a Lance bucket can be made
without going through the UI. The integration harness passes it too: its
Lance tests were creating Iceberg buckets and getting away with it only
because nothing checked.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record that a bucket declares its format

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: drop managed versioning; the store already orders commits

The namespace offered itself as an external manifest store, so that a
commit could reserve a version through a real put-if-not-exists. That was
designed around a gateway that no longer exists: If-None-Match: * is
reduced to a filer WriteCondition and evaluated at the object's owner
under its per-path lock, or under the object write lock on the fallback
path. Sixteen writers racing one fresh key get a single 200 and fifteen
412s, every time.

Lance needs nothing else. commit_handler_from_url hands every s3:// dataset
a ConditionalPutCommitHandler, which puts with PutMode::Create, which
object_store sends as If-None-Match: *. So the feature solved a problem
this store does not have, while moving a table's version history out of
the dataset and into the catalog - and lance could not use it past the
first commit anyway, since its own namespace-backed store answers
"put_if_not_exists is not supported" to the second.

The version operations answer Unsupported with the rest, managed_versioning
is false, and the flag is gone. In place of the reserve-once test there is
one that races eight writers at the manifest key through S3, which is the
path a commit actually takes.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: honour the version floor, the slot limits, and a shutdown

Five findings from review, all of them things the worker claimed to do and
did not.

The version floor was checked when a cleanup job was proposed and ignored
when it ran, so a table whose versions had aged past the retention window
in between could be taken below the count the operator asked to keep.
Execution now computes the floor itself and passes it as before_version;
CleanupPolicy ANDs its clauses, so a version has to be both too old and
below the floor to go. Both settings are clamped to the range the form
offers, since Duration::hours panics on a large enough value and a
negative min-versions wraps to a huge usize.

Admin's shutdown was answered by returning from the stream, which the
reconnect loop read as a healthy close and logged straight back in: the
worker could not be stopped. serve_once now says which of the two
happened.

The advertised concurrency limits bounded nothing - every request spawned
a task - and the heartbeat reported zero slots in use whatever was
running. Both now go through semaphores sized from the limits, with the
permits held for the life of the request and reported in the heartbeat.

A namespace call had no timeout, so a gateway that accepted the connection
and went quiet held a detection slot forever. And one table whose stats
could not be read failed the whole sweep, losing the proposals for every
table already scanned; it is now skipped and warned about, like a table
that cannot be opened.

The tests drove one shared catalog concurrently, which is why one of them
asserted "no proposals at all" and passed by luck. They now take a lock
and judge only their own tables.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: fix the review findings on the format-aware pages

The endpoint hint in Create Bucket built its HTML by concatenating the
bucket name the operator is typing, so a name like <img onerror=...> ran
in the admin origin as they typed it. It is built from DOM nodes now.

A preview reply looked its channel up under the lock and then sent outside
it, which Shutdown can close in between: a Gosched in that gap panics with
"send on closed channel" every time. The send now happens under the lock.

Observations were looked up by path alone, so a table dropped and remade
in another format at the same path was described by the observation left
behind. Lookups now have to agree on the format.

Also: the Lance namespace caps a request body rather than reading whatever
arrives; the details action no longer says "Iceberg" over a Lance table;
mini stops advertising a catalog port when it is not running S3; a format
whose server this cluster does not run cannot be picked in the modal or
accepted by the API, since a bucket nothing can reach is not worth
creating; and the unused catalogPortFor helper is gone.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: let the control stream use mTLS

The channel was hardcoded to http://, so off loopback the stream carried
preview rows and execution commands in the clear - and a cluster with grpc
TLS turned on would refuse the worker outright.

--tls-ca, --tls-cert and --tls-key take the same certificates the Go
worker reads from the [grpc.worker] section of security.toml, and must be
given together: a CA on its own would quietly mean one-way TLS, which a
mutual setup rejects anyway. Without them the stream stays plaintext,
which is what the Go worker also does when nothing is configured.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: answer null properties rather than an empty map

The catalog does not keep a table's properties. Declare echoed the
request's back and describe answered {}, both of which claim they were
stored and are empty. Null says the catalog does not keep them, which is
what the spec distinguishes and what is true here.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: test the slot accounting

The heartbeat reporting and the waiting are the two things the semaphores
are for, and neither is observable from outside without catching a sweep
mid-flight.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: fix the mixed-format catalog test, and name the binary it drives

The integration suite passed locally and failed in CI on
TestLanceRefusesIcebergTables. Both were right: CI builds the binary
first, my tree had one from the day before, so locally the test drove a
gateway with no format enforcement at all.

The test itself no longer holds as written. It made a bucket, put an
Iceberg table in it, and checked the Lance surface hid it - but a bucket
that declares LANCE now refuses the Iceberg table outright. The invariant
still matters from the other side, so it starts from an Iceberg bucket
instead: Lance must not describe or list a table whose format it does not
serve, and must refuse to declare one beside it.

The harness now prints which weed binary it is about to run and when that
was built. `make test` rebuilds first; a plain `go test` will happily
drive a weeks-old binary and report a pass for code it never ran, which is
exactly what happened here.

Also make the row-limit conversion in the preview request explicitly
bounded: CodeQL flagged the int-to-int32 conversion, and clamping by
reassignment beforehand is not a form it recognises.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: prove concurrent commits are kept, and preselect the only format on offer

Two more from review.

The commit test asserted that exactly one writer wins the conditional PUT,
which is the mechanism, not the claim. The claim is that nothing is lost:
the losers see the conflict, rebase and commit again. So there is now a
test that has eight writers append to one dataset at once and counts the
rows afterwards - all eight batches survive. That is also the sequence
managed versioning could not finish, since its store refuses the second
commit outright.

And when Iceberg's endpoint is not running, the format picker offered two
options with neither selected, so Create Bucket submitted no format at
all, fell back to ICEBERG, and was refused by the guard added last round.
Lance is preselected when it is the only format this cluster serves.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* Clamp the remaining worker settings, and bootstrap buckets in a served format

Compaction and index optimization read their thresholds and cast straight
to usize and u64, so a negative arrives as an enormous number and turns
the threshold into "never": compaction and reindexing both go quiet with
nothing to say. The cleanup job was fixed last round; these are the same
bug. Clamped to the values that stay meaningful rather than to what the
form offers - zero uncovered rows is a real setting, meaning reindex as
soon as anything is not covered, so the floor there is zero and not the
form's thousand.

mini pre-creates the buckets named by -tableBucket, and did so without a
format, which now means Iceberg. Started with the Iceberg endpoint off
and the Lance one on, that left buckets nothing could reach and which
refused every Lance table. It takes the format from the endpoint that is
actually running, and creates nothing when neither is.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3: allow-unordered is a listing parameter, not an unimplemented subresource

The guard that stops a bucket GET with an unknown subresource from being
answered with a listing does not know about allow-unordered, so it answers
501 NotImplemented - to a parameter the listing handlers already read and
already validate against delimiter.

This is why test_bucket_list_unordered and test_bucket_listv2_unordered
fail in the Ceph s3-tests suite. They fail on master too; this is not a
Lance change and can be taken on its own.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-19 22:59:56 -07:00
Chris LuandGitHub 814ee75af4 s3: allow-unordered is a listing parameter, not an unimplemented subresource (#10846)
The guard that stops a bucket GET with an unknown subresource from being
answered with a listing does not know about allow-unordered, so it answers
501 NotImplemented - to a parameter the listing handlers already read and
already validate against delimiter.

This is why test_bucket_list_unordered and test_bucket_listv2_unordered
fail in the Ceph s3-tests suite. They fail on master too; this is not a
Lance change and can be taken on its own.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-19 20:09:23 -07:00
Chris LuandGitHub e5edd8be3c s3: place multipart part chunks by the destination object's storage rule (#10845)
Multipart parts stage under /buckets/<bucket>/.uploads/<id>/, so the filer
resolved filer.conf storage rules against that path when the gateway assigned
volumes for them. A rule scoped to a key prefix - fs.configure
-locationPrefix=/buckets/b/data/ -ttl=30d - then matched a small object but not
the parts of a large one, so an object whose entry carried the rule's TTL had
its bytes spread over TTL-less volumes.

Assign part chunks against the destination object's filer path instead, the way
the x-seaweedfs-destination header made the filer resolve it before the S3 write
path moved off the filer proxy. Covers PutObjectPart and both UploadPartCopy
paths. The part entry itself is still written under .uploads, so a read-only
rule there still rejects it.

The lifecycle XML Expiration.Days TTL keeps passing 0 for parts: that rule
targets the user-visible object key and would start its clock before
CompleteMultipartUpload.
2026-08-19 20:07:57 -07:00
Chris LuandGitHub abd61de52c S3: stamp the gateway's own uid/gid on PutObject and copy entries (#10844)
* fix(s3): stamp the gateway's own ids on single-shot PutObject entries

putToFiler builds the entry in the gateway now instead of proxying a PUT to
the filer, and it hardcoded Uid/Gid 0 while every sibling write path stamps
filer_pb.OS_UID/OS_GID. On a non-root deployment that leaves single-shot PUTs
and multipart parts owned by root while directories and completed multipart
objects keep the real ids, so a mount reader can list the tree but gets EACCES
on every open once objects are not world-readable.

* fix(s3): stamp mode and ownership on copy destinations

CopyObject and UploadPartCopy build the destination attributes themselves and
then assign them over the entry filer_pb.MkFile just stamped, so the copy
landed with mode 0000 and uid/gid 0 - unreadable on a mount even by the filer's
own user. Build the destination with the same mode PutObject resolves for the
request and the gateway's own ids.
2026-08-19 18:57:37 -07:00
5d5fcdf07b fix(filer): bound aggregated metadata reads by peer watermarks (#10803)
* fix(filer): watermark-bound aggregated metadata subscription against multi-source merge races

The aggregated metadata subscription (SubscribeMetadata) merges per-filer
sources that become readable at independent paces, but tracks its progress
with a single scalar cursor. Once the cursor passes a timestamp T, anything
a source materializes below T afterwards is silently skipped: a peer
recovering from a stall re-inserts its backlog late (late ring merge), and
a source's flush can land a log file, or a later chunk of the same file,
after a subscriber's disk pass listed the files (late persisted-log
landing). This is the residual documented in #10501.

Bound the subscriber's two read paths by what every source has provably
made visible, each with its own watermark:

- Delivery low-watermark -> in-memory reads. The meta aggregator tracks,
  per subscribed peer (self included), the newest timestamp received on
  that peer's stream - real events, or idle heartbeats (peer streams now
  opt into ClientSupportsIdleHeartbeat). The aggregated ring is complete
  up to the minimum across peers; in-memory reads hold at it.
- Flush low-watermark -> persisted-log reads. Each filer reports its local
  log-buffer flush watermark on its stream: a new flushed_ts_ns response
  field, carried on idle heartbeats and on periodic flush reports (gated
  on ClientSupportsIdleHeartbeat). Disk passes freeze the minimum across
  peers before listing the log files and hold at it; the day-boundary
  cursor jump and the metadata-chunks ref listing are bounded the same
  way, the latter at minute-file granularity.
- Held reads keep the cursor at the last entry actually delivered and
  retry; the retry re-lists the log files, which is what picks up a
  late-landing file. Both watermarks are relaxed by the settled horizon
  (2 x LogFlushInterval) as a liveness escape, so a peer stalled beyond it
  delays subscribers by at most the horizon instead of forever - any loss
  that escape allows was unconditional before.

With reads held at the flush watermark, a disk advance below it is proven
complete on every peer's disk, so the unproven-crossing counter now only
counts crossings the horizon escape allowed past a stalled peer.

Live delivery on the aggregated stream may lag by up to the idle-heartbeat
interval when some peers are quiet; SubscribeLocalMetadata consumers are
unaffected.

* fix(filer): resume evicted aggregated readers from an original-space disk anchor

The aggregated ring rewrites out-of-order peer arrivals to its head, so a
subscriber tailing it advances its cursor in bumped (arrival) timestamps,
while persisted logs keep original timestamps. When a slow reader's unread
window is evicted (e.g. a peer backlog flooding in after a stall) and the
reader falls back to disk, resuming from the bumped cursor skips every
original-space entry below it that memory never delivered - reproduced as
a ~66% silent loss on a 3-filer cluster with one peer's stream frozen for
~70s while the subscriber lagged.

Track a disk anchor: the newest original-space position the stream is
proven complete through. Disk passes advance it directly; contiguous
memory reads advance it to the peers' delivery low-watermark observed
before the read (per-peer streams are ordered, so everything with an
original timestamp at or below that watermark had already arrived and was
delivered). A reader kicked off the ring resumes the disk pass from the
anchor instead of the bumped cursor - redelivering what memory already
sent is within the subscription's at-least-once contract, skipping what
it never sent is not.

* fix(filer): close review findings on the peer-watermark subscription bounds

Four correctness holes found in review, one generated-file cleanup:

- The flush-through claim could assert durability for events still on
  their way into the buffer: an event is timestamped before notification
  work that can block, and only then appended. Track stamped-but-unappended
  events on the Filer (the stamp shares a lock with the reader, and appends
  are bumped monotonically past the buffer head), and cap the reported
  flush watermark just below the oldest in-flight stamp.

- Removing a peer deleted its watermark entries while its stream kept
  running: its next signal recreated the deleted entry, which then pinned
  the low-watermark forever once the stream died. Watermarks now advance
  only for tracked peers, and peer removal cancels the subscription
  context so the stream stops feeding the aggregated buffer promptly.

- The pipelined sender folded flush reports (TsNs 0 reads as far behind)
  into batch Events tails, where the aggregator's nil-notification guard
  dropped them - a busy backlog replay could starve the flush watermark
  until the settled-horizon escape opened a loss window. Control messages
  are now unbatchable on the sender, and the receiver also reads watermark
  state off nested batch entries as belt and braces.

- A give-up skip's cursor was not anchored, so the next eviction rewind
  undid the counted decision and re-entered the same park forever when the
  evicted window carried bumped timestamps. The anchor now follows give-up
  skips; an anchored cursor makes the rewind a no-op and keeps the gap
  machinery's re-arm onto the retained window reachable.

- Regenerated-file churn from a different protoc-gen-go-vtproto version is
  dropped: the vtproto file is upstream's, plus only the flushed_ts_ns
  marshal/size/unmarshal cases in the same generator style.

New tests pin the in-flight floor, the no-resurrection rule for removed
peers, and that control messages are never nested in batches.

* fix(filer): keep a removed peer's watermarks through a grace period

Deleting a peer's watermark entries the moment the master removes it
reopened the loss the watermarks exist to prevent: a filer frozen or
partitioned long enough to miss master heartbeats is removed from the
cluster, its unflushed events still exist, and with its entries gone the
low-watermarks snap forward to the healthy peers - subscribers advance
past the absent peer's window and its late-landing log files are silently
skipped. Reproduced on a 3-filer cluster: freezing two filers for ~70s got
them removed ~28s in, and a catching-up subscriber lost their entire
overlapping window.

Removal now only marks the peer; its watermarks keep participating in the
low-watermarks for a grace period (2 x LogFlushInterval, matching the
subscribe loops' settled horizon, which already bounds a stale watermark's
influence meanwhile). A re-added peer clears the mark and continues its
values monotonically - the flap case costs nothing. A peer that stays gone
is dropped when the grace expires, so a decommission cannot pin the
low-watermarks, and a dropped peer's straggling signals cannot resurrect
its entry.

* fix(filer): cap delivery heartbeats by the in-flight floor; harden stamps

Second review pass on the watermark bounds:

- Idle heartbeats on the local stream claimed delivery-completeness
  through "now" while an event could still sit stamped-but-unappended
  behind blocking notification work. A peer aggregator turns that claim
  into its delivery low-watermark, so it could advance (and anchor
  credits with it) past an event that had not been streamed yet. The
  heartbeat timestamp is now capped just below the oldest in-flight
  stamp, like the flush claim already was.

- In-flight stamps are forced monotonic against the registry's own
  history, so a wall-clock step backwards cannot slip a new stamp under
  an already-sampled floor. The cross-goroutine ordering still shares
  the meta log's global forward-clock assumption; the comments now say
  so instead of overclaiming.

- Duplicate removal notifications no longer refresh a removed peer's
  grace deadline: the first removal time wins, so a decommissioned peer
  cannot sit in the watermark sets forever on repeated updates.

- A failed buffer append clears the event's in-flight stamp on purpose:
  the event is dropped from the change stream entirely (a pre-existing
  defect of the append path, loudly logged), and a watermark waiting for
  it would pin this filer's claims forever. The comments now state the
  decision instead of implying the failure cannot happen.

* docs(filer): tighten the watermark comments

Comment-only: compress the narrative comments added on this branch down
to their load-bearing invariants, and fix one stale sentence (peer
removal no longer deletes the watermark entries immediately). No code
changes.

* fix(filer): subscribe to the local filer before remote peers

Self's events reach the aggregated buffer only through the aggregator's
own subscription to it, but bootstrap only seeded the peers the master
already listed - and self's master registration races that listing, so
the watermark set could hold remote peers without self. Once the remotes
signalled, the low-watermarks would claim completeness for a stream that
was still missing a merge source, letting aggregated subscribers advance
past the local filer's events before its subscription started.

Seed self first, unconditionally: before that the watermark set is empty
(a documented safe state - reads hold at the settled horizon), and after
it the set can never be remotes-only. The later master update for self,
or a duplicate in the listed peers, is a no-op via the already-followed
check in OnPeerUpdate.

* fix(filer): fence watermark claims against wall-clock regression

Record issued heartbeat/flush claims in the in-flight registry and stamp
later events above them, so a backward clock step cannot land an event
under a watermark a peer has already advanced to.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(filer): re-check the buffer head after fencing heartbeat claims

An event appended between the caught-up check and the delivery claim
was covered by the claim but not yet sent on the stream. The claims
fence later stamps, so re-checking the head after them proves every
covered event was already sent before the heartbeat.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(filer): cross the aggregated ring's pre-subscription range only on proof

The eviction gate and the gap proofs read "nothing evicted yet" as "memory
holds everything after the cursor". That is false for the merge-fed
aggregated ring, which is born empty while every peer's history sits on
disk: before the ring's first real eviction, a subscriber whose cursor was
still below the bounded chunk pass's listing stop was served the ring's
earliest entry inclusively, silently skipping the withheld pre-restart
files - and the idle-wait callback credited the delivery low-watermark to
the disk anchor in the same disconnected state.

Mark everything at or below the subscriptions' start as evicted when the
aggregator is built, credit the anchor only once the run is connected to
the ring, and give the aggregated gap pass a real proof to cross the
marked boundary with: each disk pass's proven coverage (the peer flush
low-watermark capped by the pass's listing bound). An empty pass whose
proof reaches the eviction watermark crosses to it silently - no park, no
loss counter - so the mark costs a bounded catch-up delay instead of the
15-minute give-up.

* fix(filer): keep shipped chunk tails at or below the hold point

A log file spans past its named minute (window start plus up to a flush
interval), and chunk-mode clients apply a shipped file whole - so a file
tail past the hold point can become a persisted client checkpoint beyond
what every peer has proven, and a crash inside that window resumes past
another peer's late-but-in-contract flush. Stop the ref listing a minute
plus a flush interval below the hold; the withheld band is served by the
memory pass (ring retention far exceeds it) or by later passes as the
hold advances, so freshness is unchanged. A frozen peer flushing one
window that spans its whole freeze can still overshoot; that residual is
bounded by the freeze and needs a crash inside it.

* docs(filer): trim the review-fix comments

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-19 18:38:46 -07:00
Chris LuandGitHub 804111745a mount: discard a path-cache insert that raced a purge (#10842)
* mount: discard a path-cache insert that raced a purge

The Windows adapter's walk resolves a component with a Lookup RPC and inserts
the result holding no lock, so a purge can land in between - and what the walk
just resolved is then the very name the purge removed. Anything opening the
old path concurrently with a rename repopulates the cache with the vacated
name, which the next stat is served from for up to a second. The release path
already guards its equivalent insert; the walk had nothing.

The cache counts purges now. A resolve snapshots the generation before its
lookups and insert discards the entry when any purge ran in between, parking
the reference in the graveyard so the in-flight caller keeps a valid inode
either way.

Seen once in CI as TestRenameOverExisting failing with 'source survived the
rename': every SeaweedFS layer is synchronous with the rename, but a
background open of the source - an antivirus scan of the just-written file
fits - can requalify the stale name through this window. The assertion also
reports what stat returned now, and whether it persisted, so a recurrence
indicts a specific layer instead of reading as a mystery.

* mount: cover the path-cache discard by key, and let a discard rest

Review follow-ups. The generation was global, so any purge between a walk's
snapshot and its insert discarded the entry whatever its name - and an open
retries resolve-then-steal only four times before failing with EIO, so
sustained unrelated churn could fail opens of untouched paths. Purges are
remembered by key now and only one that covers the inserted name discards it;
past the remembered window the insert is discarded without a check, which only
costs a retry.

A discard that itself tripped the sweep also handed its own reference straight
to forget while the walker was still using the inode. The graveyard holds two
generations now, so an appended reference always survives the sweep of the
call that appended it - which the displaced-entry and purge paths needed too.

Also restores the original path-cache test suite this branch had overwritten
instead of extended, and rewords the semantics-test failure so it no longer
claims the source survived when stat returned a transient error.

* mount: take an open's reference directly instead of stealing it back

resolveAndSteal cached the final component only to steal it back, so an open
depended on that insert surviving whatever purges raced it - four attempts and
then EIO. The keyed purge window narrowed how often an insert is discarded,
but past the window the discard is blind again, so the cliff had only moved.

A cached entry is still stolen; anything else is now looked up directly, with
the caller owning the reference from the start. No retry loop, and no way for
churn - covered, unrelated or overflowing the window - to fail an open.

Also covers the whole-cache purge: purge of the root with prefix set clears
every entry, but the covers check tested for a '/'-prefixed key that a
normalised key never has, so it covered no in-flight insert at all.
2026-08-19 17:05:49 -07:00
Chris LuandGitHub da1e5e714f mount: move the inode table when a rename arrives from the cluster (#10822)
* mount: leave the target alone when a move has no source

MovePath cleared whatever sat at the target before it checked that the source
was still there, so a move it then declined to make had already taken the
target's mapping apart. The same rename reaching the table twice - once for an
open handle, once for the invalidation behind it - was enough to leave the
moved inode with no path at all.

* mount: move the inode table when a rename arrives from the cluster

A rename made by another client reaches this mount only as a metadata event,
and the only table update on that path sat inside the open-file-handle branch.
Every other inode the kernel still addresses by nodeid kept resolving to its
pre-rename path, so the next operation on it went to a path the filer no
longer has.

Move the entry for every rename invalidation. The filer sends one event per
moved entry, so a renamed directory's children follow their parent without a
descendant walk.

* mount: move the inode table exactly once per rename event

Making MovePath return early on a missing source was the wrong half to fix.
The source is also missing when the rename came from a client that never
visited it, and there the destination really was replaced and has to be
unlinked - the early return kept the name resolving to a file the rename
destroyed, so a dirty handle on it could still flush over what took its place.

The two cases are indistinguishable from inside MovePath, so leave it alone and
stop calling it twice: invalidateOpenFileHandle reports whether it moved, and
the handler moves only when it did not. Both paths mark a replaced file's
handle deleted, which the no-handle path previously did not do at all.

* mount: leave a rename alone unless the source is still ours to move

Two ways the fallback move could act on state it did not own. A handle whose
version guard skipped the event never reached RememberPath, so moving the table
under it left the handle flushing to the pre-rename path; the handle path owns
its inode's rename, so the fallback now runs only for an inode without one.

And the subscription can redeliver a rename once it falls out of the 4096-entry
dedup ring. A replay found no source and a live target, unlinked the mapping
the first delivery had just made, and marked the moved file's handle deleted -
worse than the stale path this set out to fix. Only a source still in the table
is moved now.

That gives up unlinking a destination the rename replaced when the source was
never visited here, which is where this started. It is what the mount already
did before this branch, and it is the safer of the two: retaining a stale name
costs a wrong lookup, while unlinking the wrong one costs a file's dirty data.

* mount: decide a rename move inside the table's lock

The source-presence check sat outside MovePath, so two invalidations for one
rename could both see the source and the loser would unlink what the winner had
just placed - the same damage the check was added to prevent. MovePath makes
the decision under its own lock now and reports that nothing moved.

The handle a rename destroyed is also marked from the caller rather than from
inside invalidateOpenFileHandle, which was setting isDeleted bare on a second
handle while holding the first one's lock. markHandleDeleted already takes the
lock the flush reads that flag under, and marking from the caller keeps it to
one handle lock at a time.

The invalidation test harness wires onEntryInvalidation now, the way the mount
does, rather than reaching past it.

* mount: apply a rename ahead of the handle's version fence

The fence exists so an old event cannot roll a handle's entry back to stale
content. A rename carries no content: it says the name the inode answered to is
gone. Skipping one on the strength of the fence left the inode and the handle
both pointing at a name the filer had vacated, and no fallback ran either,
since a handle owns its inode's rename.

Applied before the fence now, and only when MovePath reports the source was
still ours to move - which is what keeps a replayed rename from remembering a
path the handle has already moved past.
2026-08-19 14:57:45 -07:00
9a8b204a7e Add monitoring label to filer servicemonitor (#10835)
* Add monitoring label to filer servicemonitor

* helm: label the headless filer service, not the client one

The ServiceMonitor takes its job label from the service name, and the
bundled dashboard queries job="seaweedfs-filer". Selecting the client
service would have renamed the job and blanked those panels. The
headless service also publishes not-ready addresses, so it keeps
reporting while a filer is starting up or shutting down.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-19 14:21:36 -07:00
Chris LuandGitHub cd3db76eed ci: stop installing FUSE headers nothing links against (#10840)
Four workflows ran apt to install libfuse3-dev before every FUSE job. Nothing
needs it: go-fuse implements the protocol in pure Go, no cgo in the tree
references fuse, and the package does not even provide the fusermount3 the
mount actually execs - fuse3 does, and it is already on the runner image, which
is why the setuid-repair step finds it.

So the step downloaded a dev package to build against headers no compiler ever
opened, and it is the step that has been hanging whenever the Ubuntu mirror
goes slow. Configuring /etc/fuse.conf is all that is left.
2026-08-19 14:09:40 -07:00
Chris LuandGitHub e0c4732e5e rust: stop writing when a durable write's index flush fails (#10825)
* rust: stop writing when a durable write's index flush fails

A durable write flushes the .dat, publishes the needle map row, then
flushes the .idx. If that last flush failed we returned the error and
carried on: the row stayed live, the volume stayed writable, and the
handler answered 500 without replicating. The primary then served a needle
its replicas never saw, for a write the client was told had failed - and
if the unflushed row was lost on restart, the durable .dat tail took the
volume read only anyway.

Taking the row back out is not an option: it means undoing published state
on a disk that is already failing, and a truncate afterwards would leave
an .idx row pointing past the end. So the volume stops taking writes
instead, the same as when the truncate after a failed .dat flush cannot be
done. Nothing more gets appended past a record whose index is in doubt,
and the master routes writes elsewhere once the volume heartbeats read
only. The divergence against the replicas is still there, but it is
bounded and it is visible.

A failed nm.put after the .dat is down leaves the same durable but
unindexed record, so it takes the same route.

* rust: drop the import the rollback removal left behind

NeedleValue came in with rollback_unflushed_write, which went away when
the durable path moved to flushing before it publishes. Nothing has used
the type since.

* rust: mark the test-only heartbeat helper as such

collect_heartbeat has only ever been called from the tests - the send loop
uses collect_heartbeat_with_snapshot, which it wraps - so a lib build
rightly called it dead code.

* rust: flush the index on a durable write that dedups

A durable write matching content already in the volume flushed the .dat
and returned before reaching the index flush. So a fsync=true write that
deduped against an earlier non-durable one was acked with the row that
indexes it still in the page cache - the same false promise the index
flush exists to rule out, and the same read-only volume on restart if the
row is lost.

The dedup path now flushes both files, and the quarantine on a failed
index flush moved into flush_idx so it applies wherever the flush is
reached rather than only at the one call site that had it inline.
2026-08-19 14:01:19 -07:00
Chris LuandGitHub 3b18a635df ci: call the apt helper from the workflow's working directory (#10832)
The e2e workflow sets defaults.run.working-directory: docker, so the call I
added resolved to docker/docker/apt-install and every FUSE Mount run has
failed with 'sudo: docker/apt-install: command not found' since it merged.
2026-08-19 00:46:32 -07:00
Chris LuandGitHub baead6901c ci: build protoc into the crate instead of installing it per job (#10830)
Every workflow that builds the Rust volume server first installed protoc
from a package manager - twelve steps across apt, brew and choco. That is
37s per job on a good day, and this week archive.ubuntu.com stalled long
enough for four jobs to burn their whole timeout without reaching a build.

protoc-bin-vendored ships the compiler as a build-dependency, so it now
arrives through the cargo registry the workflows already cache and there
is nothing left to install. cargo build works on a machine with no protoc
at all, which is worth as much locally as it is in CI.

It also pins the version. The apt protoc on ubuntu-22.04 is 3.12, old
enough to reject proto3 optional, which is why build.rs passes
--experimental_allow_proto3_optional; the vendored one is 31.1. The flag
stays, since it costs nothing and keeps a build against an older PROTOC
working, and an explicit PROTOC still overrides the vendored binary for
packagers who supply their own.
2026-08-19 00:18:30 -07:00
Chris LuandGitHub 1564244b1a ci: install the runner's own packages through the mirror fallback too (#10831)
The e2e job overwrote the runner's sources.list with two azure-only lines and
installed fuse from it, so the same mirror outage that took out the image
builds failed the step outright - this time on the runner rather than inside
the container, where the image-side fallback cannot reach.

Install through the same helper, and widen its rewrite to match any archive
host so it works whether the pristine list came from the base image
(archive.ubuntu.com) or from a CI runner (azure.archive.ubuntu.com). Keeping
the runner's original list also restores the security and backports pockets,
which the hand-written two-line replacement dropped.

Verified against the outage itself: with the pristine list pointed at Azure,
the build logged the skip after Azure timed out for real and installed from
archive.ubuntu.com.
2026-08-19 00:17:51 -07:00
Chris LuandGitHub 05013ad3da ci: fall through to another Ubuntu mirror when one is unreachable (#10828)
The e2e image pointed both archive and security at azure.archive.ubuntu.com and
nothing else, and the samba and pjdfstest images inherit that list. When Azure
is unreachable the build has nowhere to go: Acquire::Retries just retries a dead
host, every package fails, and apt exits 100 before a single test runs. Two
different workflows lost runs to it tonight.

Install through a helper that starts from the pristine sources.list each time
and walks a list of mirrors, so Azure stays the preferred one - the reason it
was pinned in the first place - without being the only one.

Verified both paths against a real build: the normal one installs from Azure,
and with the first entry pointed at an unroutable host the fallback logs the
skip and installs from archive.ubuntu.com.
2026-08-19 00:03:04 -07:00
Chris LuandGitHub da4f06ec12 Give the local Unix socket gRPC transport room to breathe (#10824)
* Give the local Unix socket gRPC transport room to breathe

Unix socket buffers default small and never autotune: 208KB on Linux, 8KB on
macOS. Once the buffer cannot absorb what gRPC's loopyWriter emits for the
in-flight streams the writer blocks on Write, and since v1.82.1 grpc-go counts
per-RPC bookkeeping toward its control-buffer throttle, so both peers stop
reading and the connection deadlocks for good. weed mini wedged at roughly 320
concurrent S3 PUTs with every filer RPC parked in waitOnHeader and no handler
running.

Force 8MB on both ends of the sockets we open. Best effort, since a kernel may
clamp it lower; that only lowers the concurrency this survives. TCP loopback
never hit this because its buffers start large and grow.

* Set the buffer on accepted connections too

Linux does not carry the listener's SO_SNDBUF onto sockets returned by accept,
so only the dialing half was getting the headroom: measured 8388608 on the
dialed side against the 212992 default on the accepted side. Wrap the listener
and re-apply per connection. macOS inherits either way, which is why this did
not show up locally.
2026-08-18 21:32:23 -07:00
Chris LuandGitHub bb223967bd mount: fold an inode's single link into its entry (#10818)
InodeEntry held its one path in a slice, so every inode the kernel references
cost a 16-byte backing array and a second heap object on top of the 32-byte
entry. The extra links of a hard-linked file now hang off a pointer instead,
which keeps the struct in the same 32-byte size class and leaves the ordinary
single-link file with nothing to allocate.

Populating the table with 1M children: 237.5 -> 221.5 B/inode at 85-character
paths, 301.3 -> 285.6 at 148.
2026-08-18 20:56:06 -07:00
Chris LuandGitHub 9f15e3935c mount: reuse the listed entry's path instead of rebuilding it (#10817)
readdir built dirPath.Child(name) for every child while entry.FullPath was
already that exact string, from NewFullPath in the meta cache store or from
FromPbEntry on the read-through path. One allocation per entry, and on a wide
tree with long paths that is most of what a listing allocates.

BenchmarkReadDirectory/kernel_readdirplus over 200k entries: 2,039,656 ->
1,839,318 allocs/op, 174.5 -> 167.8 MB/op, 152.6 -> 135.1 ms/op.
2026-08-18 20:55:17 -07:00
Chris LuandGitHub 887910b377 rust: honor fsync on the volume server write path (#10816)
The Rust volume server ignored the fsync parameter completely: nothing
parsed it, and write_volume_needle -> write_needle -> append_needle never
flushed. So a ?fsync=true upload was acked out of the page cache, and
since ReplicatedWrite forwards the parameter, a Go primary handing a
durable write to a Rust replica got the same empty promise.

The upload handler now reads fsync the way Go's r.FormValue does, off the
decoded query fields, and threads it down to the volume. A durable write
appends, flushes the .dat, publishes the needle map entry, then flushes
the .idx, and only then is it acked. Nothing points at bytes that are not
down yet, so a failed flush only has to take its own append back off the
end - the index never moved and the volume's counters never saw the
rejected write. If that truncate cannot be done the volume stops taking
writes, rather than letting a later append bury the rejected record
mid-file where the tail integrity check cannot see it.

The .idx flush is what keeps the ack honest: load() rebuilds the map from
.idx, so an acked write whose row was lost comes back as a .dat tail the
integrity check cannot account for, and the volume loads read only.

A dedup hit flushes too: there is nothing to append, but the write it
matched may have been non-durable, and the caller is asking for the
content to be on disk.

Batched writes carry the flag per request rather than one flush per
batch, so the write queue's module doc no longer claims otherwise.
2026-08-18 20:22:42 -07:00
Chris LuandGitHub 358fd314ea test(s3/versioning): read the whole version body instead of one Read (#10815)
A single Read on the response body can return the last bytes together
with io.EOF, so asserting NoError on it fails even though the body is
complete. Use io.ReadAll, like every other test in this package.
2026-08-18 18:34:01 -07:00
Lars LehtonenandGitHub ed75a61fb0 fix(test/s3/versioning): dropped test error (#10813) 2026-08-18 17:26:51 -07:00
9575032b4c volume: forward fsync=true to replicas in ReplicatedWrite (#10805)
* volume: forward fsync=true to replicas in ReplicatedWrite

When a write request carries fsync=true, only the primary volume server
flushed to disk: the replica fan-out URL in ReplicatedWrite only carried
type/ttl/ts/cm, so replicas always wrote without fsync even when the
client explicitly requested a durable write.

Forward the fsync request parameter to the replica volume servers so a
durable write means every replica has flushed to disk, not just the
primary. Replicas without fsync are untouched (zero behavior change).

* storage: flush a durable write inline while stopping

The fsync flag on the write path really selects the async batch worker,
and it was switched off once the store is stopping. So a fsync=true write
landing during the pre-stop drain got acked without ever being flushed -
and now that ReplicatedWrite forwards fsync, that covers replicas too.

Flush it inline instead of queueing it. The drain keeps accepting writes,
which is the whole point of preStopSeconds, and the ack still means the
.dat is on disk. If the fsync fails, the append comes back off the .dat
and the needle map goes back to what it pointed at before, so nothing
resolves to an offset past the truncated end.

* storage: make the store's stopping flag atomic

SetStopping runs on the signal handler goroutine while the write and
vacuum paths read the flag, so every read of it was racy. Nothing about
the shutdown ordering changes; only the flag itself is now safe to read.

* topology: check the errors the replication test was dropping

The mock replica ignored its response write and the mock master ignored
whatever Serve returned, so a broken mock would have shown up as a
confusing timeout rather than a failure. Also drops the explicit listener
close: grpc.Server.Stop already closes the listener it was given.

---------

Co-authored-by: hzsunchao <hzsunchao@corp.netease.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-18 17:17:50 -07:00
Chris LuandGitHub 1354b58675 s3: stop unrouted bucket subresources from being answered with a listing (#10814)
* s3: answer GetBucketReplication, GetBucketWebsite and GetBucketNotificationConfiguration

None of the three had a route, so they reached the unconstrained ListObjectsV1
catch-all and a client asking for a bucket's replication config got 200 and a
<ListBucketResult> back. Replication and website report their configuration as
absent the way AWS does; notification returns the empty configuration AWS
returns for a bucket with no events wired up.

* s3: stop an unrouted bucket subresource from being answered with a listing

ListObjectsV1 is the catch-all GET on a bucket, so every subresource without a
route of its own - ?torrent today, whatever AWS adds next - came back 200 with a
<ListBucketResult>. A client that asked for a configuration and got a listing
either fails its XML decode in a way that reads like corruption, or worse,
tolerantly parses it. Refuse the request instead.

The allow-list is the ListObjects parameters rather than the subresources,
so a new one fails closed. Presigned URLs sign their credentials into the
query string, so X-Amz-* and the SigV2 trio have to stay listable.
2026-08-18 16:03:08 -07:00
68a23a4b3c filer: stop remote.unmount from deleting the remote objects (#10811)
* filer: add filer.options.disable_remote_storage_deletion for cache-only deletes

Deleting a filer entry under a remote.mount path also deletes the backing
object from the remote store (maybeDeleteFromRemote). Deployments that use a
remote mount as a read-through cache in front of an authoritative,
externally-managed object store cannot allow this: the filer typically holds
read-only credentials, so the remote delete fails and the entire delete
errors out; and even where it would succeed, it destroys data the filer does
not own.

Add filer.options.disable_remote_storage_deletion (default false, so existing
behaviour is unchanged). When enabled, maybeDeleteFromRemote is skipped for
both single-entry and recursive folder deletes: local metadata and cached
chunks are still removed, but the remote object is left intact.

* filer: assert local removal in cache-only recursive delete test

The recursive cache-only delete test only checked that no remote delete
happened; it did not verify the local child and directory entries were
removed. Add FindEntry assertions so a regression that skips local
recursive deletion is caught.

* filer: reload the remote mount mapping when /etc/remote changes

The mapping was only read at startup, so remote.unmount left the mount live
in the filer: the purge that follows the mapping delete then went to the
remote store and wiped every object under the mount.

Rebuild the rules trie and the conf map from scratch on each load, since
ptrie cannot drop a key, and swap them under a lock.

* filer: drop the filer-wide remote deletion switch

With the mapping reloaded on unmount, the purge no longer reaches the remote
store, so there is nothing left for the switch to protect against.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-18 15:38:02 -07:00
Chris LuandGitHub f41595fb10 mount: drop consumed entries when reading a directory through (#10802)
The cached readdir trims the head of the handle's entry stream as the
client walks past it; the read-through path never did, so a directory
too large to cache -- the only kind that takes that path -- was held
whole in the handle for the length of the walk. Hoist the trim to cover
both paths.
2026-08-17 21:22:03 -07:00
Chris LuandGitHub 3cf7d306a5 Give the WebDav chunk reader a bounded, invalidatable location cache (#10801)
* mount: re-resolve volume locations after a failed chunk read

NewChunkGroup passed nil as the ReaderCache's CacheInvalidator, so
retryFetchAfterCacheInvalidation was dead code on the FUSE read path. A
mount that cached a volume's locations while one server was down kept
retrying that server after it died, then returned EIO, even though the
master and filer both resolved the live replica. The S3 gateway already
passes its filerClient; do the same for the mount.

* test: FUSE integration tests for volume server failover

One mount appends while a second tails, and a volume server is killed,
started or restarted mid-stream against a 001-replicated cluster of three
volume servers. Automates the scenario matrix reported for Docker Swarm
mounts, including the large-file variant and a no-chaos control.

* test: report the filer's own view when append content mismatches

A mismatch between what the writer wrote and what the reader sees can come
from either side's cache. Read the file back through the filer's HTTP
handler as well, and let the mount verbosity be raised from the
environment, so a failing run says which layer lost the data.

* test: wait for the reader mount to converge before comparing

A mount caches metadata for about a second, so reading the file the instant
the writer's last close returned can legitimately come back short. Poll the
reader until it matches or the timeout expires; content that is wrong rather
than merely late never converges and still fails, now with the writer's
mount and the filer's own view alongside it.

* test: detect a failover cluster child that exited at startup

Signal(0) succeeds for a zombie and nothing reaped these children until
shutdown, so a process that died on startup looked alive until the readiness
timeout expired. Reap each child as it is started and consult the result.

* test: read a file the killed volume server actually holds

Placement decides which two of three servers back each volume, so killing
volume N and reading readfile-N could pass without the victim ever holding a
replica of it. Resolve each file's volumes through the filer and the master,
and pick one the victim backs, preferring a file the reader has not cached.

* ci: stop persisting checkout credentials in the failover workflow

The job does not use the token after cloning. Also tag the README's command
block as bash and match the timeout the workflow actually uses.

* test: discard the ignored errors errcheck flags in the failover harness

* test: resolve manifests when mapping a file to its volumes

A manifest chunk's own fid names the volume holding the manifest, not the
volumes holding the data, so a large enough file would point the failover
victim at the wrong server.

* test: pin the stale-location recovery path with a primed reader

Reading a file for the first time after a server dies proves nothing: the
lookup is fresh and returns the survivor. Kill one holder and wait for the
master to drop it, read a file on that volume so the reader caches the lone
survivor, restart the first server, then kill the survivor. The reader's only
cached location is now dead while the data is live elsewhere, which is the
case the invalidator exists for: EIO without it, recovery with it.

* filer: re-look-up a chunk's locations as soon as they all fail

A read that fails against every location it was given is far more likely to be
holding a stale list than to be hitting a cluster that is briefly slow, but the
retry loops spent the whole backoff ladder, about 13 s, before the caller got a
chance to invalidate and look the chunk up again. Give the loops a refresh hook
and let the reader cache invalidate on the first fully failed pass, so recovery
starts in milliseconds. Clients without an invalidator keep the old behavior.

The filer's streaming read path has its own fetch loop and is not covered.

* webdav: give the chunk reader a bounded, invalidatable location cache

WebDav resolved chunk locations through filer.LookupFn, whose own doc asks
long-running processes to prefer wdclient.FilerClient: its cache is unbounded,
and it has no way to invalidate an entry, so the reader cache was constructed
with a nil invalidator and a WebDav server that had cached a location kept
reading from it after the volume moved or died. Use FilerClient, as the mount
and the S3 gateway already do.

* filer: refresh locations on the random-read path too

readChunkSliceAt bypasses the chunk cacher in random-access mode and fetches
the range directly, which left it without the invalidation the cacher does:
a random reader parked on a stale location had no way back at all. Hoist the
refresh hook onto the reader cache so both paths share it.

* filer: compare chunk locations as a set, not in order

Lookups shuffle the locations they return, so comparing positionally reads a
reshuffle of the very same replicas as a fresh set and spends an immediate
retry on locations that just failed. weed/filer already had an
order-independent comparison for this; move it next to the retry loops so
both callers share one helper.
2026-08-17 20:19:54 -07:00
Chris LuandGitHub f3dc530919 Re-look-up a chunk's locations as soon as they all fail (#10800)
* mount: re-resolve volume locations after a failed chunk read

NewChunkGroup passed nil as the ReaderCache's CacheInvalidator, so
retryFetchAfterCacheInvalidation was dead code on the FUSE read path. A
mount that cached a volume's locations while one server was down kept
retrying that server after it died, then returned EIO, even though the
master and filer both resolved the live replica. The S3 gateway already
passes its filerClient; do the same for the mount.

* test: FUSE integration tests for volume server failover

One mount appends while a second tails, and a volume server is killed,
started or restarted mid-stream against a 001-replicated cluster of three
volume servers. Automates the scenario matrix reported for Docker Swarm
mounts, including the large-file variant and a no-chaos control.

* test: report the filer's own view when append content mismatches

A mismatch between what the writer wrote and what the reader sees can come
from either side's cache. Read the file back through the filer's HTTP
handler as well, and let the mount verbosity be raised from the
environment, so a failing run says which layer lost the data.

* test: wait for the reader mount to converge before comparing

A mount caches metadata for about a second, so reading the file the instant
the writer's last close returned can legitimately come back short. Poll the
reader until it matches or the timeout expires; content that is wrong rather
than merely late never converges and still fails, now with the writer's
mount and the filer's own view alongside it.

* test: detect a failover cluster child that exited at startup

Signal(0) succeeds for a zombie and nothing reaped these children until
shutdown, so a process that died on startup looked alive until the readiness
timeout expired. Reap each child as it is started and consult the result.

* test: read a file the killed volume server actually holds

Placement decides which two of three servers back each volume, so killing
volume N and reading readfile-N could pass without the victim ever holding a
replica of it. Resolve each file's volumes through the filer and the master,
and pick one the victim backs, preferring a file the reader has not cached.

* ci: stop persisting checkout credentials in the failover workflow

The job does not use the token after cloning. Also tag the README's command
block as bash and match the timeout the workflow actually uses.

* test: discard the ignored errors errcheck flags in the failover harness

* test: resolve manifests when mapping a file to its volumes

A manifest chunk's own fid names the volume holding the manifest, not the
volumes holding the data, so a large enough file would point the failover
victim at the wrong server.

* test: pin the stale-location recovery path with a primed reader

Reading a file for the first time after a server dies proves nothing: the
lookup is fresh and returns the survivor. Kill one holder and wait for the
master to drop it, read a file on that volume so the reader caches the lone
survivor, restart the first server, then kill the survivor. The reader's only
cached location is now dead while the data is live elsewhere, which is the
case the invalidator exists for: EIO without it, recovery with it.

* filer: re-look-up a chunk's locations as soon as they all fail

A read that fails against every location it was given is far more likely to be
holding a stale list than to be hitting a cluster that is briefly slow, but the
retry loops spent the whole backoff ladder, about 13 s, before the caller got a
chance to invalidate and look the chunk up again. Give the loops a refresh hook
and let the reader cache invalidate on the first fully failed pass, so recovery
starts in milliseconds. Clients without an invalidator keep the old behavior.

The filer's streaming read path has its own fetch loop and is not covered.

* filer: refresh locations on the random-read path too

readChunkSliceAt bypasses the chunk cacher in random-access mode and fetches
the range directly, which left it without the invalidation the cacher does:
a random reader parked on a stale location had no way back at all. Hoist the
refresh hook onto the reader cache so both paths share it.

* filer: compare chunk locations as a set, not in order

Lookups shuffle the locations they return, so comparing positionally reads a
reshuffle of the very same replicas as a fresh set and spends an immediate
retry on locations that just failed. weed/filer already had an
order-independent comparison for this; move it next to the retry loops so
both callers share one helper.
2026-08-17 20:19:28 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>Chris Lu
93227c6dc3 build(deps): bump go.etcd.io/etcd/client/v3 from 3.6.12 to 3.7.1 (#10789)
Bumps [go.etcd.io/etcd/client/v3](https://github.com/etcd-io/etcd) from 3.6.12 to 3.7.1.
- [Commits](https://github.com/etcd-io/etcd/compare/v3.6.12...v3.7.1)

---
updated-dependencies:
- dependency-name: go.etcd.io/etcd/client/v3
  dependency-version: 3.7.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-08-17 19:57:31 -07:00
Chris LuandGitHub 606a90b3b1 filer: close the empty-folder race by checking after each mutation (#10799)
* filer: re-list a folder after deleting it, and put it back if it is not empty

The emptiness check inside the delete and the removal of the folder entry are
not atomic, so an entry can land between them and be left reachable by its own
path but out of every listing. Looking again after the delete catches the ones
whose create event has not arrived yet, and does not depend on the event stream
or on the observation window holding.

* filer: create the directories holding an entry after the entry

A parent checked before the insert can be taken by the empty-folder cleaner
before the entry lands, which leaves the entry reachable by its own path but out
of every listing. Creating the parents afterwards cannot be undone by a delete
that was authorised before the insert, and pairs with the cleaner re-listing
after its own delete: whichever of the two acts second sees what the other did.

Going second means the entry is already stored when the parent fails, so it is
taken back out and the caller still sees the error it used to get.

* filer: narrow a directory that came back wider than the one it replaced

A writer recreating its own missing parent has only the entry it is inserting to
go on, so the directory it mints can grant access the deleted one denied - a
0700 folder comes back 0751. The cleaner read the real attributes before
deleting, so its restore now puts the original mode back instead of leaving the
inferred one in place. It only ever narrows, so a directory deliberately
tightened since is left as it is.
2026-08-17 19:57:05 -07:00
Chris LuandGitHub 1ddec72707 Recover from a dead volume server on the mount read path (#10798)
* mount: re-resolve volume locations after a failed chunk read

NewChunkGroup passed nil as the ReaderCache's CacheInvalidator, so
retryFetchAfterCacheInvalidation was dead code on the FUSE read path. A
mount that cached a volume's locations while one server was down kept
retrying that server after it died, then returned EIO, even though the
master and filer both resolved the live replica. The S3 gateway already
passes its filerClient; do the same for the mount.

* test: FUSE integration tests for volume server failover

One mount appends while a second tails, and a volume server is killed,
started or restarted mid-stream against a 001-replicated cluster of three
volume servers. Automates the scenario matrix reported for Docker Swarm
mounts, including the large-file variant and a no-chaos control.

* test: report the filer's own view when append content mismatches

A mismatch between what the writer wrote and what the reader sees can come
from either side's cache. Read the file back through the filer's HTTP
handler as well, and let the mount verbosity be raised from the
environment, so a failing run says which layer lost the data.

* test: wait for the reader mount to converge before comparing

A mount caches metadata for about a second, so reading the file the instant
the writer's last close returned can legitimately come back short. Poll the
reader until it matches or the timeout expires; content that is wrong rather
than merely late never converges and still fails, now with the writer's
mount and the filer's own view alongside it.

* test: detect a failover cluster child that exited at startup

Signal(0) succeeds for a zombie and nothing reaped these children until
shutdown, so a process that died on startup looked alive until the readiness
timeout expired. Reap each child as it is started and consult the result.

* test: read a file the killed volume server actually holds

Placement decides which two of three servers back each volume, so killing
volume N and reading readfile-N could pass without the victim ever holding a
replica of it. Resolve each file's volumes through the filer and the master,
and pick one the victim backs, preferring a file the reader has not cached.

* ci: stop persisting checkout credentials in the failover workflow

The job does not use the token after cloning. Also tag the README's command
block as bash and match the timeout the workflow actually uses.

* test: discard the ignored errors errcheck flags in the failover harness

* test: resolve manifests when mapping a file to its volumes

A manifest chunk's own fid names the volume holding the manifest, not the
volumes holding the data, so a large enough file would point the failover
victim at the wrong server.

* test: pin the stale-location recovery path with a primed reader

Reading a file for the first time after a server dies proves nothing: the
lookup is fresh and returns the survivor. Kill one holder and wait for the
master to drop it, read a file on that volume so the reader caches the lone
survivor, restart the first server, then kill the survivor. The reader's only
cached location is now dead while the data is live elsewhere, which is the
case the invalidator exists for: EIO without it, recovery with it.
2026-08-17 17:30:33 -07:00
Chris LuandGitHub 6fda8c67f3 Guard the gcs credential path in FetchAndWriteNeedle like the other backends (#10796)
* volume: accept only static-key gcs credentials on the fetch request

An inline credentials document of a federated type points the SDK at a url,
file or executable of the caller's choosing for the token exchange, so the
request-supplied value is no longer just a key.

* volume: guard the gcs token endpoint like the other remote endpoints

Inline credentials pick where the token request goes, so route the gcs client
through the same deny-list and rebinding-safe dialer used for S3 and azure.

* rust volume: pin that gcs has no credential-driven dial path

* volume: only check gcs credentials on a gcs remote conf

Only the gcs backend reads that field, so another backend carrying a stale
value should not fail the request.

* gcs: load credentials with the type the caller expects

The untyped loader is deprecated because it reads whatever the document
claims to be; callers handling credentials they do not control now name the
types they accept.
2026-08-17 16:40:56 -07:00
Chris Lu 9d8acbd244 Create icon.svg 2026-08-17 16:11:27 -07:00
Chris LuandGitHub 1bcd55eba2 go 1.26 (#10797) 2026-08-17 15:39:20 -07:00
Chris LuandGitHub e383ee47cb filer: use bind variables for request-controlled values in the arangodb store (#10795)
* arangodb: bind list prefix, start file name and collection into the AQL query

Concatenating them into the query text let a caller-supplied prefix or
start name close the string literal and append arbitrary AQL, which runs
with the filer's ArangoDB credentials against any collection.

* arangodb: bind the folder path and collection into the recursive delete query

A trailing-slash S3 key reaches DeleteFolderChildren through the
directory-marker cleanup, so quotes in the path could turn the filter
into a match-everything REMOVE over the whole bucket collection.

* arangodb: match the real directory prefix in the recursive delete

The prefix was built by re-joining the path segments with commas, so it
never matched a stored directory and the subtree sweep did nothing.
2026-08-17 15:15:26 -07:00
Chris LuandGitHub 5d5ea63b3f Fix what the Go 1.26 language bump breaks (#10794)
* worker: log the balance move stage through a constant format string

Go 1.26's printf analyzer now follows printf wrappers reached through an
interface, so passing the stage straight to Logger.Info is a vet failure.

* s3api: bracket the IPv6 host in the signature test URL

A bare IPv6 literal is legal in a Host header but never in a URL. Go 1.26
stopped parsing it leniently, so carry the two forms separately and set
r.Host to the value the client would actually have signed.

* mini: bracket IPv6 addresses in the readiness probe URLs

An IPv6-only host hands mini a bare literal, and %s:%d pasted it into a URL
unbracketed. Under Go 1.26 that URL no longer parses, so waiting for the
admin server never succeeds and mini refuses to start.
2026-08-17 14:46:25 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
518da712b5 build(deps): bump github.com/rabbitmq/amqp091-go from 1.11.0 to 1.13.0 (#10791)
Bumps [github.com/rabbitmq/amqp091-go](https://github.com/rabbitmq/amqp091-go) from 1.11.0 to 1.13.0.
- [Changelog](https://github.com/rabbitmq/amqp091-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/rabbitmq/amqp091-go/compare/v1.11.0...v1.13.0)

---
updated-dependencies:
- dependency-name: github.com/rabbitmq/amqp091-go
  dependency-version: 1.13.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-17 14:44:06 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d36073f8d7 build(deps): bump golang.org/x/net from 0.57.0 to 0.58.0 (#10790)
Bumps [golang.org/x/net](https://github.com/golang/net) from 0.57.0 to 0.58.0.
- [Commits](https://github.com/golang/net/compare/v0.57.0...v0.58.0)

---
updated-dependencies:
- dependency-name: golang.org/x/net
  dependency-version: 0.58.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-17 14:43:39 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
42f2a48c90 build(deps): bump google.golang.org/api from 0.289.0 to 0.293.0 (#10792)
Bumps [google.golang.org/api](https://github.com/googleapis/google-api-go-client) from 0.289.0 to 0.293.0.
- [Changelog](https://github.com/googleapis/google-api-go-client/blob/main/CHANGES.md)
- [Commits](https://github.com/googleapis/google-api-go-client/compare/v0.289.0...v0.293.0)

---
updated-dependencies:
- dependency-name: google.golang.org/api
  dependency-version: 0.293.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-17 13:15:45 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
42699b0f98 build(deps): bump github.com/aws/aws-sdk-go-v2 from 1.43.4 to 1.43.5 (#10787)
Bumps [github.com/aws/aws-sdk-go-v2](https://github.com/aws/aws-sdk-go-v2) from 1.43.4 to 1.43.5.
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/v1.43.4...v1.43.5)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2
  dependency-version: 1.43.5
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-17 10:17:00 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
52ebc9ea4e build(deps): bump golang.org/x/crypto from 0.54.0 to 0.55.0 (#10786)
Bumps [golang.org/x/crypto](https://github.com/golang/crypto) from 0.54.0 to 0.55.0.
- [Commits](https://github.com/golang/crypto/compare/v0.54.0...v0.55.0)

---
updated-dependencies:
- dependency-name: golang.org/x/crypto
  dependency-version: 0.55.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-17 10:16:46 -07:00
Chris LuandGitHub f4bcec60d7 readme: fold RustFS into the MinIO comparison (#10788)
* readme: add RustFS to the file system comparison

* readme: note RustFS write amplification and rigid layout

* readme: correct RustFS version, parity and protocol details

* readme: merge the RustFS comparison into the MinIO section
2026-08-17 10:11:48 -07:00
github-actions[bot] f04da8e9ad 4.42 2026-08-17 07:15:32 +00:00
Chris LuandGitHub 5c43c03b76 filer: restore a folder that received an entry while it was deleted (#10783)
* filer: restore a folder that received an entry while it was deleted

The empty-folder cleaner checks that a folder is empty and then deletes it,
and those two steps are not atomic. An entry created in between survives the
delete but loses the directory holding it: still readable by its own path, yet
absent from every listing until a later write happens to recreate the parent.

Record the folders deleted in each pass and re-check them on the next one,
putting back any that turned out to hold entries. The check waits a pass on
purpose - a writer looks up the parent before inserting the child, so checking
straight after the delete can still run ahead of the insert and see nothing.

Restoring a directory that holds entries is always correct, and restoring one
whose entry went away again just leaves an empty folder for a later pass to
collect, so the repair needs no locking or coordination.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* filer: keep failed restores queued and inherit the ancestor's ownership

Two gaps in the restore pass.

A folder whose count or restore hit a transient store error was dropped from
the tracking list and never looked at again, leaving its entries out of
listings until some later write recreated the folder - the very thing the pass
exists to avoid. Put those back for the next pass, still under the cap.

A restored folder was minted with a fixed mode and no owner, so a directory
that had been private came back world-readable and owned by root. Take the
mode and ownership from the nearest ancestor still present instead.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* filer: let the redis stores keep a directory listing that still has entries

On the redis stores the listing is not derived from the entries, it is the only
record that they sit under that directory. DeleteEntry opened by dropping it
outright, so an entry that arrived after the caller judged the directory empty
lost its membership and became unreachable: readable by exact path, absent from
every listing, and invisible to any later check, since counting the directory
reads the listing that was just destroyed. Nothing could detect or repair it.

Drop the listing in DeleteFolderChildren instead, alongside the children it
describes, and leave it alone in DeleteEntry. redis3 needs it explicitly, since
removeChildren clears the skip list nodes but not the list itself, and the plain
redis store was leaking the key entirely.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* filer: restore folders with their own attributes, and observe them for a window

Five gaps in the restore pass.

The restored directory was reconstructed from whatever ancestor happened to
still be present, and the mode was ORed with 0111 on the way. A private
directory under a world-traversable parent came back granting traversal it had
denied. Read the folder's own attributes before deleting it and put exactly
those back. That also removes the ancestor walk, which treated a transient
store error as "not found" and silently fell through to a broader ancestor.

A single check a pass later was not a delay at all. Ticker sends coalesce, so
when a pass runs long the next one starts immediately, and a writer already
past its parent lookup can insert after the check has read zero - after which
the folder was discarded for good. Keep each folder under observation for a
bounded wall-clock window and re-check it on every pass until it expires. This
narrows the exposure rather than closing it; only making the emptiness check
and the delete atomic would do that.

A delete that returned an error was never observed at all, though the redis
stores drop the folder before its parent-list member, so a failure return is
not proof the folder survived. Record the folder before the delete instead.

Restores now run shallowest first, so a folder taken by the parent cascade is
rebuilt with its own attributes before anything below it needs it as a parent.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* filer: recover a deleted folder from the create event for the entry that raced it

Checking each deleted folder on a timer was the wrong instrument. It cost a
listing per folder per pass, and it could only ever be a guess about when the
racing write would land.

The metadata stream already carries the answer. A folder is recorded before it
is deleted, so any entry that can be orphaned is created after that record and
its create event names that exact directory. Match the event against the
recently deleted folders and the folder is known to need putting back, rather
than inferred to.

The window stops being a guess at the race and becomes what it should be: how
far behind the event stream is allowed to run before a folder stops being
watched. Listing is now done once, for a folder an event has already named, to
skip the restore when the entry has since gone away again.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* filer: bound how long a folder is watched, and rebuild ancestors from themselves

Four gaps found reviewing the restore pass.

A folder whose restore kept failing was never let go: the written-to check ran
before the age check, so it was picked up, retried, put back, and counted again
on every pass for the life of the process. Apply the window first, whatever
state the folder is in.

At the cap, the folder being recorded was the one turned away, though it is the
one whose race is still live - the older entries are already close to ageing
out. Give up one of those instead, picked as the oldest of a small sample so
the cost stays flat under heavy deletion rates.

An ancestor taken by the same cascade was left to the descendant's restore to
recreate, which minted it from the descendant's attributes and handed back
access the ancestor never granted. Rebuild those from what they were, ahead of
anything below them.

Reading a directory's attributes assumed an entry came back. Some stores return
nothing with no error, so treat that as not found. The mode is also taken whole
rather than through Perm(), which was dropping setgid, setuid and sticky.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* redis3: take a directory listing left behind by a failed delete

Removing the last name deletes the list, and if that delete fails the header
survives pointing at a name that is gone. The retry finds nothing to remove,
reports no changes, and returns before reaching the delete, so the key stays
for good. Take it on that path too.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r
2026-08-17 00:04:07 -07:00
Chris LuandGitHub f530102c45 filer: do not sweep children when deleting a folder non-recursively (#10782)
* filer: do not sweep children when deleting a folder non-recursively

doBatchDeleteFolderMetaAndData lists a folder and bails out if it has any
children, then calls Store.DeleteFolderChildren unconditionally. On the
non-recursive path that bulk sweep has nothing legitimate to remove: it only
runs once the listing came back empty, so the sole rows it can delete are
ones inserted after the check.

The S3 empty-folder cleaner deletes through this path, so a PUT landing
between the listing and the sweep loses its entry after the write was already
acknowledged. Neither side sees an error - the client has its 200 and the
cleaner logs an ordinary empty-folder deletion - and the chunks leak, since
the cleaner passes shouldDeleteChunks=false and nothing was enumerated to
collect. Workloads that scatter objects over many shallow prefixes empty and
refill those folders constantly, which is what makes the window reachable.

Sweep only when the delete is recursive, or when the whole-bucket shortcut
skipped the listing and depends on it.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r

* filer: pin the folder entry removal left by the racing-child test

The surviving entry is reachable by path but drops out of listings until the
folder comes back, and nothing in the test said so. Assert it, so the exposure
that remains after this change is visible rather than implied.

Claude-Session: https://claude.ai/code/session_01HdLXMUopwgofPb1ZEmiE6r
2026-08-16 22:09:23 -07:00
Chris LuandGitHub 7522e17b6d iceberg: vend table-scoped credentials to clients that ask for delegation (#10777)
* iceberg: vend table-scoped credentials to clients that ask for delegation

The catalog recognised X-Iceberg-Access-Delegation: vended-credentials
and then deliberately said nothing, because it had nothing to vend: it
withheld even the S3 endpoint so the client would keep the credentials it
was configured with. That left every engine expecting the catalog to hand
out access - Snowflake, Databricks, Trino with vending, any multi-tenant
setup - needing static S3 keys distributed out of band.

Mint an STS session per request instead, scoped by a session policy to
the table's own prefix plus the bucket listing needed to resolve it, and
return it in the load response config and storage-credentials. The role
to assume is named by -s3.iceberg.credentialRole; its trust policy is
what decides whether a caller may assume it, and vending stays off until
it is set. A failed mint falls back to the old silence rather than
handing back an endpoint the client cannot sign for.

* iceberg: keep vended credentials inside the table prefix

Review follow-ups on credential vending:

Listing was granted on the bucket ARN with no condition, so a credential
vended for one table could enumerate every other table's object names.
Constrain s3:prefix to the table's own prefix, which the S3 gateway
already populates for list requests.

A table location carrying * or ? would have gone into the policy's
resource pattern unescaped and widened the session to sibling prefixes.
Refuse to vend for such a location rather than escaping it; nothing the
catalog generates contains those characters.

DurationSeconds skipped the 900..43200 bounds the other assume-role paths
enforce, so -s3.iceberg.credentialDurationSeconds could ask for a session
outside them. The check is now shared by all three entry points.

* iceberg: return the vended credentials from buildFileIOConfig itself

buildStorageConfig was a second name for what buildFileIOConfig already
did; it now returns the storage credentials alongside the properties, and
callers that only want the properties drop them.

* iceberg: split the vended bucket grants, and refuse a whole-bucket scope

The prefix condition sat on a statement that also granted
GetBucketLocation and ListBucketMultipartUploads, neither of which carries
an s3:prefix to satisfy it, so both were denied for every vended
credential. GetBucketLocation moves to its own unconditioned statement.
ListBucketMultipartUploads is dropped: Iceberg writers complete and abort
by upload id, and granting it either leaks in-flight keys bucket-wide or
breaks on the same missing prefix.

A table whose location has no prefix - one registered at the bucket root -
would have been vended read and write over every other table in the
bucket. Refuse, the way a location with wildcards is refused.
2026-08-16 12:57:12 -07:00
Chris LuandGitHub ec37ef5aaa iceberg: add view rename, scan-report and snapshots=refs to the catalog (#10776)
* iceberg: add view rename, scan-report and snapshots=refs to the catalog

Three gaps against the REST spec that clients hit in normal use:

Views had no rename, though tables did and views are stored the same way,
so the move is the same catalog-only pointer move. Tables and views share
a namespace directory, so both renames now refuse the other kind instead
of moving it.

Engines POST a scan or commit report after planning; a 404 there turns
into an error line per query. Accept the report and discard it - the
catalog keeps no metrics store.

LoadTable ignored ?snapshots=refs and always returned the whole snapshot
history, which is what clients use the parameter to avoid on long-lived
tables.

* iceberg: authorize view rename against the view ARN, tighten the metrics endpoint

Review follow-ups:

The shared rename checked the source against a table ARN whatever the
kind, so a policy scoped to a view's own ARN never matched and one
written for a table ARN was evaluated for a view. The entry kind now
carries the ARN builder.

The metrics endpoint truncated a report at 1 MiB and then failed to parse
it, answering 400 for a query that had actually succeeded. Read one byte
past the limit to tell "fits" from "cut short", and discard an oversized
report instead of rejecting it. Empty bodies and reports without a
report-type are now rejected, which the REST schema requires.

?snapshots= is defined for LoadTable, so it no longer filters what
CreateTable echoes back.
2026-08-16 12:56:45 -07:00
Chris LuandGitHub d044839ab2 iceberg: make a table commit a compare-and-swap (#10775)
* iceberg: make a table commit a compare-and-swap

The catalog validated the caller's version token, ran its authorization
checks, and only then wrote the new metadata xattr. Two engines
committing against the same base both passed that check and both wrote,
so the second silently dropped the first one's snapshot. Both also derive
the same v{N}.metadata.json name and the file write overwrote, leaving
the surviving pointer aimed at the loser's metadata - and the loser's
conflict cleanup then deleted the winner's file.

Write the metadata file with an exclusive create and update the xattr
conditionally on the bytes the handler read, the way the maintenance
worker already commits. A writer that lost the race re-reads and retries,
and reports 409 CommitFailedException once out of attempts.

* iceberg: stage a commit under a unique name when the versioned one is taken

Two follow-ups from review of the commit compare-and-swap:

Refusing to overwrite v{N}.metadata.json also refused to get past a file
left behind by a commit that died between staging and updating the
pointer. Every later commit derived the same name, saw the collision, and
reported a conflict, so the table stayed uncommittable until an orphan
sweep removed the file. Stage under v{N}-{uuid} instead: neither writer's
file is overwritten and the catalog pointer still decides who won, which
is how the maintenance worker has always staged its own metadata.
metadataVersionFromLocation learned to read the version back out of that
name.

The conditional update guarded only the metadata attribute while the
write replaced the whole entry, so a policy or tag written in the same
window was silently reverted. Guard every catalog attribute, which turns
that into a conflict the caller retries on fresh state.

* iceberg: give saveMetadataFile the exclusive flag instead of a second name

saveNewMetadataFile, saveMetadataBlobExclusive and uniqueMetadataFileName
were three new names around one existing helper. The flag now rides on
saveMetadataFile and saveMetadataBlob, and the unique-name construction
sits where it is used.

* iceberg: reuse the filer CAS helpers #10773 added, and stage transactions exclusively

#10773 landed mutateEntryExtended, which already writes an entry back under a
whole-entry precondition and retries. Drop the helper this branch added and
route the table commit through it: the check that the metadata is still the
one this request read now lives in the mutation, where it sees current state.

The policy the request was authorized against is asserted too, so an
administrator restricting it mid-commit sends the caller back through
authorization instead of having a stale decision applied. Bucket and
namespace policies live on other entries and a single-entry precondition
cannot cover them.

Multi-table transactions stage their metadata exclusively for the same
reason single-table commits do, and carry the name they landed on into the
pointer flip.
2026-08-16 12:55:42 -07:00
Chris LuandGitHub a80259d362 iceberg maintenance: fix the test build master merged broken (#10780)
#10774 gave buildTestMetadata its refs and age parameters while #10773
added a caller with the old arity. Each was green against a master that
did not yet have the other, and the merge of both does not compile, so
vet and the unit tests fail on master.
2026-08-16 12:06:43 -07:00
Chris LuandGitHub 5f6dd4d3e5 iceberg maintenance: keep the snapshots that branches and tags pin (#10774)
* iceberg maintenance: keep the snapshots that branches and tags pin

expireSnapshots only ever protected the current snapshot, so a snapshot
held by a tag or a non-main branch was expired once it aged out of the
retention window. iceberg-go's RemoveSnapshots drops any ref whose
snapshot is gone without complaint, so the tag disappeared and the files
behind it were deleted as unreferenced.

Protect every ref target, and honour a branch's own
min-snapshots-to-keep / max-snapshot-age-ms over the ancestors behind its
head. Detection skips pinned snapshots for the same reason: proposing a
job whose only outcome is a no-op keeps the worker busy forever.

* iceberg maintenance: re-plan when a ref appears mid-commit, and stop proposing no-op expiry

Three follow-ups from review of the ref-aware expiry:

The commit guard only compared the table head, so a tag created between
planning and commit could pin a snapshot the plan was about to expire.
Re-check the refs against the metadata the commit actually reads.

Detection now asks snapshotsToExpire what execution would remove instead
of approximating with its own count-and-age rules. Expiry always requires
a snapshot past the retention window, so a table over the quota whose
snapshots are all young was being proposed for a job that could only
no-op.

The branch retention test could not tell "retained the whole lineage"
from "honoured min-snapshots-to-keep", because the branch had exactly as
many ancestors as the count. Give it one more, and cover
max-snapshot-age-ms too. Both need snapshots genuinely older than a
retention window, which iceberg-go will not accept at build time, so the
fixture backdates the metadata after building it.

* iceberg maintenance: fold the metadata test builders back into one

buildTestMetadata, buildTestMetadataWithRefs, buildTestMetadataAged and
buildTestMetadataNow were four names for one thing. Keep the original and
give it the refs and age it needs.
2026-08-16 10:48:19 -07:00
Chris LuandGitHub eef6f3d1e6 s3tables: add the maintenance configuration APIs (#10773)
* s3tables: add the maintenance configuration APIs

Stores the configuration verbatim as the wire shape under a new
s3tables.maintenance extended attribute, so Get hands back what Put took
and no translation layer can drift from the AWS model.

Nothing reads the configuration yet.

Put merges a single type into the stored map so configuring compaction
does not drop snapshot management, and asserts the attribute's prior value
so two concurrent Puts cannot silently clobber each other.

* iceberg: apply the maintenance configuration in the worker

The worker now reads the per-table and per-bucket maintenance
configuration written by the control plane, so the wildcard plugin config
is a default rather than the only setting a table can have.

Table properties still win by default, since a table declaring its own
layout is what every engine honours and the compactor has to agree with
whoever writes the files. Clearing table_properties_override makes the
maintenance configuration authoritative instead.

Status is not part of that contest: a disabled type drops its operations
and no property can re-enable them, so the operator's kill switch always
holds. Manifest and delete-file rewrites have no AWS equivalent and ride
with compaction.

Detection reads both attributes from entries it already lists.

* s3tables: report maintenance job status

The worker records the outcome of each run in its own extended attribute,
separate from the configuration so operator and worker writes do not
contend, and GetTableMaintenanceJobStatus reads it back.

Only the types a run touched are written, so a partial run cannot erase
what an earlier one recorded. The reader fills in the rest: Disabled when
the configuration switched a type off, Not_Yet_Run otherwise.

Status is advisory, so a lost race is logged rather than failing a job
whose work already committed.

* s3tables: route the maintenance APIs over REST

The five actions were only reachable by X-Amz-Target dispatch, which the
AWS CLI and SDK do not use for this service. They address the operations
by path, so the APIs were unreachable from any official client.

* s3tables: fix the table bucket ARN field name

GetTableBucketMaintenanceConfiguration emitted tableBucketArn where the
wire field is tableBucketARN, as every other response in this package
already spells it. Official SDK deserializers ignore the unknown key, so
the required field came back unset.

* s3tables: carry the compaction strategy through to the worker

IcebergCompactionSettings modelled only targetFileSizeMB, so a request
naming a strategy was accepted and then dropped on the way to storage.
The worker now maps binpack and sort onto its own rewrite strategy and
lets auto defer to the worker configuration.

z-order is rejected rather than accepted and quietly binpacked.

* s3tables: report bucket-level maintenance status

GetTableMaintenanceJobStatus read only the table's configuration, so
unreferenced file removal — which is configured on the bucket — reported
Not_Yet_Run or a stale success after an operator disabled it.

The merge helper now lives in this package and the worker shares it.

* iceberg: delete orphans only after the non-current window

AWS marks a file non-current once it has been unreferenced for
unreferencedDays, then deletes it a further nonCurrentDays later.
The cutoff was taken from unreferencedDays alone, so a 3/10 configuration
hard-deleted on day three and threw away the ten day recovery window.

remove_orphans deletes in one step rather than marking, so the cutoff is
now the sum of the two.

* s3tables: assert every attribute when rewriting an entry

UpdateEntry writes the whole entry back from the snapshot the caller
read, and its precondition only covers the keys the caller names. Both
maintenance writers named one key, so a job status write could revert a
maintenance configuration an operator had just disabled, turning an
advisory write into a silent re-enable.

Both now assert the entry's full attribute set, including the target key
when absent so a concurrent create also fails the precondition.

* s3tables: assert absent attributes when rewriting an entry

The precondition covered the attributes present when the writer read the
entry, so an attribute created between that read and the write was absent
from it. A first-time PutTableMaintenanceConfiguration disabling a type
therefore lands, passes the per-key checks, and is then deleted by the
stale whole-entry write.

Every attribute this package stores is now asserted, absent ones
included. The metadata commit and planning index writers rewrite the same
entries and had the same exposure, so both use the shared snapshot too.

* iceberg: implement the auto compaction strategy

auto was accepted, stored and read back, but left the worker on its own
default, so a sorted table configured as auto was compacted with binpack.

AWS defines auto as sorting tables that declare a sort order and
bin-packing the rest. That needs the table metadata, so the choice is made
where the rewrite plan is resolved: an unsorted table falls back to
binpack rather than failing the way an explicit sort request does.

* s3tables: validate the maintenance setting ranges

PUT accepted zero, negative and oversized values for every numeric
setting. The worker then ignores a non-positive value and saturates an
oversized one, so the configuration read back was not the one that ran.

AWS bounds all five to 1..2147483647, which is now enforced. The fields
are pointers so an explicit zero is distinguishable from an omitted one
and can be rejected rather than silently ignored.

* s3tables: give every entry writer the same compare-and-swap

updateExtendedAttribute asserted the entry's attributes, but the helpers
behind the metadata, policy and tag handlers still wrote the whole entry
unconditionally. Any of them could land on a stale snapshot and delete a
maintenance configuration an operator had just written.

They all share one read-modify-write loop now, so the precondition and
the bounded retry apply wherever an entry is rewritten.

* s3tables: move the maintenance configuration with a renamed table

RenameTable carried the metadata, version, policy and tags to the new
name but left the maintenance configuration and job status behind. A
table with snapshot management disabled came back enabled under its new
name, and the stale configuration stayed on the old name where a table
created there would inherit it.

The decoupled-delete cleanup left the same two attributes behind.

* s3tables: accept every AWS partition in ARNs

The route regexes and the ARN patterns both hardcoded arn:aws, so valid
aws-cn and aws-us-gov ARNs never reached a handler. The router now shares
the partition-tolerant prefix with the parser, and a generated ARN uses
the partition its region belongs to so it parses back.

* s3tables: generate ARNs in the region's partition

The handler's own ARN generators still formatted arn:aws directly rather
than going through the partition-aware builder, so a China or GovCloud
deployment routed the request but then returned a commercial ARN and
matched IAM policies against it.

The round-trip test missed this because parsing accepts any partition, so
it now asserts the prefix the region implies.

* s3tables: complete the ARN partition table

aws-iso-e, aws-iso-f and aws-eusc were missing, so eu-isoe-*, us-isof-*
and eusc-* regions fell through to the commercial partition.

* s3tables: do not let a rename swallow a concurrent maintenance write

Rename copied the source attributes early and cleared the source at the
end, so a Put landing in between missed the copy to the destination and
was then deleted by the cleanup. It succeeded and vanished.

The cleanup now clears the source only while it still holds exactly what
was copied, and returns a conflict otherwise. Put checks the catalog
identity inside the same conditional mutation, so it also cannot write to
a name that a rename or delete has already soft-deleted.
2026-08-16 10:36:59 -07:00
Chris LuandGitHub a1d3fe236f iceberg: let table properties override the worker config (#10772)
* iceberg: carry snapshot retention in milliseconds

Config stored retention as hours, so any sub-hour value would have to be
truncated to 0 and then clamped back up to the 168 hour default. Keep the
plugin config key in hours and convert once at parse time.

* iceberg: let table properties override the worker config

Every other Iceberg implementation lets a table's own properties win over
engine defaults; the worker ignored them entirely. A writer honouring
write.target-file-size-bytes and a compactor rewriting to the plugin
config's size would rewrite each other's output forever.

Resolved once per job rather than per operation, so compaction committing
new metadata mid-job cannot change the settings underneath it.

* iceberg: clamp the orphan cutoff so it cannot overflow

collectOrphanCandidates converts the cutoff to a time.Duration. Past
roughly 2.5 million hours that multiplication wraps negative, putting the
cutoff in the future so every file walked looks like an orphan and gets
deleted, including data a concurrent writer has not yet committed.

Reachable today through orphan_older_than_hours.
2026-08-16 09:16:05 -07:00
4c40ec3a9e master: align default volume size with EC rows (#10761)
Co-authored-by: joe <joe@gmail.com>
2026-08-15 21:37:20 -07:00
Chris LuandGitHub b45f8314c5 ec.encode: require the shards to agree on size before deleting the volume (#10769)
* ec.encode: require the shards to agree on size before deleting the volume

Before an encode deletes the volume it just encoded, it asks whether
enough shards exist and whether they are spread across nodes. Both are
questions about presence: nothing asks whether those shards are whole.

Every shard takes one piece of each block row, so they are all written to
the same length. One that disagrees was truncated, half copied, or landed
on a disk that filled up -- and counting cannot see it, so the source
volume is deleted on the strength of a set that cannot rebuild it.

Compare the sizes the cluster already reports (shard_sizes travels in the
heartbeat) and hold the deletion back when they disagree, naming the odd
shard and its holder. Sizes reported as zero are skipped rather than read
as a disagreement: a volume server that predates shard-size reporting, or
one that has not heartbeated them yet, must not strand every encode in
the volume-plus-shards state this check exists to avoid.

* ec.encode: judge shard sizes on the newest encode generation only

The size check collected every shard the master reports for the volume,
while the recoverability check beside it counts only the newest encode
generation. A re-encode can change the ratio, so an orphaned older
generation -- one the pre-encode sweep could not reach, but the master
still hears about -- has shards of a different length by nature. Merging
those into the comparison makes a healthy current set look inconsistent,
and because the orphan keeps being reported, every retry fails and the
encode is left holding the volume and its shards for good.

Collect sizes the way CollectEcShardBitsByNode collects bits: fenced to
the newest EncodeTsNs, with unstamped entries forming the one legacy
generation.
2026-08-15 14:21:13 -07:00
Chris LuandGitHub 76a1983c86 test: re-lock and retry every chaos command, not just the balance (#10770)
The harness kills shells mid-command, and the master releases the dead
session's lock only when it notices the connection is gone. That cleanup
lands after the harness has already re-acquired the lock, so it can clear
the lock this run holds and the next command refuses with

  need to run "lock" first to continue

recoverInterruptedBalance answered that the way an operator would -- run
lock again and retry -- but the encode and decode recoveries called
shellCommand once and required success, so the same reap failed the run
outright. Move the retry into shellCommand: the reap can land during any
command that follows a kill, not only a balance.
2026-08-15 14:13:37 -07:00
Chris LuandGitHub fbd85d31b0 ec.decode: check the rebuilt .dat is complete before the shards can be deleted (#10768)
A decode ends by deleting the shards it read, and the only thing standing
between that and a bad reconstruction is verifyDecodedVolumeBeforeDelete,
which asks whether .dat and .idx are non-empty. A .dat truncated to a
single byte passes, and the shards -- the only other copy of everything
past the cut -- are deleted on the strength of it.

The server already knows the answer it never checks: FindDatFileSize
returns the extent the EC index references, and WriteDatFile rebuilds to
it. Compare the two once the file is written and fail the decode instead
of reporting a short volume as a good one.

Longer than the extent still verifies -- padding is not missing data --
so only a genuinely short rebuild is rejected.

Needle counts cannot answer this: .idx is written from .ecx, so the count
matches by construction and a truncated .dat still reports every needle.
2026-08-15 13:28:49 -07:00
Chris LuandGitHub 829064af71 ec.decode: finish the cleanup an interrupted decode left behind (#10767)
A decode deletes the shards only after the regenerated volume is mounted
and verified, so a run interrupted in that last phase leaves the volume
in place with its shards partway through deletion. The re-run then finds
both, tries to collect the shards again to rebuild a volume that already
exists, and fails on the first shard the interrupted run had removed:

  generate normal volume 3 ...: ec volume 3 missing shard 6

Nothing recovers from there: the shard set is deliberately being
destroyed, so every retry fails the same way while the decoded volume
sits there, already complete.

Finish that cleanup instead. A volume beside the shards is not enough to
act on -- an encode interrupted before it deleted the original leaves the
same shape, as does a decode killed while generating, whose volume may be
half written -- so require a data shard to be gone. Only the deletion
phase removes one, and it is also exactly the state no decode can
recover from, so finishing is the only move left rather than a choice
between two. The deletion still runs behind
verifyDecodedVolumeBeforeDelete, the check that guards it in a normal
run.
2026-08-15 13:04:37 -07:00
Chris LuandGitHub 97a155d14d admin: show capacity per storage tier and stop counting remote-tiered bytes as local disk usage (#10766)
* admin: show capacity per storage tier and stop counting remote-tiered bytes as local disk usage

A remote-tiered volume reports its cloud object's size, so summing volume
sizes inflated the dashboard's used-vs-capacity numbers (the local .dat is
gone after volume.tier.move). Split the accounting: DiskUsage now only
counts bytes on local disks, with the cloud bytes surfaced separately per
server and per remote storage name.

The dashboard gains a Storage Tiers table breaking volumes and EC shards
down by tier (each local disk type plus each remote storage), using the
per-disk-type statfs numbers already in the VolumeList response. The
volumes page badges remote-tiered volumes with their storage name, and
the EC shards page fills in real per-shard sizes instead of hardcoding 0.

* admin: review fixes for the tier capacity display

- A disk that predates disk_total_bytes now contributes its logical
  bytes to the tier's DiskUsed, so a tier mixing old and new volume
  servers doesn't underreport usage; the usage bar always reflects the
  displayed Disk Used value (the DataSize fallback in UsagePercent is
  gone, and the percent math is overflow-safe).
- getTopologyViaGRPC defaults a zero VolumeSizeLimitMb to 30000 MB like
  GetClusterVolumeServers, keeping slot-based capacities consistent.
- The dashboard volume-servers column reads Usage / Capacity to match
  its cell content, and the hdd disk-type default is shared between the
  volumes-page badge and countUniqueDiskTypes.
2026-08-15 12:35:16 -07:00
Chris LuandGitHub 1c926e8fac test: systematic EC interruption verification — exhaustive model check + deterministic kill matrix (#10764)
* ec: bounded-exhaustive model check of the volume lifecycle

The randomized chaos harness samples the state space; this enumerates
it. The lifecycle is a state machine whose steps mirror the pipelines in
this package, and the checker explores every schedule within the bound:
a crash at every step boundary, an error return running the rollback
(itself crashable at every step), a volume-server restart applying the
startup reconciliation rules in every quiescent state, and the
prescribed restart-based recovery from every crashed state.

Checked in every reachable state: durability (a readable copy always
exists), at most one generation mounted, and — a property the sweep
discipline turns out to guarantee — at most one generation's files on
disk. From every quiescent state the recovery must converge to a clean
volume. Runs in well under a second.

* test: deterministic EC interruption matrix

Enumerate every phase of every interruptible EC operation and kill a
real weed shell exactly when the phase announces itself on the command
output, instead of at a random moment: four encode phases, four decode
phases, and the balance's move phase (set up with -rebalance=false so a
move is guaranteed). Each scenario prepares its precondition, kills at
the marker, runs the prescribed recovery, and verifies every stored byte
still reads back identical.

The interruption recoveries move out of the randomized ops into shared
chaosRun helpers both drivers use.

* test: make the randomized EC chaos walk opt-in

The systematic layers — the interruption matrix and the lifecycle model
check — carry the CI coverage deterministically; the randomized walk
stays for exploratory runs, behind EC_CHAOS_SEED.

* ci: bound the EC integration suite by the job budget, not go test's default

The suite with the interruption matrix runs close to the default 10m
binary timeout on slower runners.

* test: require every interruption-matrix marker to appear

A marker that never prints means a pipeline refactor renamed or dropped
the progress line; silently degenerating into a no-interruption run
would let CI pass without exercising the boundary the scenario names.
Also recheck the marker channel after the wait: a shell that prints and
exits at once makes both channels ready, and select picking the exit
case must not report a printed marker as missed.
2026-08-14 17:45:11 -07:00
Chris LuandGitHub 602746f51d test: EC lifecycle chaos harness, with four fixes it found (#10763)
* ec: let the encode's balance see a migrating volume's shards across disk-type buckets

Shard generation writes beside the source .dat, so a cross-tier encode
(source on hdd, -diskType=ssd) leaves the fresh shards in the source
disk-type bucket. The encode's internal balance ingested only the target
bucket, saw no shards, and planned no moves; the spread guard then
correctly aborted the encode (and before that guard existed, the shards
silently stayed clumped on the generation host in the wrong tier).

EcBalance now takes the encode batch as migratingVolumeIds and ingests
those volumes' shards from every bucket, while everything else keeps the
bucket filter so a plain ec.balance never drags deliberately tiered
shards onto another disk type. The in-memory model delete also becomes
bucket-agnostic: a node holds a given shard in exactly one bucket, and a
bucket-scoped delete missed cross-bucket moves in the dry-run model.

* volume: decode reads shard 0 from its resolved path, not the EC volume's base dir

On a multi-disk server a volume's shards can sit on several disks; the
store registers each shard with its own path and CollectEcShards resolves
them, but FindDatFileSize derived the .ec00 path from the EcVolume's base
directory. When shard 0 lived on a sibling disk, VolumeEcShardsToVolume
failed with 'open ...ec00: no such file or directory' and ec.decode
aborted.

* ec: decode re-copies shards the topology claims but the target does not hold

An interrupted earlier decode or balance can leave the master believing
the decode target holds a shard whose file never landed: the mount
registered but the partial copy was cleaned, or the file was swept. The
collect step took the topology's word for it, excluded the shard from
the copy set, and the decode failed with 'missing shard'. Probe the
target's live inventory (VolumeEcShardsInfo) and treat anything it
cannot serve as still-to-copy.

* ec: decode discovers shards across disk-type buckets

Shards sit wherever encode generation and balance left them: a
cross-tier encode leaves them in the source disk-type bucket, a partial
migration straddles buckets. ec.decode scoped its shard discovery to the
-diskType bucket and reported a decodable volume as having no shards at
all. Union across buckets, the way the encode's shard verification
already does.

* test: EC chaos lifecycle harness

Randomized, seeded sequences of the EC lifecycle against a live cluster
in the production-shaped layout: multiple data disks per server, a
separate -dir.idx directory so .ecx/.ecj sidecars are shared across
disks, and a tagged ssd tier. Operations cover encode (hdd and ssd
targets), balance, shard damage plus rebuild, decode, re-encode,
deletes, scrub, tier moves, crash-restarts, sidecar fault injections
(a data-dir .vif pushed into the shared idx dir; a stale-generation
shard planted beside a newer encode), and interruptions: a real weed
shell subprocess killed mid-encode, mid-decode, and mid-balance, with
the recovery re-run required to converge.

One invariant holds after every step: every stored byte reads back
identical and every deleted needle stays deleted. EC_CHAOS_SEED and
EC_CHAOS_STEPS make runs reproducible and scalable.

A known gap is tolerated and logged rather than fixed here: a shard
mounted on two disks of one node (orphan adoption after an interrupted
copy) is invisible to ec.balance's dedup and unaddressable by
ec.shard.unmount's shard@address form, so no cleanup path exists yet.

* test: fail payload-corruption checks on the test goroutine

t.Fatalf inside require.Eventually's condition runs on the poller's
goroutine, where Goexit kills only that goroutine and the corruption
message can be lost behind a generic timeout. Record the mismatch, end
the polling, and fail on the test goroutine. Also assert the full shard
count in the cross-bucket decode-discovery test.
2026-08-14 17:26:54 -07:00
Chris LuandGitHub 944d967502 refactor: extract EC orchestration into a shared weed/ec package (#10760)
* shell: move ErrorWaitGroup to weed/util

* shell: remove unused CandidateEcNode and EcRack types

* ec: extract EC orchestration logic from weed/shell into weed/ec

Move the EC node/topology model, balance engine, encode pipeline, decode
pipeline, and rebuild engine into a new weed/ec package so shell commands
and maintenance workers can share the logic. Shell commands keep flag
parsing and delegate through a small ec.Env (dial option, topology fetch,
volume locations, lock check). Tests move along with the code.

* shell: remove unused proportional-rebalance type stubs

* ec: move scrub, replication check, and shard unmount engines into weed/ec

* worker: share the EC generation-aware shard counter from weed/ec

* ec: gofmt

* shell: drop EC aliases with no remaining callers

* ec: guard a missing topology hook and nil disk entries in topology helpers

* ec: drop trailing newlines from decode error strings

* ec: re-check the shell lock before applying shard unmounts

* shell: trim -node entries in ec.scrub
2026-08-14 13:54:12 -07:00
Chris LuandGitHub f66d6ffc4a s3: option to disable bucket auto-creation on upload (#10759)
* s3: add option to disable bucket auto-creation on upload

* command: expose -autoCreateBucket in s3, filer, server, and mini

* s3: apply the bucket auto-create policy to directory marker uploads

* s3: validate the bucket name before the auto-create disabled check

* s3: cover the disabled auto-create gate at all three upload entry points
2026-08-14 10:58:45 -07:00
Chris LuandGitHub 02b3ec6e90 sftp: url-encode the upload path (#10758)
sftp: url-encode the upload path so filenames can't inject filer query commands

The SFTP put handler concatenated the user-controlled filename straight into
the filer upload URL, so a name containing "?" was parsed as a query string.
Build the URL via url.URL{Path: ...} so "?" becomes %3F and stays a literal
path character.
2026-08-14 09:19:37 -07:00
Chris LuandGitHub c2ea452b9d skiplist: fix TestFindGreaterOrEqual flake (compare against largest key, not value) (#10757)
skiplist: compare against the largest key, not its value, in TestFindGreaterOrEqual
2026-08-14 08:31:31 -07:00
Chris LuandGitHub d713ab49f9 volume: validate replica targets and restrict gcs credentials in FetchAndWriteNeedle (#10755)
* volume: validate replica upload targets in FetchAndWriteNeedle

The replica leg forwarded the fetched needle to a caller-supplied address
without checking it, so a malformed target could redirect the upload to an
unintended host or path. Require each replica target to be a bare host:port
whose host is not loopback / link-local / unspecified, reusing the address
deny-list; cluster peers legitimately sit on private networks, so RFC 1918 /
CGNAT stay allowed and -volume.allowUntrustedRemoteEndpoints still opts out.

Validate every target up front so a bad one fails the request before the local
write, and upload through a client that re-checks the resolved address at
connect time so a replica hostname cannot rebind to a blocked address after
validation. Mirrored in Rust (validation moved ahead of the local write; the
Rust S3 path's connect-time re-check is still a follow-up there).

* volume: only accept inline gcs credentials in FetchAndWriteNeedle

The gcs credentials value on this request could name a local filesystem path,
which the SDK reads from disk. Accept only inline JSON here; the server-side
GOOGLE_APPLICATION_CREDENTIALS env var still supplies a path. The Rust volume
server has no gcs backend, so there is nothing to mirror.
2026-08-13 23:33:01 -07:00
Chris LuandGitHub 9125b9c835 volume: extend the remote-endpoint guard to the azure backend (#10754)
* remote_storage/azure: allow a per-request HTTP client

Thread an optional *http.Client through NewAzBlobClient and add
azure.MakeWithHTTPClient, mirroring the S3 backend. When set, the client
overrides the azblob transport so a caller can pin the dial path. The
existing makers pass nil, so behavior is unchanged.

* volume: extend the remote-endpoint guard to the azure backend

The endpoint validation and rebinding-safe dialer in FetchAndWriteNeedle
covered the S3-SDK backends. The azure backend also dials a caller-supplied
AzureEndpoint, so route both families through a single guardedRemoteClient
helper that returns the endpoint each backend dials and a constructor bound
to the guarded HTTP client. azure is guarded only when AzureEndpoint is set;
an empty endpoint derives the public host from the account.
-volume.allowUntrustedRemoteEndpoints still opts out.

* rust volume: assert the azure endpoint has no remote-client path

The Rust volume server has no azure backend, so make_remote_storage_client
rejects the type before any client is built. Add a regression test pinning
that invariant.
2026-08-13 22:32:59 -07:00
Chris LuandGitHub 94f8e2caf9 EC: handle zero-sized shard files uniformly (moves, rebuilds, startup cleanup) (#10753)
* volume_move: treat zero-sized EC shards as absent in move verification

A zero-sized shard file is residue of a failed operation (issue 10730),
not a shard - but VerifyEcShards only checked presence, so a copy that
landed as an empty file passed verification and the source was deleted
behind it. Size zero now reads as absent, with a distinct error naming
the zero-sized shard so the operator can tell a broken copy from a
missing one.

* storage: exclude zero-sized EC shards from rebuilds and clean up stale ones

The reproducer in issue 10730: a zero-sized shard file left by a failed
operation was selected as a Reed-Solomon input and failed the whole
rebuild with an input size mismatch, because input discovery checked
existence, not substance.

- RebuildEcFiles treats a zero-sized shard file as missing and
  regenerates over it in place (the reclassified-corrupt path: temp
  file beside the residue, atomic rename).
- The startup/rescan shard loader, which always skipped zero-sized
  files, now deletes them once they are older than an hour - young
  enough files can be an in-flight copy's just-created file, since the
  same scan runs from LoadNewVolumes while serving.

Regression tests: a rebuild with one emptied shard regenerates it
byte-identical; the loader deletes a stale zero-sized shard and leaves
a fresh one alone.

* storage: age-check each zero-shard cleanup candidate individually

The shard scan merges the data and idx directory listings, so the
age-checked entry and a deletion candidate can be different files
sharing one name - a stale zero-sized file in one directory next to a
fresh same-named file in the other (possibly an in-flight copy's
just-created one) could get the fresh file deleted. Each candidate's
own modification time now decides, both directories are handled in one
pass, and the split-directory case is pinned by a test.
2026-08-13 21:38:22 -07:00
9386a25a4a feat(shell): parallelize volume balance moves (#10737)
* chore: volume.balance parallelization

* chore: volume.balance add ioBytePerSecond

---------

Co-authored-by: Konstantin Lebedev <whitefox@mayflower.work>
2026-08-13 19:11:38 -07:00
Chris LuandGitHub 0de7ff5eb8 ci: run the gated redis store tests (#10746)
* redis2: route the orphan cleanup existence checks to the master

* scaffold: the redis_cluster2 read routing key is useReadOnly

* ci: run the gated redis store tests

* redis2: poll for the redis expiry instead of a fixed sleep

* redis2: assert the value key exists before testing its expiry
2026-08-13 13:36:31 -07:00
Chris LuandGitHub 0481f712b1 redis2: orphan cleanup existence checks must not read replicas (#10745)
* redis2: route the orphan cleanup existence checks to the master

* scaffold: the redis_cluster2 read routing key is useReadOnly
2026-08-13 13:33:15 -07:00
Chris LuandGitHub ae2cc8225e rust volume: mirror the VolumeConsolidateIndex RPC from Go (#10752)
The Go volume server has VolumeConsolidateIndex, which moves a volume's
.idx out of the data directory into the configured -dir.idx directory
(where an EC decode/reconstruct can leave it co-located) and reloads the
volume in place. The Rust port's proto omitted the RPC entirely, so its
generated VolumeServer trait was one method short of Go's.

Add the proto message and rpc, the gated grpc handler, and
Store::consolidate_volume_index / Volume::relocate_index_to, mirroring
Go's Store.ConsolidateVolumeIndex and Volume.RelocateIndexTo -- including
the cross-device copy fallback and the reopen-against-the-old-dir path
when the move fails.

Integration tests cover the real move (index relocated, volume still
serves reads and the move is idempotent), the no-op paths (index already
in place, no separate idx dir) and the not-found error, plus the grpc
handler end to end.
2026-08-13 13:30:58 -07:00
Chris LuandGitHub 7f27c572c4 log_buffer: end bounded reads that find the buffer empty (#10750)
A bounded LoopProcessLogData (stopTsNs set) on a buffer that never took a
write since process start fell into the ResumeFromDiskError branch, which
never checks stopTsNs when ReadFromDiskFn is nil and HasData() is false.
The read parked on the notification loop forever while the subscription's
idle heartbeats kept the stream looking alive, so a bounded
SubscribeMetadata pass on a freshly restarted idle filer never completed.

Terminate like the caught-up path does, returning a nil error: leaking
the pending ResumeFromDiskError would latch the filer's outer loop into
its gap machinery, which parks the bounded subscriber all over again.
2026-08-13 13:25:52 -07:00
Chris LuandGitHub 4f50c5b0d4 feat: throughput limits for replicate, EC shard, and worker-driven moves (#10749)
* feat: throughput limits for replicate, EC shard, and worker-driven moves

VolumeCopy was the only rate-limitable transfer; EC shard copies,
replica creation, and worker-driven moves all ran at whatever the
receiving server's maintenance rate allowed, with no per-operation
control.

- proto: VolumeEcShardsCopyRequest and the balance / ec_balance task
  params and configs gain io_byte_per_second; 0 keeps today's behavior
  (the volume server's own maintenance rate governs).
- volume server: VolumeEcShardsCopy throttles with one WriteThrottler
  per request, shared across the shard, .ecx, .ecj, .vif, and .ecsum
  copies so the limit caps the transfer as a whole - the same shape as
  VolumeCopy.
- volume_move: ReplicateVolume accepts the limit; EcMoveOptions carries
  it through MoveEcShards/CopyAndMountEcShards into the copy request,
  with fake-client tests asserting propagation.
- shell: ec.balance gains -ioBytePerSecond; volume.tier.move's
  replication top-up honors the command's existing -ioBytePerSecond
  instead of running unthrottled.
- worker: balance and ec_balance configs gain io_byte_per_second
  (surfaced in the admin config schema), carried through detection and
  plugin job parameters into task params and handed to the shared
  mover; batch balance jobs inherit the limit from their detection
  results.

The limit is per copy stream, so maxParallelization multiplies the
aggregate ceiling.

* worker plugins: expose io_byte_per_second in the plugin config and derive it

The plugin-driven detection path derives its task Config from the
plugin configuration values, and both balance and ec_balance left
IoBytePerSecond at zero there - a configured limit silently reverted
to the server maintenance rate. Both derive functions now read the
field (clamped at zero), and the plugin descriptors expose it with
defaults so the configuration form carries it.
2026-08-13 13:22:58 -07:00
Chris LuandGitHub 7d0fff32db redis2: expire entries without destroying a concurrent recreate (#10744)
* redis2: expire entries without destroying a concurrent recreate

* redis2: repair the member when redis expiry wins the compare-and-delete race
2026-08-13 13:18:31 -07:00
Chris LuandGitHub c0f33d599b rust volume: mirror Go volume server logic to gate the admin RPCs (#10748)
rust volume: gate the remaining admin RPCs behind check_grpc_admin_auth

The Go volume server gates 29 destructive VolumeServer RPCs on the
-whiteList admin check; the Rust port only gated 14. Add the gate to the
other 15 -- batch_delete, read_all_needles, fetch_and_write_needle, the
EC-shard generate/rebuild/copy/unmount/to-volume RPCs, both tier-move RPCs,
volume_copy, volume_tail_receiver, set_state, scrub_ec_volume and
volume_needle_status -- so a configured whitelist restricts them the same
way it already does on the Go side.

check_grpc_admin_auth also required peer info before checking whether any
control was configured, unlike Go's `if vs.guard == nil { return nil }`.
Short-circuit when no whitelist and no signing key are set, so in-process
callers keep working with security inactive and only the gate ordering
changes for configured servers.

tests/admin_auth_coverage.rs mirrors the Go coverage test: every handler
must either gate or be listed as intentionally open with a reason, so the
two implementations can't silently drift apart again.
2026-08-13 13:17:23 -07:00
Chris LuandGitHub 4500bdf88e iceberg: accept lowercase parquet file format when planning compaction (#10751)
* iceberg: accept lowercase parquet file format when planning compaction

* iceberg: expect absolute added-file paths in compaction integration test
2026-08-13 13:16:07 -07:00
Chris LuandGitHub 76d3fd0e9d grpc: optional client_cert/client_key for outgoing mTLS connections (#10747)
* grpc: optional client_cert/client_key for outgoing mTLS connections

* scaffold: list client_cert/client_key in each grpc section
2026-08-13 13:15:20 -07:00
Chris LuandGitHub abd36cbf92 redis2: harden the orphaned index member cleanup (#10743)
* redis2: derive the orphan cleanup keys inside the helper

* redis2: skip orphan cleanup in super large directories

* redis2: detach orphan cleanup from the request context and log a failed restore

* redis2: keep a directory member whose child index is still live

* redis2: run restore-path tests under both key prefixes and fix the test harness

* redis2: check cleanup errors in tests
2026-08-13 13:08:26 -07:00
Chris LuandGitHub 4fb5d15019 redis: remove orphaned directory index members on listing (#10742)
* redis: remove orphaned directory index members on listing

* redis: check cleanup errors in tests
2026-08-13 10:54:51 -07:00
f7ae2d4dd5 fix(redis2): remove orphaned directory index members on listing (#10735)
* fix(redis2): remove orphaned directory index members on listing

ListDirectoryEntries skipped index members whose value key was gone and
left them in the ZSET, so the per-directory child index grew without
bound under any TTL workload. Mirror the ZRem the logical-expiry branch
already performs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(redis2): keep the index member when a concurrent insert recreates the value

The orphan cleanup removed the member unconditionally, so an InsertEntry
landing between FindEntry and the ZRem left a live value with no index
member, invisible to listings until another InsertEntry on that path.
UpdateEntry does not re-add it, so the loss persisted.

Restore the member when the value is present again after the removal.
The value key and the directory index key hash to different slots, so a
Lua script or MULTI over both is not available to the cluster store.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 10:53:33 -07:00
Chris LuandGitHub 8714f42abf erasure_coding: share the EC shard teardown primitive (#10740)
The unmount+full-teardown of EC shards was duplicated: the plugin-worker EC
task had unmountAndDeleteEcShards and the shell had unmountAndDeleteEcShardsQuiet,
byte-identical apart from a fence parameter and a sentinel error. That
duplication is how the teardown fence semantics drifted between the two paths.

Distribute, mount and verify already live in weed/storage/erasure_coding and are
shared by both callers; move the teardown there too, as UnmountAndDeleteEcShards
plus the shared ErrFullTeardownNotAcked sentinel. Both paths now call the one
function, so the fence semantics cannot diverge again. The shell keeps a thin
type-converting wrapper and aliases the sentinel; behavior is unchanged.
2026-08-13 10:37:25 -07:00
Chris LuandGitHub 6408f32232 EC worker: clear stale/interrupted shards at task start and on failure (#10738)
* EC worker: clear stale/interrupted shards at task start and on failure

The EC encode task cleared stale shards from a prior interrupted encode only
at 55% progress (after mark-readonly, copy, and generate), and used a
generation-fenced teardown. Two gaps left orphan shards behind:

  - a retried encode's prior attempt carries the same admin-issued encodeTsNs,
    and the server's teardown fence preserves same-or-newer generations, so the
    prior attempt's shards were never cleared;
  - shards left by an interrupted distribute often have an unreadable .vif
    generation (the sidecar never landed), which the fence also preserves.

Both survive the next volume-server restart as orphans and make detection
refuse the volume (Manual intervention required).

Move the cleanup to a Step 0 preflight that runs before any destructive step,
and switch it to the server's blanket (generation-independent) teardown -- the
same wipe the shell ec.encode pre-cleanup uses. The admin dedupe key already
prevents a concurrent newer encode of the volume, and the blanket path aborts
rather than clobber a live newer mount.

Add rollbackDistribute: a failure after distribute begins but before verify
commits the EC copy now tears down the shards it wrote and restores the sources
to writable, so a terminally-failed encode (a single-attempt job, or the last
of a retry series, which has no successor preflight) leaves nothing behind.

The preflight also rejects a plan with no targets or no source before marking
the source readonly.

* EC worker: reject malformed targets and keep source readonly on incomplete teardown

Address review feedback:

- ensureCleanEcStart only rejected an empty target slice; a target with an
  empty Node (or no shard ids) passed the length check, was then silently
  skipped by cleanupStaleEcShards, and let Execute mark the source readonly
  with nothing to distribute to. Validate each target before the first
  destructive step. Add regression cases.

- rollbackDistribute marked the source writable even when the shard teardown
  returned an error, exposing a writable source beside stale (possibly mounted)
  shards -- reads/writes could diverge and orphan cleanup will not remove a
  writable source. On an incomplete teardown, leave the source readonly for the
  next preflight or an operator to reconcile.
2026-08-13 10:25:48 -07:00
Chris LuandGitHub fa48ce20fc shell: roll back a failed ec.encode instead of leaving readonly volumes and orphan shards (#10741)
* shell: roll back a failed ec.encode instead of leaving readonly volumes and orphan shards

ec.encode marks the source volumes readonly and generates EC shards before it
verifies the shards and deletes the originals. If any step in between failed,
the command just returned the error: the volumes were left readonly and the
partially-produced EC shards survived as orphans, cleaned up only by the next
ec.encode run (via clearPreexistingEcShards) if the operator retried.

Add a deferred rollback that runs when the batch fails before the originals are
deleted: it tears down the EC shards produced this run and restores the sources
to writable, reusing the existing clearPreexistingEcShards and
markVolumeReplicaWritable helpers. Once the shards are verified recoverable the
batch is committed to the EC copy and does not roll back. Both rollback steps
are idempotent, so a failure before the volumes were marked readonly is safe.

* shell: re-read volume locations when restoring writable in ec.encode rollback

Address review: rollbackFailedEcEncode restored writable using the location
snapshot taken before doEcEncode, but doEcEncode re-reads locations and marks
every replica of that later snapshot readonly. A replica added or moved in
between would be left readonly. Re-read locations in the rollback and fall back
to the pre-encode snapshot only if the re-read fails.
2026-08-13 10:25:30 -07:00
Chris LuandGitHub db5a086d04 read cold remote objects straight from the origin while caching (#10731)
* refactor: extract remote mount resolution into shared helpers

* refactor: share the adaptive remote cache wait policy

* filer: stream cold remote reads from the origin while caching

* s3: stream cold remote reads from the origin instead of 503 retries

* test: cover the S3 origin stream-through path

* remote mounts: match on path components and prefer the longest mount

* fail short origin streams instead of silently truncating

* s3: try the origin before failing a cold read on a local cache error

* s3: gate origin streaming on the entry's resolved version

* return the cache RPC's NotFound as a canonical status and classify it everywhere

* filer: keep multipart-range cold reads on the retry path
2026-08-12 23:00:10 -07:00
Chris LuandGitHub a0347ca545 test: assert EC shard identity and empty-view in multi-disk lifecycle tests (#10723)
test: assert EC shard identity and empty-view, not just counts, in lifecycle

Follow-up to the multi-disk EC lifecycle tests (#10721), addressing review
feedback.

The phase checks compared shard counts. A reconcile that put a shard on the
wrong disk, or loaded a different shard than the file on disk, keeps 6/5/3
right while corrupting the mapping. Compare the exact registered shard set per
disk at every phase instead, via a shared assertRegistered helper. The
cross-disk mount phase now also pins that shard 0 landed on disk2 with the
existing shards, not merely that it is findable.

The sidecar-disk-lost scenario only logged the registered view, so a change
that registered shards without reachable sidecars would pass despite the
documented expectation that the view stays empty. It now asserts
countRegistered == 0: a registered-but-unreadable shard is worse than an
unregistered one, because the master advertises it.

The first store's closer is now deferred as a closure the moment the store is
created, so a Fatalf in an early phase no longer leaks it and its
notification-drainer goroutine; the closure reads the reassigned variable so it
also covers the post-restart store.
2026-08-12 20:18:39 -07:00
Chris LuandGitHub 78e7e04377 plugin scheduler: drain started jobs past the window close instead of killing them (#10728)
* plugin scheduler: drain started jobs past the window close instead of killing them

* plugin scheduler: never drain-cap an attempt below its declared estimated runtime

* plugin scheduler: cap estimated_runtime_seconds before the Duration conversion
2026-08-12 19:38:22 -07:00
Chris LuandGitHub 0799084e98 refactor: share volume and EC shard move logic between shell and workers (#10727)
* operation: add shared volume_move package for volume and EC shard moves

The shell commands (volume.move, volume.balance, ec.balance, tier moves)
and the maintenance workers (balance, ec_balance) each carried their own
copy of the move RPC sequences, and the copies had drifted: the worker
verified the target before deleting the source but dropped the disk
type and IO throttle; the shell passed those but deleted the source
unverified.

volume_move.Mover carries the merged sequences, keeping the stricter
behavior from each side:

- LiveMoveVolume: check-then-hard-freeze the source (VolumeStatus's
  IsReadOnly also covers low-disk and readonly-but-can-delete states,
  which still accept needle deletes), copy with disk type and IO
  throttle, tail, verify the target is not behind the source before the
  destructive source delete (a target that is ahead holds writes it
  accepted during the tail and the move commits to keep them), and
  restore the source's writability when a failure precedes the delete
  and this move did the freezing. Aborts clean up the incomplete target
  copy; a failed cleanup or an ambiguous source delete keeps the source
  readonly (ErrSourceKeptReadonly) so callers do not thaw a source next
  to a possibly-authoritative copy. With a readonly source, an existing
  or unknown-state target refuses the move outright: no client-side
  observation can prove such a copy is a stale remnant rather than the
  authoritative copy of an unfinished move.
- MoveEcShards: copy with the .ecx/.ecj/.vif/.ecsum sidecars, mount,
  verify the target registered every shard before unmount+delete on the
  source, and reject same-server moves (the EC delete is server-wide).

Server identity is the grpc endpoint (SameServer), so node:8080 and
node:8080.18080 compare equal while test servers sharing a degenerate
HTTP address stay distinct; addresses are validated non-fatally before
dialing and before being embedded in copy/tail requests, since both the
client dialer and the receiving server normalize them through a parser
that aborts the process on a malformed port. The Rust volume server's
codes.NotFound counts as a definitively absent probe answer alongside
the Go server's plain-error code Unknown.

All RPCs go through an injectable ClientFunc, so the sequences are unit
tested against a fake volume server client: RPC order, request fields,
and that verification failures keep the source intact.

* shell, worker: delegate volume and EC shard moves to operation/volume_move

LiveMoveVolume and the copy/tail/delete/mark-writable helpers become
thin wrappers over the shared mover, keeping their signatures; the EC
helpers keep their per-step output and delegate the RPCs. BalanceTask
and ECBalanceTask keep their parameter validation, progress reporting,
and guards (same-node cross-disk rejection, dedup keep-node
verification, shard ids range-checked before the uint8 narrowing) and
hand the RPC sequences to the mover. volume.tier.move skips its
thaw-on-failure when the mover deliberately kept the source readonly,
since reopening the replicas beside a possibly-authoritative target
copy would fork the volume.

The tail-failure tolerance moves inside the mover: a failed tail is
tolerated only when the volume was already readonly before the move
began, backstopped by a stability re-read across the idle window, so
volume.balance's -skipTailError-by-readonly heuristic and tier-move's
unconditional skip both become the same authoritative rule.

* volume_move: keep the source readonly when a failed copy leaves a target of unknown origin

A failed copy can leave a complete, mounted copy on the target (the
server finishes after the client loses the stream). The abort probed
the target only when its pre-copy state was known-absent; an unknown
prior state skipped both the probe and the cleanup and then reopened
the source - two writable replicas of one volume, diverging from the
next write on.

The abort now probes the target on every failed copy and restores the
source only when the target provably holds nothing. A copy whose
provenance cannot be proven (unknown prior state, a pre-existing
replica, or an unreachable target) is never deleted, and the source
stays readonly with ErrSourceKeptReadonly naming the recovery.

* test: teach the plugin worker harness the shared move sequence

The fake volume server lacked VolumeStatus, which the shared mover now
issues before freezing the source, and the batch execution test's
status-read accounting predates the pre-copy target probe and the
verification reads. Mirrors the harness the enterprise tree already
carries.
2026-08-12 12:29:40 -07:00
Александр МелентьевandGitHub c55ef964d9 docs: update example link (#10725) 2026-08-12 06:57:25 -07:00
Chris LuandGitHub 2a513e71a4 test: drive ec.encode/balance/rebuild E2E with a byte-identical payload check (#10722)
The existing multi-disk EC integration test asserts on shard counts. Counting
cannot tell a healthy volume from one a repair reassembled out of the wrong
inputs — both have fourteen shards. This drives the real shell commands
(ec.encode, ec.balance, ec.rebuild) against a live three-node, four-disk
cluster and reads the stored bytes back after every step, so a rebuild that
produced fourteen plausible-but-wrong shards fails here.

An 8 KB random payload is stored, then encoded, balanced, damaged (two shard
files removed and the servers restarted so the master relearns the reduced
set from disk), and rebuilt. The rebuild output matches the shape of the
support case that motivated this — "rebuildOneEcVolume", "missing shard N.0",
"copied N.1 from ..." — and the payload is verified identical after each of
upload, encode, balance, shard loss, and rebuild.

Two ordering facts the test pins, both of which cost real debugging time:
ec.rebuild is driven by the master's topology, not disk truth, so shards must
be relearned (via restart) before a repair can target the right set; and the
shell lock is dropped when the restart disconnects the master, so it has to be
retaken before the rebuild.
2026-08-11 22:44:25 -07:00
Chris LuandGitHub 3dfe4bdaaa test: walk an EC volume through a multi-disk node's whole life (#10721)
A multi-disk volume server keeps one .ecx / .ecj / .vif set per volume on a
single disk while ec.balance scatters the shards across the others. Every EC
operation on such a node crosses that split: startup registration, balancing
the sidecar disk's shards away, rebooting in that state, and mounting a shard
delivered to a disk that has no local sidecars.

Each of those transitions is handled by a different mechanism (per-disk scan,
cross-disk reconcile, mount-time .ecx lookup), individually tested but never
as the sequence a production node actually lives through — where the output
state of one transition is the input of the next. A regression in any hop
shows up as shards that exist on disk while the master's view says otherwise,
and every topology-driven repair then works against the wrong shard set.

The layout, volume id and collection mirror a support case. The second test
pins the failure floor when the sidecar disk itself dies: shards on the
surviving disks may drop out of the registered view, since nothing can read
them without the .ecx, but their files must survive so restoring the sidecars
restores the volume.
2026-08-11 21:26:55 -07:00
Chris LuandGitHub 65114575eb mount: invalidate hot directory listings by section (#10712)
* mount: invalidate hot directory listings by section

A cached directory used to be dropped whole when it saw 64 changes in
2s: with a continuous writer the listing cycled through wipe, direct
listing and full rebuild for as long as the writer kept going, and
every sibling lookup fell through to the filer in between.

Split each cached listing into name-range sections of 1024 entries. A
burst of foreign changes invalidates just the section it lands in;
entries stay served and events keep applying, and the next readdir
re-lists only that range from the filer, reconciled through the version
gate so it cannot roll back newer applied events. Lookups in an
invalidated section read through until then. The mount's own writes no
longer invalidate anything: they are ground truth for its cache.

* meta_cache: drop the version floor with a deleted or moved directory

The other teardown paths already clear both maps; a floor left behind
here would fence the listing of a directory re-created at the same
path.

* mount: harden section refresh

An unversioned listing (pre-upgrade filer) now only fills gaps instead
of reconciling: without a snapshot to order against, an overwrite or
the deletion sweep could roll back an event applied after the listing.

The section table can be rebuilt or re-split between the listing and
its apply, so the refresh only marks fresh or splits when the section
still covers the range it read. Splicing bounds from a stale range
into a rebuilt table could leave them unsorted.

Bound the wait: a readdir gives a refresh five seconds before serving
the maintained-but-unverified cache. Bound the size: a range grown
past four sections aborts the refresh and drops the directory cache,
re-tiling it with a full rebuild, with that request served direct.

Cover the filer-facing path with a listing server: paging with the
snapshot pinned across pages, the section cutoff, no calls for a
fresh section, and the overgrown-range abort.

* meta_cache: make the section table a self-contained state machine

Churn counting, freshness, stale-range scanning and the refresh
completion with its guard and split now live on dirSections itself,
free of the lock, the store and the apply loop, so they test directly
with synthetic clocks and tables. MetaCache keeps thin wrappers that
hold its mutex and find the directory's table.

* meta_cache: keep section internals out of the apply request

The request now carries the completed build's table and one refresh as
opaque values built by section code, and the boundary-derivation rule
moves out of the build loop into a collector next to the rest of the
section logic.

* mount: fence refreshed sections with a snapshot floor

A refresh versioned the entries it fetched and tombstoned the ones it
swept, but a name absent from both cache and listing kept the old
directory floor, so a delayed event between the two snapshots could
resurrect it into a section already marked fresh. The section now
carries its own floor, consulted next to the directory floor, covering
every name in the range, present or absent — which also retires the
refresh's per-entry version stamps and sweep tombstones.

An unversioned listing sets no floor and vouches for nothing: it may
still fill gaps, but the section stays stale and reads through until a
filer that stamps snapshots re-validates it.

A listing's reach is unknowable up front — a resumed handle can skip
far ahead, and shrunken sections let one batch span many — so a
readdir now re-validates every stale section from its start name to
the end of the directory instead of the next two.

* mount: fence tombstoned names with floors and gate the reconcile

A tombstone answered for its name before the floors were consulted, so
one at an old position let through events the newer listing floor
should have fenced; a build never hit this because it prunes
superseded tombstones, which a section refresh does not. The version
gate now raises a tombstone to the floors like any other record.

With no per-entry versions, only the section floor fences a
reconcile's work, so a range the rebuilt or re-split table no longer
has must not touch the store either: the range check moves ahead of
the mutations, under the same lock the floor install holds.

An unversioned refresh no longer retries: the section is remembered as
unverifiable and skipped by the stale scan, or every batch of every
readdir would re-list the same ranges against a filer that cannot
vouch for them.

* mount: clear beaten unversioned markers and skip refresh mid-build

An unversioned marker outliving the snapshot write that replaced its
content bypassed the section floor the same way an old tombstone did,
letting a delayed pre-snapshot event roll the entry back. The refresh
now clears the marker when its write wins; pinned local-only entries
are not replaced at all, keeping their content and marker.

A rebuild wipes and repopulates the store off the apply loop, so a
refresh reconciling meanwhile could sweep children the build had
already inserted and let it publish the directory incomplete. The
refresh now skips a building directory, as events (buffered) and
purges (skipped) already do; its staleness dies with the build's
fresh table.

* mount: clear the unversioned marker only after its replacement lands

Clearing before the insert meant a failed write left the old local
content claiming the listing floors, fencing the very events that were
still entitled to correct it.

* meta_cache: rename the section state machine to sectionList

dirSections named both the type and the map of them.

* mount: raise the default cacheDirMaxEntries to 100000

The low ceiling guarded against whole-listing rebuild churn: a big
cached directory under writes kept re-streaming everything. Sectioned
invalidation ended that — a burst now costs one range listing — so the
remaining cost of caching a large directory is its one-time build,
comparable to the single direct listing that read-through mode pays on
every enumeration instead.

* meta_cache: cover section border and edge cases

A bound-named entry belongs to the section starting at the bound: the
neighboring refresh's sweep stops before it, its own section's covers
it. Churn past everything the build saw lands in the tail section, a
rename spanning two sections invalidates both, and a listed entry at
the section's end name is cut off with the ones beyond it.
2026-08-11 21:11:18 -07:00
Chris LuandGitHub a7d5443125 ec: confirm a surviving copy before deleting a duplicate EC shard (#10719)
* ec: confirm a surviving copy before deleting a duplicate EC shard

The dedup phase of EC balancing removes a shard it believes exists elsewhere.
It copies nothing first, so the shard surviving on another node is the only
thing that makes the delete safe -- and it took the plan's word for that.

The plan is built from the master's topology, which can name a location that
holds nothing: such a server answers "CopyFile not found ec volume id N" when
something later tries to read the shard there. A shard listed on a phantom
location and on a real one looks duplicated, so dedup deletes one of them. When
it picks the real one the last copy is gone, and the job reports success -- the
loss only surfaces later, as a rebuild that cannot assemble enough shards.

The move phase already refuses to work on trust: it verifies the shard
registered on the destination before removing the source. Dedup now holds to
the same standard. The planner records which node it chose to keep, and both
executors -- the worker task and the shell's ec.balance -- confirm that node
really holds the shard before deleting. A keep node that cannot be queried is
unknown rather than confirmed, and blocks the delete.

Tests drive the destructive path against an in-process volume server that
tracks what is actually on disk separately from what the plan claims, which is
the distinction the bug turns on. Without the guard, two of them fail by
deleting the only copy and returning success.

* ec: check the collection and bound the wait when confirming a survivor

Two gaps in the dedup survivor check.

The inventory RPC is keyed by volume id alone, so a server holding the same
number for a different collection answers "yes, I have that shard" to a
question about this one. Accepting that deletes the last real copy on the
strength of an unrelated volume. The response already carries the collection,
so verify against it rather than widening the RPC.

The shell path also queried on a background context, so a keep node that
accepts the connection but never answers would hang the whole balance run
instead of reporting that the survivor could not be confirmed. Bound it.

The check moves into VerifyShardsOnServer next to the existing helper, shared
by both executors, so the two paths cannot drift.
2026-08-11 20:15:34 -07:00
Chris LuandGitHub 5b519489c1 remote_storage: build all S3-compatible clients through one constructor (#10720)
* remote_storage: build S3-compatible clients through one constructor

The eight non-s3 S3-SDK providers each duplicated the AWS session setup
and only the s3 maker could take a custom *http.Client. Route every
S3-compatible type (s3, wasabi, b2, aliyun, tencent, baidu, filebase,
storj, contabo) through MakeWithHTTPClient with a single options table,
and add S3CompatibleEndpoint so callers can resolve the endpoint a given
type dials. No behavior change.

* volume: apply the remote-endpoint check to all S3-compatible providers

FetchAndWriteNeedle validated the endpoint and used the pinned dialer only
for type "s3". Every S3-SDK backend (wasabi, b2, aliyun, tencent, baidu,
filebase, storj, contabo) dials a caller-supplied endpoint through the same
client, so gate on S3CompatibleEndpoint to apply the same check uniformly.
-volume.allowUntrustedRemoteEndpoints still opts out.

* volume: don't route the guarded remote-endpoint client through a proxy

The guarded client exists to dial the validated endpoint directly and
re-check the resolved IP at connect time. With http.ProxyFromEnvironment
set, the dialer only validates the proxy's address while the proxy
re-resolves the endpoint host, which reopens the rebinding window. Drop
the proxy on this path; operators that need one can opt out with
-volume.allowUntrustedRemoteEndpoints.
2026-08-11 19:06:12 -07:00
Chris LuandGitHub 980471c818 storage: count a volume's needles in uint32 (#10718)
FileCount and DeleteCount were int, so each cost a word on every replica the
master holds. A volume caps at 30GB on a 4-byte-offset build and 8TB on a
5-byte one, and neither holds 4.29 billion needles.

That takes VolumeInfo from 120 bytes to 112, which is its own size class rather
than rounding up into the 128 one, so a replica costs 135.7 bytes in the map
instead of 151.7 -- about 25MB across the 1.6M replicas in a cluster the size
of the one this came from.

Counts are narrowed where they are read rather than assigned across, so a
report claiming more than a volume can hold pins at the ceiling instead of
wrapping to a small number.
2026-08-11 16:36:36 -07:00
9fd7075bea filer: retry and surface metadata replay failures from peers (#10714)
* filer: stop silently dropping metadata replay failures from peers

When two filers do not share a store (e.g. one leveldb3 per pod), each
subscribes to its peers' metadata streams and replays their events
locally (meta_aggregator.go's maybeReplicateMetadataChange, wired into
doSubscribeToOneFiler). A failed Replay() was logged and then treated
as done anyway: processEventFn always returned nil regardless of the
replay outcome, and processOne advanced lastTsNs unconditionally. The
offset is the only record of subscription progress, so a dropped event
was gone for good - no retry, and nothing else ever observed it.

An entry that fails to replay this way diverges from its peer
permanently. This is how a bucket's quota (entry.Quota, carried on
peer events like everything else - see entry_codec.go's EqualEntry
comparing Quota, and FromPbEntry copying it in entry.go) can end up
different across filers indefinitely: one replay hiccup on one filer,
and its enforcement and any metric reading its own store diverges from
the others' with no signal anything went wrong.

Fix: replicateMetadataChange now retries a failure with util.Retry,
which already distinguishes transient errors (timeouts, connection
resets, throttling, ...) from everything else and bounds the backoff.
That covers the common case - a busy store, a blip talking to a
remote-backed backend - without changing behavior when replay
succeeds. An error that is not transient, or outlives the retry
budget, is not retried further: propagating it so the offset never
advances would stall this peer's entire stream behind one event that
may never replay, which is worse than the one entry staying stale.
Instead it is skipped, loudly - counted in a new
stats.FilerMetaAggregatorReplayFailures metric and logged at error
level - so the divergence is discoverable instead of silent.

Tested: go build ./... and go test ./weed/filer/... ./weed/server/...
Added meta_aggregator_replay_test.go: one test fails against the old
one-shot Replay call (a single transient failure is never retried, so
the store never converges) and passes with the fix; a second covers a
permanently-failing event completing quickly and being counted instead
of retried forever.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* filer: keep the test quota constant int64 for 32-bit builds

An untyped shift constant passed to t.Fatalf's ...any defaults to int and
overflows on 32-bit, failing go vet there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* filer: name the diverged entry in the give-up log line

event.Directory is only the parent (typically /buckets), so for any
directory with more than one child the previous log line could not say
which entry failed to replay - the exact thing the change exists to
make discoverable. Name comes from NewEntry, falling back to OldEntry
for deletes; both getters are nil-safe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(filer): trim metadata replay comments to the non-obvious why

Compress the added comments on replicateMetadataChange and its tests down to
the reasoning a maintainer cannot get from the code: why a retry-exhausted
failure is skipped rather than propagated, what the old one-shot Replay body
did that the test pins, and why the quota constant is typed int64. Drops
deployment-specific narration and restatement of the code. No behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: restore load-bearing clauses trimmed in the comment pass

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* filer: document and test the multi-step DeleteEntry replay hazard

CodeRabbit flagged that FilerStoreWrapper.DeleteEntry skips the delete
once FindEntry reports the path already gone, and that the redis store
families remove the primary key before parent-directory membership.
Chained together, a delete that fails between those two steps is
retried as a no-op: the stale membership is never revisited, and
replicateMetadataChange now reports overall success for it without
incrementing FilerMetaAggregatorReplayFailures, whereas before this PR
every such failure was unconditionally logged. The underlying store
inconsistency is pre-existing (a single non-retried Replay already
leaves the same stale membership behind); what retry adds is that this
one case no longer surfaces it.

Making Replay atomic or teaching every store to repair secondary
mutations on retry is out of scope here. Instead: document the hazard
at Replay, filerstore_wrapper.go's DeleteEntry, and
replicateMetadataChange, and add a test against the real
FilerStoreWrapper (not a strawman) that pins down the current,
documented behavior.

* filer: trim replay retry comments and tests

Drop the comment-only hunks documenting the pre-existing DeleteEntry
partial-failure hazard, the test that asserted that hazard still exists,
and the second hand-rolled fake store. Reuse stubFilerStore for the two
retry tests.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-11 12:36:22 -07:00
Chris LuandGitHub 5b145fe646 shell: send read jwt when downloading chunks in fs.mergeVolumes and fs.distributeChunks (#10717)
* shell: fs.mergeVolumes sends read jwt when downloading chunks

* shell: fs.distributeChunks sends read jwt when downloading chunks
2026-08-11 11:57:53 -07:00
Cavus700andGitHub 01e97955f9 Helm: Add support for schedulerName (#10716) 2026-08-11 05:55:20 -07:00
340f9951ac admin: make paths relative (#10709)
* admin: make paths relative

* admin: make filer browser link and nav path checks prefix-relative

* admin: add isCurrentPath and currentPathStartsWith helpers

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-08-10 22:31:55 -07:00
Chris LuandGitHub 790e8d3fd6 clickhouse catalog test: cover latest ClickHouse and catalog-side CREATE TABLE (#10707)
* clickhouse catalog test: cover latest ClickHouse and catalog-side CREATE TABLE

* verify catalog registration structurally and fix README image wording
2026-08-10 20:04:33 -07:00
Chris LuandGitHub 214d3599d3 windows mount: cache file data, resolved paths and attributes (#10703)
* benchmark tool for mounted filesystems

* ci: on-demand mount benchmark, native WinFsp vs rclone plus a Linux reference

* windows mount: let the Windows cache manager cache file data

WinFsp only turns the cache manager on for a file when FileInfoTimeout
is infinite; at any finite value every application read and write is a
synchronous trip into the mount process at whatever size the application
issued. Metadata events already reach FspFileSystemNotify, which purges
a changed file's cached pages and attributes, so an infinite timeout
stays coherent. The dir listing, volume info and EA timeouts are pinned
to one second so they do not silently inherit the infinity.

* windows mount: cache resolved paths and attributes in the adapter

WinFsp addresses every operation by path and has no FORGET, so the
adapter walked the whole path through Lookup on each one, and in a
directory the filer has not listed yet every walk was a filer round
trip; nothing played the part of the kernel's dentry and attribute
caches. The path cache owns one lookup reference per entry the way the
kernel holds one until FORGET, serves attribute reads for files without
an open handle, and is purged by the mount's own mutations and by
metadata events, with the timeout as backstop.

* windows mount: keep a closed file's attributes cached

Open steals the path's cache entry for its handle and Release returned
the reference with a purge, so the stat that follows every copied file
walked to the filer again. Reading the handle's final attributes before
it goes away and moving the reference back into the cache serves that
stat locally, the way the kernel's attribute cache does after a close.

Only if the path still names that inode, though: WinFsp reports the
path the handle opened with, and after a delete-on-close or a rename
caching it would resurrect an entry that is gone.

* windows mount: persist entries at create, and let the flush stay at close

WinFsp posts the cleanup and close that carry the flush after
CloseHandle has returned, so deferring the filer entry to the flush let
everything that reads through the filer race an unflushed close: a
listing missed just-written files, and a directory rename moved a
directory on the filer before its newest child existed there, leaving
the straggler flush to recreate the child under the dead path.

Flush-at-cleanup is not the answer either: it makes every handle's
cleanup flush, and those flushes race the unlinks of delete-on-close,
re-inserting the entry the unlink just removed. Persisting the entry at
create takes the ordering question away.

* mount: flush written pages before a truncate shrinks past them

The shrink trims chunks, but written pages that have not become chunks
yet are invisible to it, so the next flush wrote them back and the file
grew again, resurrecting the truncated bytes. Windows hits this on
every write-then-shrink because its flush runs after CloseHandle, but
the gap is platform-neutral.

* mount: order a file's unlink against its in-flight flush

Unlink set the handle's deleted flag bare, so a flush already past its
own check of that flag wrote the entry back right after the delete
removed it, and a delete-on-close file outlived its last handle. The
flag is now set under the handle's flush lock and re-checked under it,
so a flush either completes before the delete or sees the flag and
skips. An eagerly created handle also starts clean: the dirty mark
existed to make the deferred filer create happen at flush, and eager
creates have nothing to flush.
2026-08-10 18:46:18 -07:00
Chris LuandGitHub c6e1387f59 shell: multi-target fs.mergeVolumes and volume.mark -readonlyCanDelete (#10706)
* shell: fs.mergeVolumes distributes one volume across multiple -toVolumeId targets

* volume: volume.mark -readonlyCanDelete rejects writes but keeps accepting deletes

* seaweed-volume: mirror readonlyCanDelete volume state
2026-08-10 16:31:26 -07:00
Chris LuandGitHub 0b1f0cafee shell: keep the source readonly when the incomplete target copy cannot be deleted (#10705) 2026-08-10 12:56:45 -07:00
Chris LuandGitHub d4d8e097dd shell: volume.move cleans up when aborted after the copy phase (#10704)
* shell: volume.move restores source writability when aborted after the copy phase

* shell: volume.move removes the incomplete target copy when aborted before the source delete

* shell: give each abort cleanup RPC its own timeout
2026-08-10 12:35:40 -07:00
Chris LuandGitHub 365d3e9e87 filer: TUS concatenation extension (#10702)
* filer: TUS creation accepts Upload-Concat partial uploads

* filer: TUS final uploads concatenate completed partials

* filer: TUS concatenation tests

* filer: consumed marker pins TUS chunk ownership on completion

* filer: TUS session delete decides chunk ownership after removing the session info

* filer: TUS completion persists the consumed marker before creating the entry

* filer: TUS completion re-verifies the session after persisting the consumed marker

* filer: serialize TUS session ownership transitions per filer

* filer: surface failed TUS consumed-marker rollbacks
2026-08-10 12:32:45 -07:00
Chris LuandGitHub 89e6f9a16e shell: volume.delete and volume.move accept a -timeout (#10701)
* shell: volume.delete accepts a -timeout

* shell: volume.move accepts a -timeout
2026-08-10 11:10:29 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
6d25ccc357 build(deps): bump github.com/ydb-platform/ydb-go-sdk/v3 from 3.146.3 to 3.147.1 (#10694)
build(deps): bump github.com/ydb-platform/ydb-go-sdk/v3

Bumps [github.com/ydb-platform/ydb-go-sdk/v3](https://github.com/ydb-platform/ydb-go-sdk) from 3.146.3 to 3.147.1.
- [Release notes](https://github.com/ydb-platform/ydb-go-sdk/releases)
- [Changelog](https://github.com/ydb-platform/ydb-go-sdk/blob/master/CHANGELOG.md)
- [Commits](https://github.com/ydb-platform/ydb-go-sdk/compare/v3.146.3...v3.147.1)

---
updated-dependencies:
- dependency-name: github.com/ydb-platform/ydb-go-sdk/v3
  dependency-version: 3.147.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 11:07:51 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
3ff6ce3c69 build(deps): bump github.com/pierrec/lz4/v4 from 4.1.27 to 4.1.28 (#10695)
Bumps [github.com/pierrec/lz4/v4](https://github.com/pierrec/lz4) from 4.1.27 to 4.1.28.
- [Release notes](https://github.com/pierrec/lz4/releases)
- [Commits](https://github.com/pierrec/lz4/compare/v4.1.27...v4.1.28)

---
updated-dependencies:
- dependency-name: github.com/pierrec/lz4/v4
  dependency-version: 4.1.28
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 11:07:43 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2e86dcf725 build(deps): bump modernc.org/sqlite from 1.53.0 to 1.56.0 (#10696)
Bumps [modernc.org/sqlite](https://gitlab.com/cznic/sqlite) from 1.53.0 to 1.56.0.
- [Changelog](https://gitlab.com/cznic/sqlite/blob/master/CHANGELOG.md)
- [Commits](https://gitlab.com/cznic/sqlite/compare/v1.53.0...v1.56.0)

---
updated-dependencies:
- dependency-name: modernc.org/sqlite
  dependency-version: 1.56.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 11:07:30 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f25e415c7b build(deps): bump cloud.google.com/go/kms from 1.31.0 to 1.33.0 (#10697)
Bumps [cloud.google.com/go/kms](https://github.com/googleapis/google-cloud-go) from 1.31.0 to 1.33.0.
- [Release notes](https://github.com/googleapis/google-cloud-go/releases)
- [Changelog](https://github.com/googleapis/google-cloud-go/blob/main/documentai/CHANGES.md)
- [Commits](https://github.com/googleapis/google-cloud-go/compare/kms/v1.31.0...kms/v1.33.0)

---
updated-dependencies:
- dependency-name: cloud.google.com/go/kms
  dependency-version: 1.33.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 11:07:21 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a9bf8d76a7 build(deps): bump github.com/aws/aws-sdk-go-v2/credentials from 1.19.32 to 1.19.34 (#10698)
build(deps): bump github.com/aws/aws-sdk-go-v2/credentials

Bumps [github.com/aws/aws-sdk-go-v2/credentials](https://github.com/aws/aws-sdk-go-v2) from 1.19.32 to 1.19.34.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/credentials/v1.19.32...credentials/v1.19.34)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/credentials
  dependency-version: 1.19.34
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 11:07:13 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
6dd83c823d build(deps): bump github/codeql-action from 4.37.4 to 4.37.6 (#10699)
Bumps [github/codeql-action](https://github.com/github/codeql-action) from 4.37.4 to 4.37.6.
- [Release notes](https://github.com/github/codeql-action/releases)
- [Changelog](https://github.com/github/codeql-action/blob/main/CHANGELOG.md)
- [Commits](https://github.com/github/codeql-action/compare/v4.37.4...v4.37.6)

---
updated-dependencies:
- dependency-name: github/codeql-action
  dependency-version: 4.37.6
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 11:07:05 -07:00
Chris LuandGitHub 7c87d78ea2 s3: a key deleted after enabling versioning must leave the listing (#10684)
* s3: a null object wins over a rescan when the latest-version pointer is absent

The read path already resolves an absent pointer this way; the listing-path
counterpart scanned .versions/ first and could surface an old version or
delete marker over the current suspended-versioning null object.

* s3: dedup a key against its .versions sibling in suspended buckets too

A suspended bucket keeps its .versions directories, so a suspended-versioning
null object and its .versions sibling emitted the same key twice.

* s3: retract a null object from the listing when a delete marker shadows it

Deleting a key whose null version predates versioning leaves the base-path
entry in place and records the delete marker under <key>.versions. The
listing appended the base-path entry and relied on the .versions sibling to
replace it, but a delete-marker current version emitted nothing, so the
deleted key stayed visible to ListObjects while GET and HEAD returned 404.

* s3: keep a key's .versions sibling on the same page as the key

When the page quota ran out between a base-path entry and its .versions
directory, the page ended with the stale entry and the next page skipped the
directory as a marker echo, so the replacement or retraction never happened.

* s3: the null version is not latest when the .versions pointer names a newer one

ListObjectVersions stamped IsLatest on every base-path null object, so a key
deleted after enabling versioning reported IsLatest on both the delete marker
and the null version.

* s3: test listing after a pre-versioning null object is delete-marked

* s3: find a key's earlier page entry by scan, not by adjacency

A key such as k.bak sorts between k and k.versions, so the entry a .versions
sibling replaces or retracts is not always the last one on the page. Scan
back through the page for the key, and insert a late resolution in sorted
position instead of at the end.

* s3: settle trailing null objects by lookup when a page fills

The quota can run out while keys still sit between a null object and its
.versions sibling, and the sibling-adjacent page-boundary exception never
fires for those. Track the trailing null objects whose sibling has not been
ruled out and look each one up before declaring the page full; a retraction
reopens the quota.

* s3: do not resolve a .versions sibling its page has already moved past

A page resuming from a marker inside the base key's extension region has
already listed and settled the base null object on an earlier page, so
resolving the .versions directory again re-emitted the key.

* s3: test listing with keys between a null object and its .versions sibling

* s3: pick the newer of the null object and the scanned versions

Making the null object win outright whenever the pointer is absent misread
multi-filer pointer lag: version files replicate ahead of the pointer, and a
key overwritten or delete-marked after pre-versioning days would list its
stale null again. The suspended-versioning write that legitimately makes the
null current is also the newer entry, so mtime tells the two apart.

* s3: a delete-marked null object no longer keeps its prefix alive

The hidden-entries probe took any plain file as proof of a listable key, but
a null object shadowed by its .versions sibling's delete marker is not one.
Hold plain files pending until the sibling settles them either way.

* s3: settle an evicted pending null instead of dropping it

Nested keys like k, k!, k!! can hold more pending nulls than the cap. A
silently evicted one could close the page unsettled, and the resume skip
would then keep the stale entry for good.

* s3: test deleted-prefix hiding and the pending-null cap

* s3: cover the reported '!' intervening key with a live version

* s3: an unstamped same-second version outranks the null object

Second-resolution mtimes cannot order same-second writes, so the tie went to
the stale null when the pointer lagged. The suspended write that makes a null
current stamps the version it displaces before clearing the pointer, so the
stamp is the authoritative signal and a tie without it goes to the version.

* s3: a pointer-less versions listing still checks what replicated

ListObjectVersions took a missing pointer as proof the null object is latest,
but under pointer lag the sibling can already hold newer replicated versions
or markers. Apply the same nullObjectWins rule as the listing recovery.

* s3: a failed null-object settlement fails the listing

Every getEntry error read as a missing sibling, so a transient filer error at
a page boundary committed the unsettled null and the next page skipped its
sibling for good. Only a definitive not-found means the null is live; other
failures are retained on eviction and fail the request at page close.

* s3: retract a CommonPrefix whose only backers were delete-marked nulls

The directory probe settles this for the / delimiter, but any other delimiter
derives prefixes from base-path keys directly, and a prefix built solely from
null objects survived their delete markers. Count the unsettled null backers
behind the newest prefix and retract it when the last one settles as a marker;
a live resolution or any listable contributor confirms the prefix instead.

* s3: test custom-delimiter prefix retraction

* s3: an explicit signal marks the null object current, not the demotion stamp

The NoncurrentSinceNs stamp survives promotion: delete the version that
demoted another and the promoted one is current yet still stamped, so a
lagging replica would resurrect the stale null. A suspended-versioning write
now records Seaweed-X-Amz-Null-Version-Is-Latest on the .versions directory
when it clears the pointer, every pointer update removes it, and the
recovery paths trust the signal instead of the stamp.

* s3: a filer failover retry rebuilds the listing page from scratch

The failover wrapper reruns the callback on another filer after a transport
error, and the partially built page, spent quota, and advanced marker leaked
into the retry, which could then return a stale or duplicated page as
success.

* s3: only a prefix's own backers can debit it

A delete marker for a version-only key (no base object) derived the same
prefix as its neighbors and decremented backing it never contributed,
retracting a prefix that a live null object still backed. Track backers by
key so settlement is idempotent and only debits what was counted.

* s3: test a version-only marker against a null-backed prefix

* s3: a pointer recompute clears the null-current signal

The routed finalize for delete markers, COPY, and multipart rewrites the
.versions pointer through RECOMPUTE_LATEST, which left a suspended-era
null-current signal in place. Version files never carry the signal, so
mapping it in CopyExtended deletes it whenever the pointer recomputes.

* s3: the pointer outranks the null-current signal in the versions listing

The signal check guarded the pointer check, so a stale signal a recompute
had not cleared yet would have let the null claim IsLatest alongside the
pointed-at version.
2026-08-10 11:04:06 -07:00
Chris LuandGitHub 753cb8cda8 master: stop copying the cluster to name it (#10700)
* topology: name a node's volumes without copying them

ToVolumeLocations reads a volume id off every volume in the cluster, and got
there through GetVolumes, which copies a whole storage.VolumeInfo per volume to
be read for four bytes of it. Every client that connects asks for this.

At 800k volumes the walk goes from 94.6MB to 16.0MB, which is the ids
themselves.

* master: log why a client send failed, not what was sent

The message names every volume on a newly connected node, so a client going
away had the master format a protobuf that size into text -- through the one
log level that is always on. The error is the part worth having.
2026-08-10 11:02:09 -07:00
Chris LuandGitHub 5e9b7833ee topology: keep the volume map's values out of its slots (#10680)
Go stores a map value inline once it fits in 128 bytes, and indirectly above
that. Shrinking storage.VolumeInfo to 120 bytes therefore moved 1.6M of them
into the map's own slots, which are allocated to capacity rather than to
occupancy, and the map grew by 149MB even though each volume got smaller.

Hold pointers, so the size of a volume record stops deciding how it is stored.
Updates are written through the pointer already there, so a heartbeat that
changes a volume allocates nothing, as it did when the struct was above the
threshold and Go was doing this itself.

At 800k volume ids across 3 servers the resident topology goes from 705.3MB to
546.6MB, with heartbeat cost unchanged at 111.6MB.
2026-08-10 11:00:05 -07:00
Chris LuandGitHub 52d74df4d1 clients: stream the volume listings that ask for everything (#10679)
* master: stream volume listings

A listing of 800k volumes is 36MB on the wire but 305MB as messages, and the
master built all of it, then held it while grpc encoded it. Two of those at
once is most of a small master's heap, and the maintenance scanner asks every
30 minutes.

The topology goes out first, listing nothing, then its volumes in batches, so
the master holds a batch rather than a cluster: 341MB of live heap for one
listing becomes 4.4MB. It allocates much the same either way -- what changes is
how much of it has to be live at once, which is what sets the heap ceiling.

Batches are built under their disk's lock and sent outside it, so a slow reader
stalls the stream rather than the topology. They therefore do not share one
instant, which a single listing did not either: it takes each disk's lock in
turn, so a volume moving during either can be seen twice or not at all.

The client helper hides which kind of master answered: one too old for the
stream is asked the old way and its reply cut into the same batches. Either way
the topology handed over lists no volumes, so a caller cannot come to depend on
finding them there.

* admin: stream the listing the maintenance scan reads

It asks for every volume in the cluster every 30 minutes. Reassembling it
client-side keeps the scan identical -- ActiveTopology splits disks by the
disk ids on the volumes, so it needs them in the topology -- while the master
no longer builds the whole reply to send it.

* topology: report a disk id that does not depend on map order

A topology disk that fronts several physical disks took its reported id from
whichever volume the map yielded first, so two listings of an unchanged disk
could disagree. Take the smallest instead.

* topology: test that a streamed listing rebuilds to the whole one

The callers that stream now rebuild the listing from a topology sent without
volumes plus the batches after it, so that has to come out the same as being
sent it whole, at every batch size and under a filter.

* clients: stream the volume listings that ask for everything

The dashboard's list and export pages, the collection and ec shard pages, the
topology view, the worker metrics and two shell commands each asked the master
to build all 800k volumes into one reply. They read the same listing as before,
rebuilt on their side, so the master no longer holds it.

The three that already ask for one volume or one collection stay as they are:
their replies are small, and streaming one costs a round trip to say so.
2026-08-10 09:51:08 -07:00
Chris LuandGitHub 46ce8cbe84 master: stream volume listings (#10676)
* master: stream volume listings

A listing of 800k volumes is 36MB on the wire but 305MB as messages, and the
master built all of it, then held it while grpc encoded it. Two of those at
once is most of a small master's heap, and the maintenance scanner asks every
30 minutes.

The topology goes out first, listing nothing, then its volumes in batches, so
the master holds a batch rather than a cluster: 341MB of live heap for one
listing becomes 4.4MB. It allocates much the same either way -- what changes is
how much of it has to be live at once, which is what sets the heap ceiling.

Batches are built under their disk's lock and sent outside it, so a slow reader
stalls the stream rather than the topology. They therefore do not share one
instant, which a single listing did not either: it takes each disk's lock in
turn, so a volume moving during either can be seen twice or not at all.

The client helper hides which kind of master answered: one too old for the
stream is asked the old way and its reply cut into the same batches. Either way
the topology handed over lists no volumes, so a caller cannot come to depend on
finding them there.

* admin: stream the listing the maintenance scan reads

It asks for every volume in the cluster every 30 minutes. Reassembling it
client-side keeps the scan identical -- ActiveTopology splits disks by the
disk ids on the volumes, so it needs them in the topology -- while the master
no longer builds the whole reply to send it.
2026-08-10 09:41:00 -07:00
Chris LuandGitHub 98f9e67b4d topology: provisional volume update must not erase the reported disk id (#10687)
Volume growth registers a provisional record before it can know which
directory the server chose, while the server's own report -- pushed
during the AllocateVolume RPC -- carries the real disk id. The merge is
last-writer-wins, so whichever lands second sticks, and fresh volumes
nondeterministically show disk 0 on multi-dir servers. Keep the reported
disk id when the provisional update carries none, before the report
digest is computed so the stored record stays consistent with what the
server keeps reporting.

Claude-Session: https://claude.ai/code/session_01QdTEEPbg4MtcoEGwqbgtZC
2026-08-10 00:41:33 -07:00
Chris LuandGitHub 65b9ae7704 master: keep disk_id when registering volumes from incremental heartbeats (#10686)
The volume server names the directory index in every
VolumeShortInformationMessage, but NewVolumeInfoFromShort dropped it, so
volumes registered through the incremental new-volume path showed
disk_id 0 at the master until a full report -- misreporting multi-dir
servers in volume.list and the per-physical-disk topology views.

Claude-Session: https://claude.ai/code/session_01QdTEEPbg4MtcoEGwqbgtZC
2026-08-10 00:40:29 -07:00
Chris LuandGitHub 00c5572e8c volume: decode IPv6 transition addresses in the remote-endpoint guard (#10683)
* volume: decode IPv6 transition addresses in the remote-endpoint guard

checkBlockedIP normalized only ::ffff: mapped IPv4, so NAT64 (64:ff9b::/96),
6to4 (2002::/16), Teredo (2001:0000::/32), and IPv4-compatible (::/96) addresses
that embed an internal IPv4 (loopback, 169.254.169.254, RFC 1918) passed the
endpoint guard even though the plain IPv4 forms are refused. Extract the
embedded IPv4 from those forms and re-check it against the deny list, which
covers both the up-front validation and the dial-time guard. Mirrored in the
Rust volume server.

* volume: require the full NAT64 well-known prefix before decoding

Only 64:ff9b::/96 carries the embedded IPv4 in the low 32 bits, so also require
bytes 4-11 to be zero before treating an address as NAT64; other 64:ff9b:
prefixes place the IPv4 elsewhere and are left untouched. Add public-target
coverage for 6to4, Teredo, and IPv4-compatible so every decoder is exercised on
both a blocked and an allowed destination. Mirrored in the Rust volume server.
2026-08-09 23:22:13 -07:00
Chris LuandGitHub 3911e4c548 master: keep a racing registration out of a dying collection (#10677) 2026-08-09 22:20:34 -07:00
Chris LuandGitHub c8cc56be91 iceberg: route unprefixed requests to the first table bucket (#10675) 2026-08-09 22:20:13 -07:00
Chris LuandGitHub a2ff9cca27 master: let VolumeList ask for the volumes it wants (#10674)
* master: let VolumeList ask for the volumes it wants

The request carried nothing, so every caller was answered with the whole
cluster. A dashboard opening one volume's page, or a capacity probe adding up
one bucket, was served all 800k of them and threw away the rest -- and the
master built every one of those messages first.

The topology, its disks and their counters are still reported in full: a caller
reading free space or replica placement needs the cluster whichever volumes it
asked about. Only what is listed under a disk is selected, ec shards included.

An empty collection and a zero volume id take everything, the way volume.list
already reads its own -collectionPattern and -volumeId, so a caller that
forgets to narrow is answered too much rather than answered wrongly. That
leaves the default collection unnameable, since it is the one the empty string
names, so it gets a field of its own.

An older client sends none of it and is answered exactly as before.

* admin: ask the master for the volume the page is showing

A volume's detail page was pulling every volume in the cluster to find one and
its replicas, and discarding the rest.

* admin: ask the master for the ec volume the page is showing

Same as the volume detail page: one volume's shards were found by pulling every
ec shard in the cluster.

* s3: ask the master for the bucket's own collection

The SOSAPI capacity probe summed one collection's volumes out of a listing of
every volume in the cluster. Cluster capacity still comes out the same: it is
read from the disk counters, which a filtered listing reports in full.

* topology: read the disk usage counters atomically

They are written with atomic.AddInt64 from heartbeats but were read plainly by
the two listings and by FreeSpace, and the map they sit in was iterated without
the lock its neighbour takes. Under -race a listing concurrent with a heartbeat
trips on both.
2026-08-09 21:59:42 -07:00
Chris LuandGitHub e428b05224 test: let the vacuum shell session outlive the vacuum (#10682) 2026-08-09 21:59:25 -07:00
Chris LuandGitHub f09e8345c6 storage: stop keeping the remote storage key on the master (#10672)
A master decides nothing from it. Every caller that read it was asking whether
a volume is remote, which the backend name answers, and the value itself is
reported on demand by the server holding the volume, through the volume info in
ReadVolumeFileStatus.

It is also the one string here that cannot be shared: unique per volume, so
unlike the collection and backend names it carries its own characters for every
volume a master tracks.

VolumeInfo goes from 136 bytes to 120. 800k volumes registered from a heartbeat
that has been over the wire go from 214 to 163 B/volume when tiered.

The volume server's own status page keeps showing the key, now read from the
volume it holds rather than relayed through a master, which is also where the
other volume server implementation reads it.

The heartbeat digest drops it on the same grounds: a change to something the
master does not hold cannot make its copy stale. Both implementations and their
shared vectors move together, and the field-coverage test now names what is
deliberately not retained rather than being loosened.
2026-08-09 12:43:31 -07:00
Chris LuandGitHub 0f7a64c596 storage: order VolumeInfo by alignment (#10669)
* storage: order VolumeInfo by alignment

The struct is held for every volume replica in the cluster, so the padding the
compiler inserts is multiplied by however many volumes a master tracks. Two
one-byte fields each sat at the head of a word and left the rest of it empty,
which was ten of the eighteen wasted bytes.

Grouping by size rather than by meaning takes the struct from 152 bytes to 136,
and the map holding them shrinks with it, since a Go map's slack scales with
the size of the value.

800k volumes registered from a heartbeat that has been over the wire:
211 -> 195 B/volume, 214 -> 198 tiered.

* trim the comments on this change to the parts that are not evident
2026-08-09 09:41:50 -07:00
Chris LuandGitHub 7d6c55dedb topology: build the volume list without copying the volume map first (#10668)
* topology: build the volume list without copying the volume map first

ToDiskInfo copied every VolumeInfo on the disk into a fresh slice, walked it to
build a protobuf message for each, and threw the copy away. The copy was as
large as the messages it produced.

Building them straight from the map holds the disk's read lock for the walk
rather than just the copy, so a heartbeat updating that disk waits for it. It
is a read lock on a call that is now infrequent, against an allocation of the
same size as the response.

ToTopologyInfo over 550k volumes  193617502 B/op -> 110011017 B/op, and faster
for not making the copy.

* trim the comments on this change to the parts that are not evident
2026-08-09 09:38:57 -07:00
Chris LuandGitHub e5dc98dcb2 ec.balance: add a -volumeIds filter (#10667)
* ec.balance: add a -volumeIds filter

Collection scope is often too broad for maintenance. -volumeIds narrows the
plan to the given ec volume ids by leaving every other volume out of the
topology handed to the planner, so no phase, dedup included, can plan against
them. Ids with no ec shard in the selected collection, dataCenter and disk type
are rejected rather than silently skipped.

* ec.encode: key the orphan sweep without narrowing the volume id

int is 32-bit on 32-bit builds, so int(vid) wraps for volume ids above
MaxInt32. Format the id as the uint32 it is.
2026-08-09 09:37:49 -07:00
652 changed files with 65302 additions and 10137 deletions
+3 -3
View File
@@ -27,7 +27,7 @@ jobs:
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v4.37.4
uses: github/codeql-action/init@v4.37.6
# Override language selection by uncommenting this and choosing your languages
with:
languages: go
@@ -35,7 +35,7 @@ jobs:
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
# If this step fails, then you should remove it and run the build manually (see below).
- name: Autobuild
uses: github/codeql-action/autobuild@v4.37.4
uses: github/codeql-action/autobuild@v4.37.6
# ℹ️ Command-line programs to run using the OS shell.
# 📚 See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
@@ -49,4 +49,4 @@ jobs:
# make release
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v4.37.4
uses: github/codeql-action/analyze@v4.37.6
-3
View File
@@ -31,9 +31,6 @@ jobs:
- name: Checkout
uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
+1 -4
View File
@@ -76,9 +76,6 @@ jobs:
with:
ref: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.source_ref || github.ref }}
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -367,7 +364,7 @@ jobs:
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
uses: github/codeql-action/upload-sarif@v4.37.4
uses: github/codeql-action/upload-sarif@v4.37.6
if: always()
with:
sarif_file: trivy-results.sarif
@@ -59,9 +59,6 @@ jobs:
- name: Checkout
uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -400,7 +397,7 @@ jobs:
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4.37.4
uses: github/codeql-action/upload-sarif@v4.37.6
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
+3 -6
View File
@@ -61,13 +61,10 @@ jobs:
- name: Install dependencies
run: |
# Use faster mirrors and install with timeout
sudo rm -f /etc/apt/sources.list.d/azure-cli.list /etc/apt/sources.list.d/microsoft-prod.list
echo "deb http://azure.archive.ubuntu.com/ubuntu/ $(lsb_release -cs) main restricted universe multiverse" | sudo tee /etc/apt/sources.list
echo "deb http://azure.archive.ubuntu.com/ubuntu/ $(lsb_release -cs)-updates main restricted universe multiverse" | sudo tee -a /etc/apt/sources.list
sudo apt-get update --fix-missing
sudo DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends fuse
# Same helper the e2e image installs through: the runner's own list is
# azure-only too, and an outage there fails this step outright.
sudo ./apt-install fuse
# Verify FUSE installation
echo "FUSE version: $(fusermount --version 2>&1 || echo 'fusermount not found')"
+5 -2
View File
@@ -30,7 +30,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code into the Go module directory
@@ -43,7 +43,10 @@ jobs:
- name: Run EC Integration Tests
working-directory: test/erasure_coding
run: |
go test -v
# The suite now includes the interruption matrix and runs close to Go's
# default 10m binary timeout on slower runners; bound it by the job's
# 30m budget instead.
go test -v -timeout 25m
- name: Collect server logs on failure
if: failure()
+4 -3
View File
@@ -40,10 +40,11 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install FUSE dependencies
- name: Configure FUSE
run: |
sudo apt-get update
sudo apt-get install -y libfuse3-dev
# Nothing to install: fuse3 ships fusermount3 and is pre-installed,
# and go-fuse is pure Go, so the libfuse headers were never linked
# against.
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
+76
View File
@@ -0,0 +1,76 @@
name: "FUSE Volume Server Failover Tests"
on:
pull_request:
paths:
- 'weed/command/mount*.go'
- 'weed/mount/**'
- 'weed/filer/**'
- 'weed/wdclient/**'
- 'weed/operation/upload_content.go'
- 'test/fuse_failover/**'
- '.github/workflows/fuse-failover.yml'
- '.github/actions/fix-fusermount-setuid/**'
push:
branches: [master]
paths:
- 'weed/command/mount*.go'
- 'weed/mount/**'
- 'weed/filer/**'
- 'weed/wdclient/**'
- 'weed/operation/upload_content.go'
- 'test/fuse_failover/**'
- '.github/workflows/fuse-failover.yml'
- '.github/actions/fix-fusermount-setuid/**'
concurrency:
group: ${{ github.head_ref || github.ref }}/fuse-failover
cancel-in-progress: true
permissions:
contents: read
jobs:
fuse-failover:
name: FUSE Volume Server Failover
runs-on: ubuntu-22.04
timeout-minutes: 40
steps:
- name: Check out code
uses: actions/checkout@v7
with:
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Configure FUSE
run: |
# Nothing to install: fuse3 ships fusermount3 and is pre-installed,
# and go-fuse is pure Go, so the libfuse headers were never linked
# against.
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
- name: Repair the fusermount3 setuid bit
uses: ./.github/actions/fix-fusermount-setuid
- name: Build SeaweedFS
run: go build -o weed/weed -buildvcs=false ./weed
- name: Run failover integration tests
timeout-minutes: 35
env:
WEED_BINARY: ${{ github.workspace }}/weed/weed
run: go test -v -count=1 -timeout=30m ./test/fuse_failover/...
- name: Upload logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: fuse-failover-test-logs
path: /tmp/seaweedfs-fuse-failover-logs/
retention-days: 3
+3 -5
View File
@@ -38,12 +38,10 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install FUSE and dependencies
- name: Configure FUSE
run: |
sudo apt-get update
# fuse3 is pre-installed on ubuntu-22.04 runners and conflicts
# with the legacy fuse package, so only install the dev headers.
sudo apt-get install -y libfuse3-dev
# Nothing to install: fuse3 ships fusermount3 and is pre-installed, and
# go-fuse is pure Go, so the libfuse headers were never linked against.
# Allow non-root FUSE mounts with allow_other
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
+4 -3
View File
@@ -46,10 +46,11 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install FUSE dependencies
- name: Configure FUSE
run: |
sudo apt-get update
sudo apt-get install -y libfuse3-dev
# Nothing to install: fuse3 ships fusermount3 and is pre-installed,
# and go-fuse is pure Go, so the libfuse headers were never linked
# against.
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
+12
View File
@@ -111,6 +111,16 @@ jobs:
test:
name: Test
runs-on: ubuntu-latest
services:
redis:
image: redis:8
ports:
- 6379:6379
options: >-
--health-cmd "redis-cli ping"
--health-interval 10s
--health-timeout 5s
--health-retries 5
steps:
- name: Check out code into the Go module directory
uses: actions/checkout@v7
@@ -119,6 +129,8 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Test
env:
RUN_REDIS_TESTS: "1"
run: cd weed; go test -tags "elastic gocdk sqlite ydb tarantool tikv rclone" -v ./...
test-32bit:
+77
View File
@@ -58,10 +58,52 @@ jobs:
grep -q "kind: Deployment" /tmp/s3.yaml && grep -q "seaweedfs-s3" /tmp/s3.yaml
echo "S3 deployment renders correctly"
echo "=== Testing S3 credentials from an existing secret ==="
credential_args=(
--set s3.credentials.admin.existingSecret=minio-root
--set s3.credentials.admin.accessKeyKey=root-user
--set s3.credentials.admin.secretKeyKey=root-password
--set s3.credentials.read.existingSecret=minio-root
)
for workload in \
"s3.enabled=true,s3.enableAuth=true" \
"filer.s3.enabled=true,filer.s3.enableAuth=true" \
"allInOne.enabled=true,allInOne.s3.enabled=true,allInOne.s3.enableAuth=true"
do
helm template test $CHART_DIR --set "$workload" "${credential_args[@]}" > /tmp/s3-existing-credentials.yaml
# The identities file names the variables, and each is bound to the key it was pointed at.
grep -q 'accessKey":"${SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID}' /tmp/s3-existing-credentials.yaml
grep -q 'secretKey":"${SEAWEEDFS_S3_ADMIN_SECRET_ACCESS_KEY}' /tmp/s3-existing-credentials.yaml
grep -q 'accessKey":"${SEAWEEDFS_S3_READ_ACCESS_KEY_ID}' /tmp/s3-existing-credentials.yaml
for pair in \
"SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID root-user" \
"SEAWEEDFS_S3_ADMIN_SECRET_ACCESS_KEY root-password" \
"SEAWEEDFS_S3_READ_ACCESS_KEY_ID read_access_key_id" \
"SEAWEEDFS_S3_READ_SECRET_ACCESS_KEY read_secret_access_key"
do
set -- $pair
grep -A 4 -- "- name: $1\$" /tmp/s3-existing-credentials.yaml | grep -q "name: \"minio-root\""
grep -A 4 -- "- name: $1\$" /tmp/s3-existing-credentials.yaml | grep -q "key: \"$2\""
done
# The keys stay in the user's secret rather than being copied into the chart's.
! grep -qE "^ (admin|read)_(access_key_id|secret_access_key):" /tmp/s3-existing-credentials.yaml
echo "S3 credentials reference the existing secret for $workload"
done
echo "=== Testing with all-in-one mode ==="
helm template test $CHART_DIR --set allInOne.enabled=true > /tmp/allinone.yaml
grep -q "seaweedfs-all-in-one" /tmp/allinone.yaml
echo "All-in-one deployment renders correctly"
echo "=== Testing the all-in-one s3 secret is created by every flag that mounts it ==="
for auth in allInOne.s3.enableAuth s3.enableAuth filer.s3.enableAuth; do
helm template test $CHART_DIR \
--set allInOne.enabled=true --set allInOne.s3.enabled=true --set "$auth=true" \
> /tmp/allinone-s3-auth.yaml
grep -q "secretName: test-seaweedfs-s3-secret" /tmp/allinone-s3-auth.yaml
grep -q "name: test-seaweedfs-s3-secret" /tmp/allinone-s3-auth.yaml
echo "All-in-one s3 secret renders for $auth"
done
echo "=== Testing with security enabled ==="
helm template test $CHART_DIR --set global.seaweedfs.enableSecurity=true > /tmp/security.yaml
@@ -440,6 +482,41 @@ jobs:
--set filer.s3.enableAuth=true > /tmp/filer-s3.yaml
echo "Filer S3 gateway renders correctly"
echo "=== Testing the mysql filer store gates its secret and env ==="
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MONGODB_ENABLED=true \
--set-string filer.extraEnvironmentVars.WEED_LEVELDB2_ENABLED=false > /tmp/filer-mongodb.yaml
! grep -q "db-secret" /tmp/filer-mongodb.yaml
! grep -q "WEED_MYSQL" /tmp/filer-mongodb.yaml
grep -q "name: WEED_MONGODB_ENABLED" /tmp/filer-mongodb.yaml
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MYSQL_ENABLED=true > /tmp/filer-mysql.yaml
grep -q "name: test-seaweedfs-db-secret" /tmp/filer-mysql.yaml
grep -q "name: WEED_MYSQL_USERNAME" /tmp/filer-mysql.yaml
grep -q "name: WEED_MYSQL_PASSWORD" /tmp/filer-mysql.yaml
grep -q "name: WEED_MYSQL_HOSTNAME" /tmp/filer-mysql.yaml
# Secret-backed keys follow the same rule as the plain ones.
helm template test $CHART_DIR \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.name=db \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.key=password > /tmp/filer-mysql-off-secret.yaml
! grep -q "WEED_MYSQL" /tmp/filer-mysql-off-secret.yaml
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MYSQL_ENABLED=true \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.name=db \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_PASSWORD.secretKeyRef.key=password > /tmp/filer-mysql-on-secret.yaml
grep -A 4 -- "- name: WEED_MYSQL_PASSWORD$" /tmp/filer-mysql-on-secret.yaml | grep -q "name: db"
# A flag the chart cannot read counts as selected, not as off.
helm template test $CHART_DIR \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_ENABLED.secretKeyRef.name=store \
--set filer.secretExtraEnvironmentVars.WEED_MYSQL_ENABLED.secretKeyRef.key=enabled > /tmp/filer-mysql-secret.yaml
grep -q "name: test-seaweedfs-db-secret" /tmp/filer-mysql-secret.yaml
grep -q "name: WEED_MYSQL_HOSTNAME" /tmp/filer-mysql-secret.yaml
helm template test $CHART_DIR \
--set-string filer.extraEnvironmentVars.WEED_MYSQL2_HOSTNAME=other > /tmp/filer-mysql2.yaml
grep -q "name: WEED_MYSQL2_HOSTNAME" /tmp/filer-mysql2.yaml
! grep -q "WEED_MYSQL_" /tmp/filer-mysql2.yaml
echo "The mysql secret and env follow the selected filer store"
echo "=== Testing SFTP enabled ==="
helm template test $CHART_DIR --set sftp.enabled=true > /tmp/sftp.yaml
grep -q "seaweedfs-sftp" /tmp/sftp.yaml
+1 -1
View File
@@ -42,7 +42,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
cache-dependency-path: |
**/go.sum
+14 -14
View File
@@ -43,7 +43,7 @@ jobs:
matrix:
container-id: [unit-tests-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 1.0 --memory 1g --hostname kafka-unit-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -53,7 +53,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code
@@ -87,7 +87,7 @@ jobs:
matrix:
container-id: [integration-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 2.0 --memory 2g --ulimit nofile=1024:1024 --hostname kafka-integration-${{ matrix.container-id }}
env:
GOMAXPROCS: 2
@@ -98,7 +98,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code
@@ -134,7 +134,7 @@ jobs:
matrix:
container-id: [e2e-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 2.0 --memory 2g --hostname kafka-e2e-${{ matrix.container-id }}
env:
GOMAXPROCS: 2
@@ -148,7 +148,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
cache-dependency-path: |
**/go.sum
@@ -313,7 +313,7 @@ jobs:
matrix:
container-id: [consumer-group-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 1.0 --memory 2g --ulimit nofile=512:512 --hostname kafka-consumer-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -327,7 +327,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
cache-dependency-path: |
**/go.sum
@@ -475,7 +475,7 @@ jobs:
matrix:
container-id: [client-compat-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 1.0 --memory 1.5g --shm-size 256m --hostname kafka-client-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -489,7 +489,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
cache-dependency-path: |
**/go.sum
@@ -633,7 +633,7 @@ jobs:
matrix:
container-id: [smq-integration-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 1.0 --memory 2g --hostname kafka-smq-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -647,7 +647,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
cache-dependency-path: |
**/go.sum
@@ -794,7 +794,7 @@ jobs:
matrix:
container-id: [protocol-1]
container:
image: golang:1.24-alpine
image: golang:1.26-alpine
options: --cpus 1.0 --memory 1g --tmpfs /tmp:exec --hostname kafka-protocol-${{ matrix.container-id }}
env:
GOMAXPROCS: 1
@@ -805,7 +805,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code
+203
View File
@@ -0,0 +1,203 @@
name: "mount: benchmark"
# Manual benchmark: native WinFsp mount vs rclone+WebDAV on the same Windows
# runner, with a Linux FUSE mount of the same build as a reference. Numbers
# from shared runners are noisy; this is for finding factor-of-N gaps, not
# regressions of a few percent.
on:
workflow_dispatch:
push:
branches: [ 'winfsp-bench**' ]
concurrency:
group: ${{ github.workflow }}-${{ github.head_ref || github.run_id }}
cancel-in-progress: true
permissions:
contents: read
jobs:
bench-windows:
name: Windows native vs rclone
runs-on: windows-latest
timeout-minutes: 60
env:
CGO_ENABLED: 0
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Install WinFsp and rclone
run: choco install winfsp rclone -y --no-progress
- name: Build weed.exe
run: go build -o weed.exe ./weed
- name: Benchmark both mounts
shell: pwsh
run: |
$ErrorActionPreference = 'Stop'
function Test-Port($port) {
$client = New-Object System.Net.Sockets.TcpClient
try { $client.Connect('127.0.0.1', $port); return $client.Connected }
catch { return $false }
finally { $client.Dispose() }
}
function Wait-Drive($drive, $what) {
$deadline = (Get-Date).AddMinutes(2)
while ((Get-Date) -lt $deadline) {
if (Test-Path "${drive}\") { Write-Host "$what is mounted on $drive"; return }
Start-Sleep -Seconds 2
}
throw "$what never appeared on $drive"
}
function Invoke-Bench($dir, $label, $out) {
Write-Host "::group::bench $label"
& go run ./test/mount_bench -dir $dir -label $label -filer 127.0.0.1:8888 -out $out
$code = $LASTEXITCODE
Write-Host "::endgroup::"
if ($code -ne 0) { throw "bench $label failed with exit $code" }
}
New-Item -ItemType Directory -Force -Path C:\seaweed-data | Out-Null
Start-Process -FilePath .\weed.exe `
-ArgumentList '-logtostderr','mini','-dir=C:\seaweed-data','-ip=127.0.0.1' `
-RedirectStandardOutput C:\seaweed-mini.log -RedirectStandardError C:\seaweed-mini.err.log
$deadline = (Get-Date).AddMinutes(3)
while ((Get-Date) -lt $deadline) {
if ((Test-Port 8888) -and (Test-Port 18888) -and (Test-Port 7333)) { break }
Start-Sleep -Seconds 3
}
if (-not ((Test-Port 8888) -and (Test-Port 18888) -and (Test-Port 7333))) {
Get-Content C:\seaweed-mini.log, C:\seaweed-mini.err.log -ErrorAction SilentlyContinue
throw "mini cluster never came up"
}
Write-Host "filer on 8888/18888, webdav on 7333"
# --- native WinFsp mount ---
Start-Process -FilePath .\weed.exe `
-ArgumentList '-logtostderr','mount','-filer=127.0.0.1:8888','-dir=S:' `
-RedirectStandardOutput C:\seaweed-mount.log -RedirectStandardError C:\seaweed-mount.err.log
Wait-Drive 'S:' 'weed mount'
Invoke-Bench 'S:\bench-native' 'winfsp-native' 'C:\results-native.json'
Get-CimInstance Win32_Process -Filter "Name = 'weed.exe'" |
Where-Object { $_.CommandLine -like '*mount*' } |
ForEach-Object { Stop-Process -Id $_.ProcessId -Force }
$deadline = (Get-Date).AddMinutes(1)
while ((Get-Date) -lt $deadline -and (Test-Path S:\)) { Start-Sleep -Seconds 2 }
# --- rclone + WebDAV on the same WinFsp ---
# rclone serves listings from a directory cache it fills lazily, so the
# big-listing files have to exist before it mounts or it never sees them.
& go run ./test/mount_bench -filer 127.0.0.1:8888 -seed bench-rclone/biglist
if ($LASTEXITCODE -ne 0) { throw "seeding failed" }
$env:RCLONE_CONFIG_SEAWEED_TYPE = 'webdav'
$env:RCLONE_CONFIG_SEAWEED_URL = 'http://127.0.0.1:7333'
$env:RCLONE_CONFIG_SEAWEED_VENDOR = 'other'
Start-Process -FilePath rclone `
-ArgumentList 'mount','seaweed:','T:','--vfs-cache-mode=writes','-v','--log-file=C:\rclone.log'
Wait-Drive 'T:' 'rclone mount'
Invoke-Bench 'T:\bench-rclone' 'rclone-webdav' 'C:\results-rclone.json'
Stop-Process -Name rclone -Force -ErrorAction SilentlyContinue
# --- comparison ---
$table = & go run ./test/mount_bench -compare C:\results-native.json,C:\results-rclone.json
$table | Write-Host
"## Windows: native WinFsp vs rclone+WebDAV" | Out-File -Append $env:GITHUB_STEP_SUMMARY
$table | Out-File -Append $env:GITHUB_STEP_SUMMARY
- name: Logs
if: always()
shell: pwsh
run: |
foreach ($f in 'C:\seaweed-mount.log','C:\seaweed-mount.err.log','C:\rclone.log','C:\seaweed-mini.log','C:\seaweed-mini.err.log') {
if (Test-Path $f) { Write-Host "===== $f"; Get-Content $f -Tail 100 }
}
- name: Results
if: always()
uses: actions/upload-artifact@v4
with:
name: results-windows
path: C:\results-*.json
if-no-files-found: ignore
bench-linux:
name: Linux FUSE reference
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- uses: actions/checkout@v7
with:
persist-credentials: false
- uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
- name: Repair the fusermount3 setuid bit
uses: ./.github/actions/fix-fusermount-setuid
- name: Allow non-root FUSE mounts with allow_other
run: |
echo 'user_allow_other' | sudo tee -a /etc/fuse.conf
sudo chmod 644 /etc/fuse.conf
- name: Build weed
run: go build -o /tmp/weed ./weed
- name: Benchmark FUSE mount
run: |
set -e
mkdir -p /tmp/seaweed-data
/tmp/weed -logtostderr mini -dir=/tmp/seaweed-data -ip=127.0.0.1 > /tmp/mini.log 2>&1 &
for i in $(seq 1 60); do
if nc -z 127.0.0.1 8888 && nc -z 127.0.0.1 18888; then break; fi
sleep 3
done
nc -z 127.0.0.1 8888 || { cat /tmp/mini.log; echo "filer never came up"; exit 1; }
mkdir -p "$HOME/mnt"
/tmp/weed -logtostderr mount -filer=127.0.0.1:8888 -dir="$HOME/mnt" > /tmp/mount.log 2>&1 &
for i in $(seq 1 60); do
if mountpoint -q "$HOME/mnt"; then break; fi
sleep 2
done
mountpoint -q "$HOME/mnt" || { cat /tmp/mount.log; echo "mount never appeared"; exit 1; }
go run ./test/mount_bench -dir "$HOME/mnt/bench-linux" -label linux-fuse -filer 127.0.0.1:8888 -out /tmp/results-linux.json
{
echo "## Linux FUSE reference"
go run ./test/mount_bench -compare /tmp/results-linux.json
} >> "$GITHUB_STEP_SUMMARY"
- name: Logs
if: always()
run: |
tail -n 100 /tmp/mount.log /tmp/mini.log 2>/dev/null || true
- name: Results
if: always()
uses: actions/upload-artifact@v4
with:
name: results-linux
path: /tmp/results-linux.json
if-no-files-found: ignore
-8
View File
@@ -175,10 +175,6 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
if: matrix.impl == 'rust'
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
if: matrix.impl == 'rust'
uses: dtolnay/rust-toolchain@stable
@@ -331,10 +327,6 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
if: matrix.impl == 'rust'
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
if: matrix.impl == 'rust'
uses: dtolnay/rust-toolchain@stable
+1 -1
View File
@@ -41,7 +41,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code
+133 -12
View File
@@ -1,4 +1,17 @@
name: "release: bump version"
name: "release: bump version and cut the release"
# One entry point for a SeaweedFS release:
# 1. bump MAJOR/MINOR in constants.go and the Helm Chart.yaml, commit to master
# 2. push the <appVersion> tag, which fans out to the workflows that trigger on
# `push: tags` (binaries_release*, container_release_unified, helm_manual_release)
# 3. create the GitHub release, with GitHub's generated notes
# 4. dispatch "Prepare release" in seaweedfs-csi-driver and seaweedfs-operator,
# which pick up the new master through `go get -u`, and wait for both
#
# Events raised by the default GITHUB_TOKEN do not start other workflows, and it
# cannot reach the other two repositories at all. Add a repo secret RELEASE_PAT
# with `contents: write` here and `actions: write` on the csi-driver and operator
# repos. Without it the tag is still pushed, but nothing downstream of it runs.
on:
workflow_dispatch:
@@ -14,15 +27,37 @@ on:
description: "Explicit MAJOR.MINOR to set, e.g. 4.36 (overrides 'bump')"
type: string
required: false
permissions:
contents: write
downstream:
description: "Also release the CSI driver and the operator"
type: boolean
default: true
dry_run:
description: "Show the version bump, but change nothing"
type: boolean
default: false
jobs:
bump-version:
release:
runs-on: ubuntu-latest
permissions:
contents: write
outputs:
app_version: ${{ steps.compute.outputs.app_version }}
sha: ${{ steps.tag.outputs.sha }}
steps:
- uses: actions/checkout@v7
with:
ref: master
fetch-depth: 0
token: ${{ secrets.RELEASE_PAT || secrets.GITHUB_TOKEN }}
- name: Check the release token
env:
HAS_PAT: ${{ secrets.RELEASE_PAT != '' }}
run: |
if [ "$HAS_PAT" != "true" ]; then
echo "::warning::RELEASE_PAT is not set. The tag will be pushed with GITHUB_TOKEN, so the binary, container and helm workflows will not start on their own."
fi
- name: Compute new version
id: compute
@@ -99,17 +134,103 @@ jobs:
sed -i -E "s/^version:.*/version: ${CHART_VERSION}/" "$CHART"
cat "$CHART"
- name: Commit and push
- name: Commit, and push the tag
id: tag
env:
APP_VERSION: ${{ steps.compute.outputs.app_version }}
TAG: ${{ steps.compute.outputs.app_version }}
DRY_RUN: ${{ inputs.dry_run }}
run: |
set -euo pipefail
if git diff --quiet; then
echo "No version change to commit."
if git ls-remote --exit-code --tags origin "refs/tags/${TAG}" >/dev/null 2>&1; then
echo "::error::Tag ${TAG} already exists."
exit 1
fi
if [ "$DRY_RUN" = "true" ]; then
git --no-pager diff --stat
echo "sha=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
exit 0
fi
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git add weed/util/version/constants.go k8s/charts/seaweedfs/Chart.yaml
git commit -m "${APP_VERSION}"
git push
if git diff --quiet; then
echo "::warning::Version files are already at ${TAG}; tagging the current HEAD."
else
git add weed/util/version/constants.go k8s/charts/seaweedfs/Chart.yaml
git commit -m "${TAG}"
git push
fi
# Push the tag with git so the `push: tags` triggers fire. Creating the
# tag through the release API alone would only emit a `create` event.
git tag "$TAG"
git push origin "$TAG"
echo "sha=$(git rev-parse HEAD)" >> "$GITHUB_OUTPUT"
- name: Create the release
if: ${{ !inputs.dry_run }}
env:
GH_TOKEN: ${{ secrets.RELEASE_PAT || secrets.GITHUB_TOKEN }}
TAG: ${{ steps.compute.outputs.app_version }}
run: gh release create "$TAG" --title "$TAG" --generate-notes --verify-tag
downstream:
needs: release
if: ${{ inputs.downstream && !inputs.dry_run }}
runs-on: ubuntu-latest
permissions: {}
strategy:
fail-fast: false
matrix:
include:
- repo: seaweedfs/seaweedfs-csi-driver
workflow: prepare_release.yaml
- repo: seaweedfs/seaweedfs-operator
workflow: prepare_release.yml
steps:
- name: Release ${{ matrix.repo }}
env:
GH_TOKEN: ${{ secrets.RELEASE_PAT }}
REPO: ${{ matrix.repo }}
WORKFLOW: ${{ matrix.workflow }}
SHA: ${{ needs.release.outputs.sha }}
MODULE: github.com/seaweedfs/seaweedfs
run: |
set -euo pipefail
if [ -z "${GH_TOKEN}" ]; then
echo "::error::RELEASE_PAT with actions:write on ${REPO} is required to release it"
exit 1
fi
# The dispatched workflow pins seaweedfs with `go get -u ...@latest`, so
# wait until the proxy serves the release commit as the tip. Asking for
# the commit by name is what makes the proxy fetch it.
for _ in $(seq 30); do
curl -sf "https://proxy.golang.org/${MODULE}/@v/${SHA}.info" >/dev/null || true
TIP=$(curl -sf "https://proxy.golang.org/${MODULE}/@latest" | jq -r '.Origin.Hash // ""' || true)
[ "$TIP" = "$SHA" ] && break
sleep 10
done
if [ "$TIP" != "$SHA" ]; then
echo "::error::the module proxy still serves ${TIP} as the tip, so ${REPO} would pin a pre-release commit. Run ${WORKFLOW} there once it catches up."
exit 1
fi
# Wait on the release the dispatched workflow publishes, not on the run
# that publishes it: a dispatch cannot be told apart from a concurrent
# one through the API, and the release is what we are here for.
released() { gh api "repos/${REPO}/releases?per_page=30" --jq '[.[].tag_name]'; }
BEFORE=$(released)
gh workflow run -R "$REPO" "$WORKFLOW" --ref master -f bump=patch -f update_seaweedfs=true
for _ in $(seq 80); do
sleep 15
NEW=$(released | jq -c --argjson before "$BEFORE" '. - $before')
[ "$(jq length <<<"$NEW")" -gt 0 ] && break
done
if [ "$(jq length <<<"$NEW")" -eq 0 ]; then
echo "::error::${REPO} published no release within 20 minutes; see https://github.com/${REPO}/actions/workflows/${WORKFLOW}"
exit 1
fi
echo "${REPO} released $(jq -r 'join(", ")' <<<"$NEW")"
@@ -36,9 +36,6 @@ jobs:
- name: Checkout code
uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -79,9 +76,6 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -164,9 +158,6 @@ jobs:
with:
go-version-file: 'go.mod'
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
-6
View File
@@ -41,9 +41,6 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
@@ -110,9 +107,6 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: brew install protobuf
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -28,9 +28,6 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -131,9 +128,6 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: brew install protobuf
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
with:
@@ -211,9 +205,6 @@ jobs:
steps:
- uses: actions/checkout@v7
- name: Install protobuf compiler
run: choco install protoc -y
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
+2 -2
View File
@@ -41,7 +41,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
- name: Set up Python ${{ matrix.python-version }}
@@ -148,7 +148,7 @@ jobs:
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
cache: true
- name: Run Go unit tests
+303 -3
View File
@@ -6,6 +6,7 @@ on:
- 'weed/s3api/**'
- 'weed/filer/**'
- 'weed/server/**'
- 'weed/worker/tasks/iceberg/**'
- 'test/s3tables/**'
- 'go.mod'
- 'go.sum'
@@ -354,9 +355,19 @@ jobs:
retention-days: 3
clickhouse-iceberg-catalog-tests:
name: ClickHouse Iceberg Catalog Integration Tests
name: ClickHouse Iceberg Catalog Integration Tests (${{ matrix.tag }})
runs-on: ubuntu-22.04
timeout-minutes: 30
strategy:
fail-fast: false
matrix:
include:
# Pinned baseline, and latest so new ClickHouse releases are
# exercised without a code change.
- clickhouse-image: clickhouse/clickhouse-server:25.8
tag: "25.8"
- clickhouse-image: clickhouse/clickhouse-server:latest
tag: latest
steps:
- name: Check out code
@@ -376,7 +387,7 @@ jobs:
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull clickhouse/clickhouse-server:25.8
pull ${{ matrix.clickhouse-image }}
pull python:3.11-slim
- name: Run go mod tidy
@@ -389,6 +400,8 @@ jobs:
- name: Run ClickHouse Iceberg Catalog Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_clickhouse
env:
CLICKHOUSE_IMAGE: ${{ matrix.clickhouse-image }}
run: |
set -x
set -o pipefail
@@ -422,7 +435,7 @@ jobs:
if: failure()
uses: actions/upload-artifact@v7
with:
name: clickhouse-iceberg-catalog-test-logs
name: clickhouse-iceberg-catalog-test-logs-${{ matrix.tag }}
path: test/s3tables/catalog_clickhouse/test-output.log
retention-days: 3
@@ -862,6 +875,293 @@ jobs:
path: test/s3tables/unity_catalog/test-output.log
retention-days: 3
lancedb-namespace-tests:
name: LanceDB Namespace Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 30
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull python:3.11-slim
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run LanceDB Namespace Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_lancedb
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting LanceDB Namespace Tests ==="
go test -v -timeout 20m . 2>&1 | tee test-output.log || {
echo "LanceDB namespace integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_lancedb
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: lancedb-namespace-test-logs
path: test/s3tables/catalog_lancedb/test-output.log
retention-days: 3
table-lifecycle-tests:
name: Table Lifecycle Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 40
steps:
- name: Check out code
uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull python:3.11-slim
pull duckdb/duckdb:latest
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run Table Lifecycle Integration Tests
timeout-minutes: 35
working-directory: test/s3tables/lifecycle
env:
# The Rust worker's own tests cover its handlers; a cold build of the
# lance crate costs more here than the layer it would be checking.
WEED_LANCE_MAINTENANCE: library
run: |
set -x
set -o pipefail
go test -v -timeout 30m . 2>&1 | tee test-output.log || {
echo "Table lifecycle integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/lifecycle
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: table-lifecycle-test-logs
path: test/s3tables/lifecycle/test-output.log
retention-days: 3
duckdb-lance-tests:
name: DuckDB Lance Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 30
steps:
- name: Check out code
uses: actions/checkout@v7
with:
# The job uploads a test log on failure; nothing here needs to push,
# so do not leave a token in the checkout for it to pick up.
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull duckdb/duckdb:latest
pull python:3.11-slim
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run DuckDB Lance Integration Tests
timeout-minutes: 25
working-directory: test/s3tables/catalog_duckdb_lance
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting DuckDB Lance Tests ==="
go test -v -timeout 20m . 2>&1 | tee test-output.log || {
echo "DuckDB Lance integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_duckdb_lance
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker|duckdb)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: duckdb-lance-test-logs
path: test/s3tables/catalog_duckdb_lance/test-output.log
retention-days: 3
spark-lance-namespace-tests:
name: Spark Lance Namespace Integration Tests
runs-on: ubuntu-22.04
timeout-minutes: 40
steps:
- name: Check out code
uses: actions/checkout@v7
with:
# The job uploads a test log on failure; nothing here needs to push,
# so do not leave a token in the checkout for it to pick up.
persist-credentials: false
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: 'go.mod'
id: go
- name: Configure Docker Hub mirror
run: |
echo '{"registry-mirrors": ["https://mirror.gcr.io"]}' | sudo tee /etc/docker/daemon.json
sudo systemctl restart docker
- name: Pre-pull images
run: |
pull() { for i in 1 2 3; do docker pull "$1" && return 0; sleep 15; done; return 1; }
pull apache/spark:3.5.1
- name: Run go mod tidy
run: go mod tidy
- name: Build SeaweedFS
run: |
cd weed && go build -buildvcs=false .
- name: Run Spark Lance Namespace Integration Tests
timeout-minutes: 35
working-directory: test/s3tables/catalog_spark_lance
run: |
set -x
set -o pipefail
echo "=== System Information ==="
uname -a
free -h
df -h
docker info
echo "=== Starting Spark Lance Namespace Tests ==="
go test -v -timeout 30m . 2>&1 | tee test-output.log || {
echo "Spark Lance namespace integration tests failed"
exit 1
}
- name: Show test output on failure
if: failure()
working-directory: test/s3tables/catalog_spark_lance
run: |
echo "=== Test Output ==="
if [ -f test-output.log ]; then
tail -200 test-output.log
fi
echo "=== Process information ==="
ps aux | grep -E "(weed|test|docker|spark)" || true
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: spark-lance-namespace-test-logs
path: test/s3tables/catalog_spark_lance/test-output.log
retention-days: 3
s3-tables-build-verification:
name: S3 Tables Build Verification
runs-on: ubuntu-22.04
@@ -36,7 +36,7 @@ jobs:
- uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
- name: Build SeaweedFS
run: |
+1 -1
View File
@@ -30,7 +30,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code into the Go module directory
@@ -30,7 +30,7 @@ jobs:
- name: Set up Go 1.x
uses: actions/setup-go@v7
with:
go-version: ^1.25
go-version: ^1.26
id: go
- name: Check out code into the Go module directory
+15 -12
View File
@@ -61,14 +61,14 @@ Table of Contents
* [Features](#features)
* [Additional Features](#additional-features)
* [Filer Features](#filer-features)
* [Example: Using Seaweed Object Store](#example-using-seaweed-object-store)
* [Example: Using Seaweed Blob Store](#example-using-seaweed-blob-store)
* [Architecture](#object-store-architecture)
* [Compared to Other File Systems](#compared-to-other-file-systems)
* [Compared to HDFS](#compared-to-hdfs)
* [Compared to GlusterFS, Ceph](#compared-to-glusterfs-ceph)
* [Compared to GlusterFS](#compared-to-glusterfs)
* [Compared to Ceph](#compared-to-ceph)
* [Compared to Minio](#compared-to-minio)
* [Compared to MinIO, RustFS](#compared-to-minio-rustfs)
* [Dev Plan](#dev-plan)
* [Installation Guide](#installation-guide)
* [Disk Related Topics](#disk-related-topics)
@@ -90,7 +90,7 @@ S3_BUCKET=my-bucket \
./weed mini -dir=/data
```
That's it — the S3 endpoint is at http://localhost:8333, `my-bucket` already exists, and `admin`/`secret` are valid credentials. `S3_BUCKET` accepts a comma-separated list (e.g. `raw,processed`); use `S3_TABLE_BUCKET` for S3 Tables (Iceberg) buckets. Drop any of the env vars to skip that piece (no AWS keys → S3 runs in unauthenticated "Allow All" mode for development).
That's it — the S3 endpoint is at http://localhost:8333, `my-bucket` already exists, and `admin`/`secret` are valid credentials. `S3_BUCKET` accepts a comma-separated list (e.g. `raw,processed`); use `S3_TABLE_BUCKET` for S3 Tables buckets, each `name` or `name:FORMAT` where the format is `ICEBERG` (the default) or `LANCE`. Drop any of the env vars to skip that piece (no AWS keys → S3 runs in unauthenticated "Allow All" mode for development).
The same command starts everything else too:
- **S3 Endpoint**: http://localhost:8333
@@ -466,7 +466,8 @@ The architectures are mostly the same. SeaweedFS aims to store and read files fa
| GlusterFS | hashing | | FUSE, NFS | | |
| Ceph | hashing + rules | | FUSE | Yes | |
| MooseFS | in memory | | FUSE | | No |
| MinIO | separate meta file for each file | | | Yes | No |
| MinIO | separate meta file per drive for each file | | | Yes | No |
| RustFS | separate meta file per drive for each file | | | Yes | No |
[Back to TOC](#table-of-contents)
@@ -508,24 +509,26 @@ SeaweedFS Filer uses off-the-shelf stores, such as MySql, Postgres, Sqlite, Mong
[Back to TOC](#table-of-contents)
### Compared to MinIO ###
### Compared to MinIO, RustFS ###
Please note, as Apr 25, 2026 MinIO ceased development. It's strongly discouraged to use that unmaintained software with multiple security bugs.
Please note, as Apr 25, 2026 MinIO ceased development. It's strongly discouraged to use that unmaintained software with multiple security bugs. RustFS is a MinIO reimplementation in Rust, Apache 2.0 licensed and still developed, keeping MinIO's storage model down to a byte-compatible on-disk format. So the points below apply to both.
MinIO followed AWS S3 closely and was ideal for testing for S3 API. It had good UI, policies, versionings, etc. SeaweedFS is trying to catch up here.
MinIO metadata were in simple files. Each file write will incur extra writes to corresponding meta file.
The metadata are in simple files. Each file write incurs extra writes to the corresponding meta file, on every drive of the erasure set. Changing only tags or retention rewrites that meta file on all of them, so the write amplification does not shrink with object size.
MinIO did not have optimization for lots of small files. The files were simply stored as is to local disks.
There is no optimization for lots of small files. The files are simply stored as is to local disks.
Plus the extra meta file and shards for erasure coding, it only amplifies the LOSF problem.
MinIO had multiple disk IO to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.
Multiple disk IO are needed to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.
MinIO had full-time erasure coding. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.
Erasure coding is full-time. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.
MinIO did not have POSIX-like API support.
No POSIX-like API support.
MinIO had specific requirements on storage layout. It is not flexible to adjust capacity. In SeaweedFS, just start one volume server pointing to the master. That's all.
There are specific requirements on storage layout, which makes it hard to scale out and to maintain. An erasure set must be 2 to 16 drives and must divide the drive list symmetrically, and capacity grows or shrinks a whole pool at a time. In SeaweedFS, just start one volume server pointing to the master. That's all.
[Back to TOC](#table-of-contents)
## Dev Plan ##
+696
View File
@@ -0,0 +1,696 @@
# Lance Catalog for SeaweedFS
A second catalog surface next to the Iceberg REST catalog, speaking the Lance Namespace
REST spec, over the same table buckets and the same filer.
## Why
Gravitino 1.1 added a Lance REST service and 1.3 ships it as a standalone server; Lakekeeper
added Lance in the same window by a completely different route. That is the useful signal:
two unrelated catalogs decided independently that Lance had to be first-class, not a niche.
The client side is already there — `lance-spark` (`LanceNamespaceSparkCatalog`
with `impl=rest`), `lance-ray`, and the generated Python/Java/Rust clients all talk the same
OpenAPI. Implementing the spec means those engines work against SeaweedFS with no
SeaweedFS-specific code on the client.
The second reason is that Gravitino's own documentation names the gap it cannot close:
DuckDB, pandas and DataFusion "do not support Lance REST natively yet" and have to fetch a
location from the catalog and then open the dataset directly. Gravitino cannot help there,
because it does not own the storage. SeaweedFS does. That is the whole design opportunity
below.
## Prior art: three families
Upstream lists twelve catalog implementations, and they fall into three shapes. Knowing
which one we are building matters more than any individual API decision.
**1. Storage-native, no service.** The Lance Directory Catalog. V1 is a directory listing
where every `<name>.lance/` child of a prefix is a table; V2 adds a `__manifest` table —
itself a Lance table — holding `object_id`/`object_type`/`location` rows, with nested
namespaces, hash-prefixed table directories, and optional managed versioning. No server, no
credentials, no governance. This is the floor every other implementation has to beat.
**2. Protocol-native server.** Someone implements the Lance Namespace REST OpenAPI and
clients connect with `impl=rest`. Gravitino is the only one of the twelve that does this,
and it is what this design proposes.
**3. Client-side adapters onto an existing catalog.** Nine of the twelve. The Lance client
translates namespace operations into whatever the backing catalog already speaks: Apache
Polaris, Unity Catalog, AWS Glue, Hive Metastore v2 and v3, Google BigLake, Dataproc,
Microsoft OneLake — and Apache Iceberg REST. Two flavors:
- Catalogs with a real non-Iceberg table concept mark the format directly. Polaris uses its
Generic Table API with `format = lance`; Unity uses an `EXTERNAL` table with
`table_type=lance` in properties and the path in `storage_location`; Glue uses
`EXTERNAL_TABLE` plus `table_type=lance` in `Parameters`, path in
`StorageDescriptor.Location`.
- Catalogs with no such concept fake one. The Iceberg REST adapter registers **a regular
Iceberg table with a dummy schema — a single nullable string column named `dummy`** —
carrying the property `table_type=lance`, and treats the Iceberg table location as the
Lance dataset root.
Every adapter in family 3 lands in the same place: `DeclareTable`/`ListTables`/
`DescribeTable`/`DeregisterTable` only, `DropNamespace` in RESTRICT mode only,
`load_detailed_metadata=false` only, and `managed_versioning=false`. They are a name-to-
location map and nothing more.
Lakekeeper is the instructive outlier. It has the same generic-table concept Polaris has,
but no upstream adapter exists for it — there is no `lance-namespace` reference anywhere in
its repository and no page for it in the supported-catalogs list. So Polaris's generic tables
are reachable from a stock Lance client and Lakekeeper's are not, despite being the same
idea. Shipping the concept is not the same as shipping the integration.
## Gravitino and Lakekeeper: the two opposite bets
Both shipped Lance support in the same window and did not build the same thing.
**Gravitino implements the protocol.** Its `lance/` module serves the Lance Namespace REST
spec on its own port (`:9101/lance`), so stock `lance-spark` and `lance-ray` connect with
`impl=rest` and no vendor-specific client. The cost is governance: storage credentials are
static properties on the catalog (`lance.storage.access_key_id`, `secret_access_key`,
`endpoint`, `region`, `allow_http`), optionally overridden per table, handed to the engine
as-is. No STS, no expiry, no per-table scoping.
**Lakekeeper refuses the protocol and governs the object instead.** There is no
`lance-namespace` anywhere in the repository; Lance arrived in 0.13.0 (2026-06-30, issue
#1673 `Generic Table API with Lance`) as one `format` string on a Lakekeeper-native Generic
Table API:
```
POST/GET/DELETE /lakekeeper/v1/{prefix}/namespaces/{ns}/generic-tables[/{table}]
GET /lakekeeper/v1/{prefix}/namespaces/{ns}/generic-tables/{table}/credentials
POST /lakekeeper/v1/{prefix}/generic-tables/rename
```
`format` is opaque, `schema` and `statistics` are stored but never validated, and the
catalog writes no format-specific metadata — engines go straight to the location. In
exchange Lance tables get everything Iceberg tables get: STS-vended prefix-scoped
credentials, OpenFGA per-action permissions (16 actions), soft-delete with undrop, a
protection flag, rename, pagination, and name uniqueness across Iceberg tables, views and
generic tables in one namespace. The price is that no stock Lance client can talk to it —
you need `pylakekeeper`, which exists mainly to translate vended credentials into
`lance_storage_options`.
So: protocol fidelity and weak governance, or strong governance and client lock-in. Both
documented their limit honestly, and it is the same limit. Lakekeeper's capability table
says it outright — "Commit coordination: the catalog does not arbitrate writes — engines
write directly." Gravitino does not claim it either. Neither of them coordinates a Lance
commit, which is exactly the thing a store can do and a control plane cannot.
We do not have to choose. Serve the Lance protocol natively the way Gravitino does, over
the `s3tables` entries that already carry ARNs, policies, tags and maintenance config, and
the governance comes from the layer underneath rather than from a proprietary API on top.
That is only available to us because we are the store, which is also what makes the third
option — arbitrating the commit — available.
## We are probably already a Lance catalog, and that is a problem
The Iceberg REST adapter does not care whose Iceberg catalog it is talking to. It needs
`/v1/config?warehouse=`, `/v1/{prefix}/namespaces`, `/v1/{prefix}/namespaces/{ns}/tables`
and unit-separator (`\x1F`) multi-level namespaces. We serve all of those, and
`parseNamespace` in `weed/s3api/iceberg/utils.go:22` already splits on `\x1F`. So a stock
Lance client pointed at our Iceberg catalog on :8181 with the Iceberg impl should already
create, list, describe and deregister Lance tables today, with no SeaweedFS change at all.
That is worth testing before writing a line of the design above, for two reasons. It is a
free baseline — and possibly a free announcement. And it is a data-loss hazard.
A Lance table registered this way is an Iceberg table whose metadata references no data
files, sitting on top of a Lance dataset that uses `data/` for its fragments — the same
subdirectory name Iceberg uses. The maintenance worker's orphan cleaner walks exactly
`<table>/metadata` and `<table>/data`, and deletes every file not referenced by a snapshot
and older than `orphan_older_than_hours`
(`weed/worker/tasks/iceberg/operations.go:331`, default 72). Against an adapter-registered
Lance table, every fragment is unreferenced by construction. Run maintenance and the
dataset is deleted.
Maintenance is disabled by default (`handler.go:334`), so this is a latent hazard rather
than a live one: it needs an operator to enable Iceberg maintenance on a bucket that also
holds adapter-registered Lance tables. But it costs nothing to close — detection should
skip any table carrying a non-Iceberg format marker (`table_type` property, or
`Format != "ICEBERG"` once the format field is honest), and that guard is worth landing on
its own regardless of whether the rest of this design ever gets built. It is the same
"catalog-only, no maintenance" marker the generic-format question needs.
## Where we differ from Gravitino
Gravitino is a metadata service in front of somebody else's object store:
```
Spark / Ray Spark / Ray / pandas / duckdb
| |
Lance REST Lance REST (direct S3)
| | |
Gravitino SeaweedFS S3 gateway ----+
| |
S3 keys handed out SeaweedFS filer + volumes
|
somebody else's S3
```
It resolves a name to a location plus `lance.storage.*` credentials, and steps out of the
way. Everything a Lance table actually is — `_versions/`, `data/`, `_indices/` — is opaque
to it.
We are the store. Three things follow that Gravitino cannot do:
1. The catalog and a plain directory listing can be made to agree, so a client with no
catalog at all still sees the right tables.
2. `_versions/` is a filer directory listing, not an object-store `LIST`. Version history
is cheap and can back the admin UI.
3. We can offer a genuinely atomic commit reservation. Lance's commit protocol needs
put-if-not-exists; our S3 layer does not currently provide one (see
[Commit safety](#commit-safety)). The filer does.
## Placement
The Iceberg catalog is a thin HTTP shell over `s3tables.Manager`; the storage work lives in
`weed/s3api/s3tables`. Table buckets live under `TablesPath = s3_constants.DefaultBucketsPath`,
i.e. the same filer tree the S3 gateway serves, so `s3://bucket/ns/table/` is simultaneously
a catalog entry and an S3 prefix. Catalog entries are filer directories carrying `s3tables.*`
extended attributes. `Table.Format` already exists and is hard-checked against `"ICEBERG"`
in `weed/s3api/s3tables/handler_table.go:48`.
So:
```
weed/s3api/lance/ new: HTTP surface, id codec, error model
weed/s3api/s3tables/ extended: Format "LANCE", lance state xattr, version entries
weed/command/s3.go new: -port.lance (default 9101), startLanceServer
```
`Format: "LANCE"` on the table entry is the whole storage-model change for phase 1.
Everything else — namespaces, ARNs, policies, tags, ownership — is shared verbatim.
```
s3tables.Manager (filer)
|
+------------------------+------------------------+
| |
weed/s3api/iceberg weed/s3api/lance
Iceberg REST :8181 Lance REST :9101
| |
Iceberg tables Lance datasets
\ /
+-------------------- s3 :8333 -----------------+
|
SeaweedFS volumes
```
## Identifier mapping
Lance identifiers are `["ns", ..., "table"]`, encoded in the URL as a single string joined
by a delimiter that defaults to `$`. The delimiter alone means the root namespace, so
`/v1/namespace/$/list` lists the root's children.
Iceberg had to invent a warehouse selector because its identifier is flat and every table
bucket is a separate catalog. Lance does not need that — its identifier is already
hierarchical, and Gravitino uses exactly three levels (`["lance_catalog", "sales", "orders"]`).
That maps onto us without inventing anything:
```
$ root -> list of table buckets
$analytics level 1 -> a table bucket
$analytics$sales level 2 -> a namespace in that bucket
$analytics$sales$orders table
```
`spark.sql.catalog.lance.parent = analytics` then makes `sales.orders` resolve, which is the
same shape Gravitino's Spark example uses.
Levels 2..N join into one `s3tables` namespace with `.`, matching what the Iceberg catalog
already does with `flattenNamespacePath`. The flattened form is only the directory name —
`namespaceMetadata.Namespace []string` in the xattr keeps the authoritative parts, so the
mapping stays invertible even though `.` is a legal character inside a namespace part.
Reject `$` in any name part with `InvalidInput`; our charsets already exclude it, so no
escaping scheme is needed.
Root-level `ListNamespaces` returning table buckets means an unauthenticated or
broadly-scoped caller can enumerate buckets. Filter it through the same
`s3tables/permissions.go` check `ListTableBuckets` uses, not a separate path.
`CreateNamespace` on a one-part identifier creates a table bucket, and it does so only if
the caller is permitted to — the namespace never creates a bucket as a side effect of
creating something inside it. A table bucket is a tenant resource with its own policy, ARN
and lifecycle, and conjuring one because a client said `CREATE SCHEMA` is a privilege
escalation dressed as a convenience. Lakekeeper draws the same line explicitly: its client
creates tables, not warehouses.
## Storage layout
Lay tables out as:
```
s3://<table-bucket>/<flattened-namespace>/<table>/
data/
_versions/
_indices/
```
**Built without the `.lance` suffix this design originally proposed.** The suffix would have
made every namespace prefix a valid Lance Directory Catalog V1 root, since V1 recognises a
table by exactly that naming. It does not survive contact with the storage layer: the
catalog entry *is* the dataset directory, `validateTableName` excludes `.` from the charset,
and a suffixed entry name would leak into ARNs, policy documents and the S3 Tables API,
where the same table would answer to two different names. Making `GetTablePath` format-aware
instead spreads an "unless it is Lance" branch through code that has no business knowing —
the exact cross-cutting cost this design rejects family 3 for.
So one name, one directory. What survives is direct access by URI, which is the larger half
of the story and needs no naming convention at all:
```python
# with the catalog
spark.sql("SELECT * FROM lance.sales.orders")
# without it, same bytes
lance.dataset("s3://analytics/sales/orders")
```
DuckDB, pandas and DataFusion still reach the data with no catalog running, which is the gap
Gravitino's documentation admits to. What they no longer get for free is *enumeration* — a
directory-catalog client pointed at the namespace prefix will not list these as tables. If
that turns out to matter, the cheapest fix is a repair-style tool that materialises `.lance`
aliases, not a rename of the catalog entry.
Note also that the directory catalog's own V2 mode puts child-namespace tables in
`<hash>_<ns$table>` directories at the root and creates no physical subdirectories for
namespaces, so full directory-catalog fidelity was never on offer anyway. We are a
server-backed catalog; the human-readable prefix layout is worth more than partial V1
lookalike behaviour.
## The table bucket was not a neutral container
This design assumed a table bucket is a place to put a table's files. It is
not: `validateTableBucketObjectPath` runs on every S3 write into one and
validated the path against Iceberg's layout, so a Lance client got 403 on
`data/*.lance`, on `_versions/`, and on `_transactions/` — a directory Lance
writes that neither the spec documentation nor this design anticipated. Nothing
about the catalog worked end to end until that changed.
The layout guard now admits the union of what the supported formats write, and
treats any underscore-prefixed top-level directory as belonging to the format,
checking only that the path stays inside the table. Enumerating Lance's
internal directories by name is exactly the mistake that missed
`_transactions`. Iceberg writes none of them, so it loses nothing.
Found by pointing the real Python client at a running gateway, not by reading
the spec. Worth remembering for the next format: the premise to check first is
whether the storage layer will accept its files at all.
## Table lifecycle
Lance has three table states, and the spec pins them to marker files:
| State | Marker | Created by | Visible in ListTables |
| --- | --- | --- | --- |
| declared | `.lance-reserved` | `DeclareTable` | yes, when `include_declared=true` |
| created | `_versions/` present | client writes, or `CreateTable` | yes |
| deregistered | `.lance-deregistered` | `DeregisterTable` | no; data preserved |
Record the state in an xattr (`s3tables.lanceState`) on the catalog entry *and* write the
marker file into the table directory. The xattr is what the catalog reads; the marker is
what keeps a directory-catalog client honest. Dual-write is the price of the interop claim
above, and it is one extra filer write on three rarely-called operations.
`DeclareTable` is the operation `lance-spark` actually calls on `CREATE TABLE` (it replaced
the legacy `create-empty`), so it is not optional in practice even though the spec marks
only a subset as required.
`DeregisterTable` preserving data is the same shape as our Iceberg rename, where the catalog
entry moves and the data stays put — reuse `TableDataDirFromMetadataLocation`'s idea rather
than re-deriving the data path from the catalog name.
## Commit safety
This is the part I got wrong, and the correction removed a feature rather than adding one.
Lance commits a version by writing `_versions/{v}.manifest` with put-if-not-exists: exactly
one writer is supposed to win, and the loser rebases. In lance 10 that path is not optional
and needs nothing bolted on — `commit_handler_from_url` hands every `s3://` dataset a
`ConditionalPutCommitHandler`, which calls `put_opts` with `PutMode::Create`, which
object_store's S3 backend sends as `If-None-Match: *`.
I originally read our gateway as evaluating that header check-then-act, and designed around
it. That was already out of date. `buildWriteCondition`
(`weed/s3api/s3api_object_routed_write.go`) reduces `If-None-Match: *` to a filer
`WriteCondition{IF_NOT_EXISTS}`, and `putToFiler` routes the create to the object's owner
filer, which evaluates the precondition under its per-path lock; when routing is not
available it falls back to the object write lock, which evaluates it under the lock too.
Either way it is atomic. Sixteen concurrent writers of the same fresh key get one 200 and
fifteen 412s, repeatedly.
So the store already has the primitive Lance needs, cluster-wide, for every conditional-PUT
client and not just this one.
### What that removed
An earlier draft of this design offered the catalog as an **external manifest store**:
`managed_versioning: true` plus `CreateTableVersion` and friends, with the reserve step as a
filer `CreateEntry` with `o_excl`. It was implemented, tested, and shipped behind a default-off
flag — and it should not exist.
- It solves a problem this store does not have. The spec offers that path for stores that
cannot order commits themselves.
- It moves a table's version history out of the dataset and into the catalog, so a reader
that does not go through this namespace no longer sees the whole picture. That is a real
cost paid for nothing.
- lance 10 cannot even use it past the first commit: `NamespaceManifestStore::put_if_not_exists`
answers "put_if_not_exists is not supported for namespace-backed stores", which is exactly
what a second `append` needs.
The version operations now answer `Unsupported` alongside the other operations the catalog
does not serve, and `managed_versioning` is answered `false`. The property they were
protecting is covered instead by a test that races eight writers at the manifest key through
S3 and asserts one wins — testing the path Lance actually takes.
## Credential vending
Iceberg needed a header (`X-Iceberg-Access-Delegation: vended-credentials`) and a bespoke
response shape. Lance has it in the spec: `vend_credentials: true` on the request,
`storage_options` on the response, with `expires_at_millis` as the well-known expiry key.
Reuse the existing vendor interface unchanged — `iceberg.CredentialVendor` /
`STSService.AssumeRoleForPrincipal` scoped to the table prefix (#10777) — and map its output
to the storage options Lance passes through to `object_store`:
```
aws_access_key_id, aws_secret_access_key, aws_session_token,
aws_region, aws_endpoint, allow_http, expires_at_millis
```
Those are the names `pylakekeeper` emits as `lance_storage_options`, which is the shape
Lakekeeper's tested S3 path actually feeds to Lance. `object_store` also accepts the
un-prefixed aliases (`endpoint`, `region`) that the directory catalog's `storage.` prefix
strips down to and that Gravitino's `lance.storage.endpoint` resolves to, but the `aws_`
forms are the ones with a tested integration behind them, so emit those. `aws_endpoint`
should come from `deriveS3AdvertisedEndpoint()`, the same source the Iceberg `FileIO` config
uses, and `allow_http` must be set when that endpoint is plain HTTP or every read fails with
a TLS error that looks like a credential problem — Lakekeeper vends both automatically for
exactly this reason, and calls out that there is then no per-vendor branch in client code.
We emit this server-side, in the `storage_options` field the Lance spec already defines,
which is strictly better than Lakekeeper's arrangement: no client library has to translate
anything, so vending works from any stock Lance client rather than only from theirs.
Guard the same way #10777 had to after review: bucket-scoped list grants need an `s3:prefix`
condition, and a location containing `*` or `?` must be refused rather than widened into a
resource pattern.
## Auth and authorization
Authentication reuses `S3Authenticator` and `CredentialValidator` as-is. The Lance spec maps
identity to headers — `api_key` to `x-api-key`, `auth_token` to `Authorization: Bearer` — and
SigV4 keeps working because it is the same authenticator the Iceberg catalog already fronts.
Authorization needs nothing new. A Lance table gets the same ARN shape,
`arn:aws:s3tables:...:bucket/B/table/NS/T`, so every existing table-bucket policy covers
Lance tables with no new policy language and no second permission model. Route it through
`s3tables/permissions.go` and inherit the `DefaultAllow` semantics the Iceberg server already
mirrors from the S3 port.
One spec quirk worth honoring: request context entries prefixed `header.` become request
headers, and every response header comes back as a `header.`-prefixed context entry. Echoing
`x-request-id` through it costs nothing and makes tracing work.
## What to take from Lakekeeper
Rejecting Lakekeeper's API shape does not mean rejecting what it learned building it.
**Deregister is soft-delete, so implement it as one.** Lakekeeper gives generic tables
soft-deletion with undrop and a `protected` flag that makes a drop require `force=true`.
Lance already has the concept — `DeregisterTable` preserves the data and hides the table —
so the `.lance-deregistered` marker is a soft-delete by another name, and a re-register is
an undrop. A protection flag on table-bucket entries is worth having regardless of Lance:
it is a few lines against the existing xattrs and it applies to Iceberg tables too.
**Enforce one identifier space across entry kinds.** Lakekeeper rejects a generic table
whose name collides with an Iceberg table or view in the same namespace. Our catalog entries
already share one filer directory and already carry `s3tables.entryType`, so this is
structurally true — but it has to be enforced deliberately on every path, or a Lance handler
happily loads an Iceberg table's directory and vice versa. That is the same crossover bug
class as the view/table rename authorization fixed in #10776; the `catalogEntryKind` pattern
from that change is the thing to reuse rather than re-derive.
**A re-vend path matters more than it looks.** Lakekeeper exposes `/credentials` separately
from load, because STS credentials expire in the middle of long jobs and re-loading the
whole table to refresh them is wasteful. In Lance the spec's answer is another
`DescribeTable` with `vend_credentials: true`, which is fine — but it means `DescribeTable`
must stay cheap when `load_detailed_metadata` is false, which is another reason not to open
the dataset on that path.
**Generic tables are a cheap orthogonal win.** Lakekeeper's real insight is that Delta,
Parquet, CSV, Vortex and Paimon all get governance for free once the catalog stops caring
what the format is. Our `Table.Format` field already exists and the only thing stopping it
is the hard `"ICEBERG"` check in `handler_table.go:48`. Loosening that and letting the S3
Tables API register a table with an arbitrary format and a location — no metadata, no
commits — is a small change that makes every format cataloguable. It is independent of this
design and probably worth doing first, since `Format: "LANCE"` is then just a value rather
than a special case.
**Skip remote signing.** It is Lakekeeper's fallback for S3-compatible stores with no STS,
and their own documentation notes that Lance will not use it — format libraries with their
own S3 client expect static credentials and do not implement the Iceberg signer protocol. We
have STS, so vended credentials are the path, and the signer is not worth building for a
client that cannot consume it.
## Errors
Lance uses `{code, error, detail, instance}` with numeric codes, not Iceberg's exception-type
strings. The mapping is mechanical:
| HTTP | code | when |
| --- | --- | --- |
| 400 | 13 InvalidInput | charset violations, malformed id, route/body mismatch |
| 401 | 16 Unauthenticated | |
| 403 | 15 PermissionDenied | |
| 404 | 1 NamespaceNotFound, 4 TableNotFound, 11 TableVersionNotFound | |
| 409 | 2/5 AlreadyExists, 3 NamespaceNotEmpty, 14 ConcurrentModification | |
| 501 | 0 Unsupported | every phase-3 data operation |
Route/body mismatch is a spec requirement, not a nicety: when the identifier appears in both
the path and the body and they disagree, the server must return 400. Cheap to get right at
the decode step, annoying to retrofit.
## Route surface
Phase 0 is not in this table: point a stock Lance client at the existing Iceberg catalog
with the Iceberg impl, see how far it gets, and land the maintenance guard either way. That
tells us what the native server actually has to beat.
Phase 1, the whole `lance-spark` and `lance-ray` contract:
```
POST /v1/namespace/{id}/create CreateNamespace mode: Create|ExistOk|Overwrite
GET /v1/namespace/{id}/list ListNamespaces
POST /v1/namespace/{id}/describe DescribeNamespace
POST /v1/namespace/{id}/drop DropNamespace mode: Fail|Skip, behavior: Restrict|Cascade
POST /v1/namespace/{id}/exists NamespaceExists
GET /v1/namespace/{id}/table/list ListTables ?include_declared, ?page_token, ?limit
GET /v1/table ListAllTables
POST /v1/table/{id}/declare DeclareTable
POST /v1/table/{id}/describe DescribeTable ?with_table_uri, ?load_detailed_metadata, ?check_declared
POST /v1/table/{id}/exists TableExists
POST /v1/table/{id}/register RegisterTable mode: Create|Overwrite
POST /v1/table/{id}/deregister DeregisterTable
POST /v1/table/{id}/drop DropTable
POST /v1/table/{id}/rename RenameTable
```
`DescribeTable` with `load_detailed_metadata=false` needs only `location`, which is the
common case and which we can answer from xattrs alone. With `load_detailed_metadata=true`
the spec wants `version`, `schema` and `stats`, which means reading the Lance manifest. For
phase 1, return the fields we can derive from the filer — `version` from the highest entry in
`_versions/`, given V2 naming is `{u64::MAX - version:020}.manifest` and V1 is
`{version}.manifest` — and omit `schema`/`stats` rather than fabricating them. The spec
tolerates a partial response here; it does not tolerate a wrong one.
Phase 2 was the five version operations plus `managed_versioning`; it was built and then
removed, for the reasons under Commit safety.
Phase 3 is the data plane: `CreateTable`, `InsertIntoTable`, `MergeInsertIntoTable`,
`UpdateTable`, `DeleteFromTable`, `QueryTable`, `CountTableRows`, and the index and tag
operations. These exchange Arrow IPC, and more to the point they require reading and writing
the Lance file format, for which no Go implementation exists. Return `Unsupported` (code 0)
and say so in the docs. `arrow-go/v18` is already an indirect dependency, so Arrow framing is
not the blocker — Lance is.
## Does a Lance table need maintenance?
Yes, and one part of it has no Iceberg equivalent. The client exposes three jobs:
- `optimize.compact_files()` — Lance writes a fragment per write batch, so a table fed by
small appends accumulates small files exactly the way an Iceberg table does.
- `optimize.optimize_indices()` — **rows written after an index was built are not covered by
it.** A vector search against a stale index silently misses recent data. That is a
correctness-shaped failure, not a slow query, and it is specific to what people use Lance
for.
- `cleanup_old_versions()` — every version is retained until something removes it. Lance can
do this itself: `optimize.enable_auto_cleanup()` sets it on the dataset, so this one need
not be an external job at all.
None of it can run in the Go worker. All three read and rewrite Lance files, which needs
Lance format code that does not exist in Go, and there is no useful subset either: deciding
which fragments an old version still references means parsing Lance manifests.
So the maintenance worker must not touch a Lance table, and it declines by reading the format
the catalog recorded rather than by failing to parse Iceberg metadata.
## The worker can be Rust, and it is not a sidecar
The Go worker is not the only worker. `weed/pb/plugin.proto` defines `PluginControlService`,
a language-agnostic gRPC stream that external maintenance workers connect on: the worker
opens `WorkerStream`, sends `WorkerHello` with the job types it can `detect` and `execute`,
answers `RequestConfigSchema` with a `JobTypeDescriptor`, replies to `RunDetectionRequest`
with `JobProposal`s and to `ExecuteJobRequest` with `JobProgressUpdate`s and `JobCompleted`.
`weed worker -admin=host:23646` is the Go reference implementation of exactly that contract,
from outside the admin process.
Nothing in it is Go-specific, and the Rust toolchain is already in the tree.
`seaweed-volume/build.rs` compiles protos straight out of `../weed/pb/` with `tonic_build`,
including `filer.proto`, on tonic 0.12 and prost 0.13. A Lance worker is that same build
with `plugin.proto` added and the `lance` crate as a dependency — the real one, no FFI and
no Python.
Three job types, one per real maintenance operation:
| Job type | Calls | Detected from |
| --- | --- | --- |
| `lance_compact` | `optimize.compact_files` | fragment count and sizes |
| `lance_optimize_indices` | `optimize.optimize_indices` | rows an index does not cover |
| `lance_cleanup_versions` | `cleanup_old_versions` | version count and age |
What the existing machinery then supplies for free is the part worth noticing. Scheduling,
retries, dedupe by `dedupe_key`, progress reporting, per-job concurrency limits and the
admin settings page all come from the protocol: a worker that answers `RequestConfigSchema`
with a descriptor gets its configuration form rendered in the admin UI without a line of Go
or templ. A Rust worker is a first-class maintenance worker, not an appendage.
The remaining wiring is small and mostly decided already. `RunDetectionRequest` carries a
`ClusterContext` with filer and S3 addresses plus a free-form `metadata` map, which is where
the Lance namespace URL goes; the worker lists Lance tables from the namespace, which is the
catalog of record and already filters by format. It gets at the data by asking
`DescribeTable` for `storage_options` with `vend_credentials`, so the worker is just another
client of the STS path rather than a component with its own credentials. And when it commits
a compaction it goes through `CreateTableVersion` like any other writer, which is what
managed versioning was for.
## The worker is also the only thing that can describe the table
Admin can render an Iceberg table because it can read Iceberg metadata. It cannot read
Lance: it knows the dataset's location and its format string, and that is the whole of it.
The details page showed a location and two empty panels, which is an honest answer and a
useless one.
The worker already knows. Detection opens every dataset to decide whether it needs
compacting, so at that moment it holds the schema, the row count, the fragment count and
the version count. It just had no way to say so — every message on the stream was about
work.
So `WorkerObservations` is a body on `WorkerToAdminMessage`: a repeated `ObjectObservation`
of `object_id`, `object_kind`, `format`, and a `ConfigValue` map the worker fills with
whatever it can cheaply say. Admin keeps the last observation per object and serves it back
with the time it was taken and the worker that took it. Nothing schedules from it, and it is
not authoritative — it is a cache with its staleness on the label, which is why the page
badges it rather than presenting it as metadata it read itself.
The keys are the worker's to choose, which keeps the protocol out of the business of knowing
what a Lance table is. A worker for any other format admin cannot parse describes itself the
same way.
## A bucket declares its format
Format was recorded per table, which is enough for the storage layer and not enough for
anything that has to answer a question about a bucket. The admin UI printed one Iceberg
endpoint for every bucket, including the ones holding Lance datasets, where that endpoint
serves nothing; an empty bucket had no format at all.
So `CreateTableBucket` takes an optional `format`, stored with the rest of the bucket
metadata. Empty means `ICEBERG` - what AWS S3 Tables serves, and therefore what an SDK
that has never heard of the field means. `CreateTable` refuses another format, and
`CreateView` refuses outright outside an Iceberg bucket, a view being Iceberg metadata.
The Lance namespace declares `LANCE` for the buckets it creates.
**Enforced rather than defaulted**, because the point of showing a format at all is the
endpoint that follows from it, and that endpoint is only truthful if the bucket holds one
format. **Buckets that already exist stay undeclared** and keep taking anything: nothing is
migrated, and the UI shows "unset" as a fact about the bucket's age rather than a fault.
That state is also the only way to hold both formats at once, which is what the
Iceberg-REST adapter path produces.
## Sample rows are fetched, not cached
The same asymmetry has a second half. Admin renders an Iceberg table's rows by
reading its Parquet files directly; for Lance it has nothing to read with, so the data
page offered a Browse Data button that led to an empty grid.
`RequestObjectPreview` / `ObjectPreviewResponse` mirror the config-schema round trip
already on the stream: admin asks, the worker scans the dataset and hands back rows it
has already rendered as text, because it is the only side that knows the types. Admin
picks the worker from the observation store, so the one that last described a table is
the one asked to read it.
The rows are deliberately not cached, and that is the line between the two channels. An
observation describes an object, so a copy with a timestamp on it is useful. Rows are the
object's contents: a copy held in admin would be stale, larger, and nobody's business.
The page fetches on load, bounded, or says why it cannot.
## The sidecar question
The data plane is a different problem, and this design previously conflated the two.
Maintenance rides the worker protocol; `QueryTable` and `InsertIntoTable` do not, because
they are synchronous REST operations on the namespace's own surface. Serving those means a
Rust process that answers HTTP, either behind the Go namespace as a proxy target or in front
of it. It would make SeaweedFS a store you can run vector search *in* rather than one you
read vectors *out of*, which is the larger prize and the reason to keep the option open.
Neither should gate phase 1. Phases 1 and 2 are pure Go over the filer and are worth
shipping on their own — they are what makes Spark and Ray work.
## Testing
Mirror the Iceberg package: `httptest` plus a fake filer client for the handler tests, in
`weed/s3api/lance`. Then an integration suite under `test/s3tables/catalog/` next to the
existing `pyiceberg_test.go`, driving the generated Python `lance-namespace` client against
a live gateway. Three things that suite must cover and unit tests cannot:
- the storage-options key names actually work, i.e. a client that gets `storage_options` from
`DescribeTable` can open the dataset;
- a table created through the catalog is visible to `lance.dataset()` by URI and to a V1
directory-catalog client rooted at the namespace prefix;
- concurrent writers do not lose a commit, which is the phase-2 acceptance test and the
thing that justifies the external manifest store.
Phase 1 is validated: `lance_namespace` 0.11.1 with `impl=rest` drives the namespace,
`lance.write_dataset` writes to the vended location with the vended `storage_options`, and
the rows read back. Note that this client version drops `check_declared` and
`include_declared` on the wire, so `is_only_declared` reads null through it however the
server behaves.
The commit path is validated at both levels. The mechanism: eight writers race the same
manifest key through S3 with `If-None-Match: *`, and exactly one wins. The property that
actually matters, which single-winner exclusivity does not by itself establish: eight
writers append to one dataset concurrently through lance, and afterwards every batch is
still there — the losers saw the conflict, rebased, and committed again. That second test
is also the sequence managed versioning could not complete at all, since its store answers
"put_if_not_exists is not supported" to the second commit.
One more that belongs in the Iceberg suite, not this one: a Lance dataset registered through
the Iceberg adapter must survive a full maintenance pass. Reading the code, that test should
fail today; it has not been run.
## Open questions
- Root-level `ListNamespaces` enumerating table buckets is convenient and is a listing
surface we do not have on the Iceberg side. Decide whether it is gated behind a flag.
- Whether the `.lance` directory suffix is worth the divergence from the Iceberg layout. I
think yes — it is what makes the catalog optional — but it means the two catalogs' tables
do not look alike on disk, and the admin UI has to know that.
- Names: our charsets are lowercase-only and Lance identifiers are arbitrary strings. Reject
and document, as Iceberg does, or case-fold. Rejecting is right, but see #10734 for how
case handling bites when only one side normalizes.
- Whether to land generic-format registration first. Dropping the `"ICEBERG"` check and
letting a table carry an arbitrary format plus a location is smaller than this whole
design, gets Delta and Parquet catalogued as a side effect, and turns `Format: "LANCE"`
into an ordinary value. The argument against is that it invites tables the maintenance
worker cannot service, so it needs a "catalog-only, no maintenance" marker to be honest.
+11 -15
View File
@@ -2,25 +2,21 @@ FROM ubuntu:22.04
LABEL author="Chris Lu"
# Use Azure's Ubuntu mirror — much faster than archive.ubuntu.com from GitHub-hosted runners,
# which have been hanging long enough on Ign:/retry to trip the 10-min step timeout.
# Note: This e2e test image intentionally runs as root for simplicity and compatibility.
# Production images (Dockerfile.go_build) use proper user isolation with su-exec.
# For testing purposes, running as root avoids permission complexities and dependency
# on Alpine-specific tools like su-exec (not available in Ubuntu repos).
RUN sed -i 's|http://archive.ubuntu.com/ubuntu|http://azure.archive.ubuntu.com/ubuntu|g; s|http://security.ubuntu.com/ubuntu|http://azure.archive.ubuntu.com/ubuntu|g' /etc/apt/sources.list && \
apt-get -o Acquire::http::Timeout=15 update && \
DEBIAN_FRONTEND=noninteractive apt-get -o Acquire::http::Timeout=15 install -y \
--no-install-recommends \
--no-install-suggests \
curl \
fio \
fuse \
ca-certificates \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/* \
&& rm -rf /tmp/* \
&& rm -rf /var/tmp/*
# apt-install prefers Azure's mirror, which archive.ubuntu.com is slow enough from
# GitHub-hosted runners to justify, and falls through to archive.ubuntu.com when
# Azure is unreachable - which it periodically is, and Acquire::Retries against a
# single mirror just retries a dead host. Images built FROM this one install
# through it for the same reason.
COPY apt-install /usr/local/bin/apt-install
RUN chmod +x /usr/local/bin/apt-install && \
apt-install curl fio fuse ca-certificates && \
rm -rf /tmp/* /var/tmp/*
RUN mkdir -p /etc/seaweedfs /data/filerldb2
COPY ./weed /usr/bin/
+1 -1
View File
@@ -1,4 +1,4 @@
FROM golang:1.25 AS builder
FROM golang:1.26 AS builder
RUN apt-get update && \
apt-get install -y build-essential wget ca-certificates && \
+1 -1
View File
@@ -1,6 +1,6 @@
# Pin the builder to the host arch and cross-compile the (CGO-free) Go binary,
# so arm64/arm/386 targets skip QEMU emulation of the whole compile.
FROM --platform=$BUILDPLATFORM golang:1.25-alpine AS builder
FROM --platform=$BUILDPLATFORM golang:1.26-alpine AS builder
RUN apk add git g++ fuse
RUN mkdir -p /go/src/github.com/seaweedfs/
ARG BRANCH=${BRANCH:-master}
+1 -1
View File
@@ -1,4 +1,4 @@
FROM golang:1.25 AS builder
FROM golang:1.26 AS builder
RUN apt-get update
RUN apt-get install -y build-essential libsnappy-dev zlib1g-dev libbz2-dev libgflags-dev liblz4-dev libzstd-dev
+1 -1
View File
@@ -1,4 +1,4 @@
FROM golang:1.25 AS builder
FROM golang:1.26 AS builder
RUN apt-get update
RUN apt-get install -y build-essential libsnappy-dev zlib1g-dev libbz2-dev libgflags-dev liblz4-dev libzstd-dev
+25
View File
@@ -0,0 +1,25 @@
#!/bin/sh
# Install packages, falling through to the next Ubuntu mirror when one is
# unreachable. See docker/Dockerfile.e2e.
set -e
# Every rewrite starts from the pristine list, so a mirror that just failed does
# not become the pattern the next rewrite has to match.
[ -f /etc/apt/sources.list.orig ] || cp /etc/apt/sources.list /etc/apt/sources.list.orig
for mirror in azure.archive.ubuntu.com archive.ubuntu.com; do
# Any archive host, so this works whether the pristine list came from the
# base image (archive.ubuntu.com) or a CI runner (azure.archive.ubuntu.com).
sed "s|http://[a-z0-9.]*archive\.ubuntu\.com/ubuntu|http://$mirror/ubuntu|g; s|http://security\.ubuntu\.com/ubuntu|http://$mirror/ubuntu|g" \
/etc/apt/sources.list.orig > /etc/apt/sources.list
if apt-get -o Acquire::Retries=3 -o Acquire::http::Timeout=15 update && \
DEBIAN_FRONTEND=noninteractive apt-get -o Acquire::Retries=3 -o Acquire::http::Timeout=15 install -y \
--no-install-recommends --no-install-suggests "$@"; then
apt-get clean
rm -rf /var/lib/apt/lists/*
exit 0
fi
echo "apt: $mirror unreachable, trying the next mirror" >&2
done
exit 1
+5 -1
View File
@@ -11,4 +11,8 @@ scrape_configs:
- 'master:9324'
- 'volume:9325'
- 'filer:9326'
- 's3:9327'
- 's3:9327'
# A plugin worker publishes its own metrics when started with
# -metricsPort (Go) or --metrics-port (Rust). 9328 continues the
# series, since 9327 is already the S3 gateway's.
# - 'worker:9328'
+34 -34
View File
@@ -1,6 +1,6 @@
module github.com/seaweedfs/seaweedfs
go 1.25.8
go 1.26
require (
cloud.google.com/go v0.123.0 // indirect
@@ -12,7 +12,7 @@ require (
github.com/bwmarrin/snowflake v0.3.0
github.com/cenkalti/backoff/v4 v4.3.0
github.com/coreos/go-semver v0.3.1 // indirect
github.com/coreos/go-systemd/v22 v22.6.0 // indirect
github.com/coreos/go-systemd/v22 v22.7.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/dustin/go-humanize v1.0.1
github.com/eapache/go-resiliency v1.6.0 // indirect
@@ -49,7 +49,7 @@ require (
github.com/kurin/blazer v0.5.3
github.com/linxGnu/grocksdb v1.10.8
github.com/mailru/easyjson v0.9.2 // indirect
github.com/mattn/go-isatty v0.0.23 // indirect
github.com/mattn/go-isatty v0.0.24 // indirect
github.com/modern-go/concurrent v0.0.0-20180306012644-bacd9c7ef1dd // indirect
github.com/modern-go/reflect2 v1.0.2 // indirect
github.com/olivere/elastic/v7 v7.0.32
@@ -84,22 +84,22 @@ require (
github.com/xdg-go/scram v1.2.0
github.com/xdg-go/stringprep v1.0.4 // indirect
github.com/youmark/pkcs8 v0.0.0-20240726163527-a2c0da244d78 // indirect
go.etcd.io/etcd/client/v3 v3.6.12
go.etcd.io/etcd/client/v3 v3.7.1
go.mongodb.org/mongo-driver v1.17.9
go.opencensus.io v0.24.0 // indirect
gocloud.dev v0.46.0
gocloud.dev/pubsub/natspubsub v0.46.0
gocloud.dev/pubsub/rabbitpubsub v0.46.0
golang.org/x/crypto v0.54.0
golang.org/x/crypto v0.55.0
golang.org/x/exp v0.0.0-20260709172345-9ea1abe57597
golang.org/x/image v0.44.0
golang.org/x/net v0.57.0
golang.org/x/net v0.58.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.47.0
golang.org/x/text v0.40.0 // indirect
golang.org/x/text v0.41.0 // indirect
golang.org/x/tools v0.48.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.289.0
google.golang.org/api v0.293.0
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94 // indirect
google.golang.org/grpc v1.84.0-dev.0.20260723093437-b6eac429d7b6
google.golang.org/protobuf v1.36.11
@@ -107,11 +107,11 @@ require (
modernc.org/b v1.0.0 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
modernc.org/sqlite v1.53.0
modernc.org/sqlite v1.56.0
)
require (
cloud.google.com/go/kms v1.31.0
cloud.google.com/go/kms v1.33.0
github.com/Azure/azure-sdk-for-go/sdk/keyvault/azkeys v0.10.0
github.com/DATA-DOG/go-sqlmock v1.5.2
github.com/Jille/raft-grpc-transport v1.6.1
@@ -122,9 +122,9 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.43.2
github.com/aws/aws-sdk-go-v2 v1.43.5
github.com/aws/aws-sdk-go-v2/config v1.32.33
github.com/aws/aws-sdk-go-v2/credentials v1.19.32
github.com/aws/aws-sdk-go-v2/credentials v1.19.34
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2
github.com/cespare/xxhash/v2 v2.3.0
github.com/cognusion/imaging v1.0.4
@@ -141,9 +141,9 @@ require (
github.com/linkedin/goavro/v2 v2.15.0
github.com/minio/crc64nvme v1.1.1
github.com/orcaman/concurrent-map/v2 v2.0.1
github.com/parquet-go/parquet-go v0.30.1
github.com/parquet-go/parquet-go v0.32.0
github.com/pkg/sftp v1.13.11
github.com/rabbitmq/amqp091-go v1.11.0
github.com/rabbitmq/amqp091-go v1.13.0
github.com/rclone/rclone v1.75.0
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.21.0
@@ -156,8 +156,8 @@ require (
github.com/tikv/client-go/v2 v2.0.7
github.com/xeipuuv/gojsonschema v1.2.0
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2
github.com/ydb-platform/ydb-go-sdk/v3 v3.146.3
go.etcd.io/etcd/client/pkg/v3 v3.6.12
github.com/ydb-platform/ydb-go-sdk/v3 v3.147.1
go.etcd.io/etcd/client/pkg/v3 v3.7.1
go.uber.org/atomic v1.11.0
golang.org/x/sync v0.22.0
golang.org/x/tools/godoc v0.1.0-deprecated
@@ -183,7 +183,7 @@ require (
github.com/antlr4-go/antlr/v4 v4.13.1 // indirect
github.com/apache/arrow-go/v18 v18.7.0 // indirect
github.com/apache/thrift v0.24.0 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.5.2 // indirect
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4 // indirect
github.com/bahlo/generic-list-go v0.2.0 // indirect
github.com/bazelbuild/rules_go v0.46.0 // indirect
github.com/biogo/store v0.0.0-20201120204734-aad293a2328f // indirect
@@ -291,7 +291,7 @@ require (
require (
cel.dev/expr v0.25.2 // indirect
cloud.google.com/go/auth v0.20.0 // indirect
cloud.google.com/go/auth v0.23.0 // indirect
cloud.google.com/go/auth/oauth2adapt v0.2.8 // indirect
cloud.google.com/go/compute/metadata v0.9.0 // indirect
cloud.google.com/go/iam v1.11.0 // indirect
@@ -322,21 +322,21 @@ require (
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.33 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.33 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.33 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.34 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.14 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.33 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.33.2 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.2 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.45.2
github.com/aws/smithy-go v1.27.5
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 // indirect
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 // indirect
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4
github.com/aws/smithy-go v1.27.7
github.com/boltdb/bolt v1.3.1 // indirect
github.com/bradenaw/juniper v0.15.3 // indirect
github.com/buengese/sgzip v0.1.1 // indirect
@@ -383,12 +383,12 @@ require (
github.com/gogo/protobuf v1.3.2 // indirect
github.com/golang-jwt/jwt/v4 v4.5.2 // indirect
github.com/google/s2a-go v0.1.9 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.18 // indirect
github.com/googleapis/enterprise-certificate-proxy v0.3.20 // indirect
github.com/gorilla/schema v1.4.1 // indirect
github.com/gorilla/securecookie v1.1.2 // indirect
github.com/gorilla/sessions v1.4.0
github.com/grpc-ecosystem/go-grpc-middleware v1.4.0 // indirect
github.com/grpc-ecosystem/grpc-gateway/v2 v2.28.0 // indirect
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 // indirect
github.com/hashicorp/go-cleanhttp v0.5.2 // indirect
github.com/hashicorp/go-hclog v1.6.3 // indirect
github.com/hashicorp/go-immutable-radix v1.3.1 // indirect
@@ -436,7 +436,7 @@ require (
github.com/pelletier/go-toml/v2 v2.2.4 // indirect
github.com/pengsrc/go-shared v0.2.1-0.20190131101655-1999055a4a14 // indirect
github.com/philhofer/fwd v1.2.0 // indirect
github.com/pierrec/lz4/v4 v4.1.27
github.com/pierrec/lz4/v4 v4.1.28
github.com/pingcap/errors v0.11.5-0.20211224045212-9687c2b0f87c // indirect
github.com/pingcap/failpoint v0.0.0-20220801062533-2eaa32854a6c // indirect
github.com/pingcap/kvproto v0.0.0-20230403051650-e166ae588106 // indirect
@@ -482,7 +482,7 @@ require (
github.com/zeebo/blake3 v0.2.4 // indirect
github.com/zeebo/errs v1.4.0 // indirect
go.etcd.io/bbolt v1.5.0 // indirect
go.etcd.io/etcd/api/v3 v3.6.12 // indirect
go.etcd.io/etcd/api/v3 v3.7.1 // indirect
go.opentelemetry.io/auto/sdk v1.2.1 // indirect
go.opentelemetry.io/contrib/detectors/gcp v1.44.0 // indirect
go.opentelemetry.io/contrib/instrumentation/google.golang.org/grpc/otelgrpc v0.68.0 // indirect
@@ -497,12 +497,12 @@ require (
golang.org/x/term v0.45.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260715232425-e75dac1f907d // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260807164820-c8921c73eeea // indirect
gopkg.in/natefinch/lumberjack.v2 v2.2.1 // indirect
gopkg.in/validator.v2 v2.0.1 // indirect
gopkg.in/yaml.v2 v2.4.0 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
modernc.org/libc v1.73.4 // indirect
modernc.org/libc v1.74.4 // indirect
moul.io/http2curl/v2 v2.3.0 // indirect
sigs.k8s.io/yaml v1.6.0 // indirect
storj.io/common v0.0.0-20260629224719-ba1bff0a7846 // indirect
+74 -74
View File
@@ -94,8 +94,8 @@ cloud.google.com/go/assuredworkloads v1.7.0/go.mod h1:z/736/oNmtGAyU47reJgGN+KVo
cloud.google.com/go/assuredworkloads v1.8.0/go.mod h1:AsX2cqyNCOvEQC8RMPnoc0yEarXQk6WEKkxYfL6kGIo=
cloud.google.com/go/assuredworkloads v1.9.0/go.mod h1:kFuI1P78bplYtT77Tb1hi0FMxM0vVpRC7VVoJC3ZoT0=
cloud.google.com/go/assuredworkloads v1.10.0/go.mod h1:kwdUQuXcedVdsIaKgKTp9t0UJkE5+PAVNhdQm4ZVq2E=
cloud.google.com/go/auth v0.20.0 h1:kXTssoVb4azsVDoUiF8KvxAqrsQcQtB53DcSgta74CA=
cloud.google.com/go/auth v0.20.0/go.mod h1:942/yi/itH1SsmpyrbnTMDgGfdy2BUqIKyd0cyYLc5Q=
cloud.google.com/go/auth v0.23.0 h1:6Gg1CMgpgubRG7DGz5Vf1pcoNo8RfiRiRAPS4crTp54=
cloud.google.com/go/auth v0.23.0/go.mod h1:4DhBRcqvtljQN3dJ57qtqbib5ZGCYE5f2crfiiC2EM0=
cloud.google.com/go/auth/oauth2adapt v0.2.8 h1:keo8NaayQZ6wimpNSmW5OPc283g65QNIiLpZnkHRbnc=
cloud.google.com/go/auth/oauth2adapt v0.2.8/go.mod h1:XQ9y31RkqZCcwJWNSx2Xvric3RrU88hAYYbjDWYDL+c=
cloud.google.com/go/automl v1.5.0/go.mod h1:34EjfoFGMZ5sgJ9EoLsRtdPSNZLcfflJR39VbVNS2M0=
@@ -298,8 +298,8 @@ cloud.google.com/go/kms v1.4.0/go.mod h1:fajBHndQ+6ubNw6Ss2sSd+SWvjL26RNo/dr7uxs
cloud.google.com/go/kms v1.5.0/go.mod h1:QJS2YY0eJGBg3mnDfuaCyLauWwBJiHRboYxJ++1xJNg=
cloud.google.com/go/kms v1.6.0/go.mod h1:Jjy850yySiasBUDi6KFUwUv2n1+o7QZFyuUJg6OgjA0=
cloud.google.com/go/kms v1.9.0/go.mod h1:qb1tPTgfF9RQP8e1wq4cLFErVuTJv7UsSC915J8dh3w=
cloud.google.com/go/kms v1.31.0 h1:LS8N92OxFDgOLg5NCo3OmbvjtQAIVT5gUHVLKIDHaFE=
cloud.google.com/go/kms v1.31.0/go.mod h1:YIyXZym11R5uovJJt4oN5eUL3oPmirF3yKeIh6QAf4U=
cloud.google.com/go/kms v1.33.0 h1:pG0X78m212b2pv9N4fdMoUO69LuZGQ9kSvn8sHBOFAo=
cloud.google.com/go/kms v1.33.0/go.mod h1:CSGvW6GnMQbY+1nOHcIzhMtHSbExXlOmCKjWtYVjcpA=
cloud.google.com/go/language v1.4.0/go.mod h1:F9dRpNFQmJbkaop6g0JhSBXCNlO90e1KWx5iDdxbWic=
cloud.google.com/go/language v1.6.0/go.mod h1:6dJ8t3B+lUYfStgls25GusK04NLh3eDLQnWM3mdEbhI=
cloud.google.com/go/language v1.7.0/go.mod h1:DJ6dYN/W+SQOjF8e1hLQXMF21AkH2w9wiPzPCJa2MIE=
@@ -707,50 +707,50 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.43.2 h1:cl+IXwWb3qazClUcm08tGSsB6OiuV83JVJO9B0jQcPc=
github.com/aws/aws-sdk-go-v2 v1.43.2/go.mod h1:WEzLKBh/mEjXvx1FtQMWgSxMSTVqxQzjkRtk5fa3wkg=
github.com/aws/aws-sdk-go-v2 v1.43.5 h1:yKT5GYnFWhuDo+DqKvE5ZPwVn3RjC4MAeBtZGlh6AVM=
github.com/aws/aws-sdk-go-v2 v1.43.5/go.mod h1:wZjAJppCntyOGgVSmgVTfDyRJK5PHOasO6Wsy8U7Axk=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14 h1:3IZY0XAJquT3aHzbkHfPzy4ACPcEjVG0x87KOwtpqGY=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.14/go.mod h1:zwM6veDkhGgQFqkBy+uT28AAYpLu+uFMlPl+rCg/73E=
github.com/aws/aws-sdk-go-v2/config v1.32.33 h1:M1m/Q6f0OKDEDGwhiNOqx1OjTdrewe3v+GDbHmKczWk=
github.com/aws/aws-sdk-go-v2/config v1.32.33/go.mod h1:fGj1iQj2QpIZzp7jE4aQQ+71TE8cd4z9K4+xCd6EqmE=
github.com/aws/aws-sdk-go-v2/credentials v1.19.32 h1:eNE0JnIblBo1NCvd3tqEYuZz9XDefn69R74CHd3nT7U=
github.com/aws/aws-sdk-go-v2/credentials v1.19.32/go.mod h1:yYJu+6tqKUYZuJSYcpSGjz/6sV/SUaAaKIufnWKx2OU=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.33 h1:MobhiR6KIerWxmO74Zit5I3379+mSc2DOdZ3DeRFB9w=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.33/go.mod h1:xu02847OdZfNr/jAfZpHtyRk0b3v4d0kaoxNHxZGG/w=
github.com/aws/aws-sdk-go-v2/credentials v1.19.34 h1:y6GkSmcv5myd1ngrYbGmiLlwQqB6TQhOuN/tbSSuWDY=
github.com/aws/aws-sdk-go-v2/credentials v1.19.34/go.mod h1:w3dTcnDVoQIewjo7JG45hduAToikiIFLC4FIO7fndvw=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35 h1:+S7kbJoLDDQ5tE+lHrUBgMkzC8NLgsaioS2F3dVoFAE=
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.35/go.mod h1:Ak7xXviIARfFdNUJ9Etb0bdVDt/KAvKjMGJVLWXDzik=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34 h1:Pn7OsMwBLbkZ6OnCxWHAjf0L/22H8cnhxZC0uPwtMtg=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.34/go.mod h1:eToXR/Gk1uqpn04eSmdgVXwfS0WvH8aG4eBFr8ygbpU=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.2.3 h1:w5OoDiMN6x53ROmiIImGzmVcxXv2q1GXY+aKV4WAJYM=
github.com/aws/aws-sdk-go-v2/feature/s3/transfermanager v0.2.3/go.mod h1:dAhgYp776bX3LuWvnSCFwQEjNs6fuFg7YXIy5PXcP3Q=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.33 h1:HAp1wLFZzch054uh3FK7rcVYg4v7J2FxVf3h3IGNZas=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.33/go.mod h1:mJk5fmqnF+WUlMdPG37pR2Fh3oh6r8F6ZGUgPKvzu0c=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.33 h1:0YA0aCKgsJyno6xkFfaIgjE3/wK08+Qxo9nQfe1UrWM=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.33/go.mod h1:UZqj4WIdTH+ga8Y/DgpAuy/8cGjM3h7gDCliJYGg2SE=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.34 h1:HQYnjFnXpX8EbPW5M1QT8mXzesRPwly0HEPTcFlS02Y=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.34/go.mod h1:tGzj56niKYZBbDIRhwPGDqrULzmWv5b6uBQGqyNaFZw=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.14 h1:SA43nfaY7+1jjMNIc2ywu99JLJLButtIdLP6j+bT870=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.14/go.mod h1:Du3llKcwbQvHsTXSLzTOGQz0DTDBMEzdg7DAGu7inrY=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35 h1:kzVuGlatQtYinwBJEEyLAbggepCoavosiaHHX9+fD+c=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.35/go.mod h1:0yLx0yEI+SfqeJMPvOtIEFoZbiQYXMGszBueiutQyaI=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35 h1:WK6CjihTuLisCjSKKbildJ79sGZZgbBz3iNa7VsKIhU=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.35/go.mod h1:KYleN57luLoe97R7vTnx8PMcVrr9gAcRECtOjl91DNg=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36 h1:jbGY4CXLzZElOXgGsexlC3Hi+3YM0rSmk4opFXKqg/k=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.36/go.mod h1:uBu/9aKsS/UQGc72RAt3y54kjgYQxmhut8ZD2dXCDNE=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15 h1:JJLBQxwY+AFwuPAi5ivGc1ChnTdUt4cXMv7e76m2c/Y=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.15/go.mod h1:lQknBIe78MVL0cQOQDlag8KGflMbMEVFx9mB6O8ENvk=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23 h1:9Fjh6fi/U5JEStVZijmaMpUwE/gvBJj7x2B/PjbO9To=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.23/go.mod h1:iMoT2f1tClxrWAAnKCXjZQ6LOmfLrMG14wmnWpM+F14=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.33 h1:mqI7OrxN/DUH85F5OqVn3cIfuZ3+HVcebUm2N8mLlgQ=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.33/go.mod h1:eZ5jdEpvaaOU8nWWE4cTAJETSEA5FZoWxvNRao4piHY=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35 h1:BBEElKh4a+rKshvjrfpajTe9CbpZvrbb4Jkg2PB7RzA=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.35/go.mod h1:zaZk983w//8beSruBVec/mr4CmDwgZitW/qzGhAAX0g=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31 h1:uao4A3QZ5UmB326V6KF+qRpv9Tjz7IlnlnTbbANntlU=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.31/go.mod h1:I/1+z0VwL1GhQyLgkoHDlygpUZ+iTAwOQ/NsftiUL2I=
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2 h1:5C00eQYpTrgQXnp6V3P6P7zPElna3AXvlukbANE6nJI=
github.com/aws/aws-sdk-go-v2/service/s3 v1.105.2/go.mod h1:zdmCoFO/dSI7GlrwsPqFJI+WlFnSU4Tc8TJnlXrM1Do=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.2 h1:EjI1CZzDcBxPkTa3j1BdtIrUDbqnOGssFMeyUS+6W0I=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.2/go.mod h1:vN3eb5H8MEAZ4dx0F5Wc9LT8eb3eW7bZZ5BjGJdbw9k=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4 h1:cOJELVNrq5Q3Udry2GLuHUM7MhwpeaQRdYaoa6GI/yI=
github.com/aws/aws-sdk-go-v2/service/signin v1.5.4/go.mod h1:f4LxzKBtaTxD7xh3PiVg3CE1tchQemfmghaJr+NbK2c=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14 h1:p8WdWDh5AwSZdp19Haa3XMyPCICi9Z375a/Nu3IIEZY=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.14/go.mod h1:NKVY7DER6VXHkt2I/ycmHakALNboi3Rqwt4eEf/1Cnk=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24 h1:JP2wjWGmUp8lTCZb13Dv0Eciyc1jbO8pd0HZVMHFlrc=
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.24/go.mod h1:Ql9ziDutk8ERAN9HMaYANCW3lop451ppebkxEJMLCTM=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.2 h1:zMP1FDFE08L7sM5f1QqkH/ZgKKg8Uc0Dz7KhSSYqWkw=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.2/go.mod h1:0LoIZSUKjdo2BleHfT1hv/jlD33LQS00IrBlzoUsoUQ=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.2 h1:9eTqUYl+SyVmaRPMyBXSO9wwqC6TRwZB82pKENK2hdQ=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.2/go.mod h1:DThweuz22kiLc7lGHop5vQ9c3bx5W6Azs/YqSHa2fu8=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.2 h1:EJd8vZO3E8SE6nmPqxuxlQ1NeSb8as50sf6eGdV4Saw=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.2/go.mod h1:OgpPvKzsO2Ranjpli/20djMkg6UrV5mw4W3pZpq1Mqo=
github.com/aws/smithy-go v1.27.5 h1:d1ro7KpYOYwP6m73YFa+Kc/A130VsAdX68SpsJwARMM=
github.com/aws/smithy-go v1.27.5/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4 h1:AMW7a7S8iQaHjBYZdU3PCq4GKRPijTPRAc7e6XtEThY=
github.com/aws/aws-sdk-go-v2/service/sso v1.33.4/go.mod h1:QQNsFV1DVXoXcZt18FS8lI8rtUrlDyAuWZLQ5shunv4=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4 h1:AsbZcJAQPRmHDJG8K1N0pof/1zPWjVT8TFlTWuGLSvo=
github.com/aws/aws-sdk-go-v2/service/ssooidc v1.38.4/go.mod h1:6imqztH0//t0mKbl6yWl7swSEl7F/w32oAmqB3vP1ag=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4 h1:w/AryDYMjSUANSQ2uoZxJovUsMTwWJNTv3IMex30Y+4=
github.com/aws/aws-sdk-go-v2/service/sts v1.45.4/go.mod h1:WeBiAa67azG7Su9Vf+ChGDBLiAozJCXzdjXiPBUwtbc=
github.com/aws/smithy-go v1.27.7 h1:Zgj5z4LfcDYoQIVk+n/yGdTkP/2y6ZT5vYxe0fp7bqE=
github.com/aws/smithy-go v1.27.7/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc=
github.com/bahlo/generic-list-go v0.2.0 h1:5sz/EEAK+ls5wF+NeqDpk5+iNdMDXrh3z3nPnH1Wvgk=
github.com/bahlo/generic-list-go v0.2.0/go.mod h1:2KvAjgMlE5NNynlg/5iLrrCCZ2+5xWbdbCW3pNTGyYg=
github.com/bazelbuild/rules_go v0.46.0 h1:CTefzjN/D3Cdn3rkrM6qMWuQj59OBcuOjyIp3m4hZ7s=
@@ -885,8 +885,8 @@ github.com/containerd/typeurl/v2 v2.2.3 h1:yNA/94zxWdvYACdYO8zofhrTVuQY73fFU1y++
github.com/containerd/typeurl/v2 v2.2.3/go.mod h1:95ljDnPfD3bAbDJRugOiShd/DlAAsxGtUBhJxIn7SCk=
github.com/coreos/go-semver v0.3.1 h1:yi21YpKnrx1gt5R+la8n5WgS0kCrsPp33dmEyHReZr4=
github.com/coreos/go-semver v0.3.1/go.mod h1:irMmmIw/7yzSRPWryHsK7EYSg09caPQL03VsM8rvUec=
github.com/coreos/go-systemd/v22 v22.6.0 h1:aGVa/v8B7hpb0TKl0MWoAavPDmHvobFe5R5zn0bCJWo=
github.com/coreos/go-systemd/v22 v22.6.0/go.mod h1:iG+pp635Fo7ZmV/j14KUcmEyWF+0X7Lua8rrTWzYgWU=
github.com/coreos/go-systemd/v22 v22.7.0 h1:LAEzFkke61DFROc7zNLX/WA2i5J8gYqe0rSj9KI28KA=
github.com/coreos/go-systemd/v22 v22.7.0/go.mod h1:xNUYtjHu2EDXbsxz1i41wouACIwT7Ybq9o0BQhMwD0w=
github.com/cosmos/go-bip39 v1.0.0 h1:pcomnQdrdH22njcAatO0yWojsUnCO3y2tNoV1cb6hHY=
github.com/cosmos/go-bip39 v1.0.0/go.mod h1:RNJv0H/pOIVgxw6KS7QeX2a0Uo0aKUlfhZ4xuwvCdJw=
github.com/cpuguy83/dockercfg v0.3.2 h1:DlJTyZGBDlXqUZ2Dk2Q3xHs/FtnooJJVaad2S9GKorA=
@@ -1236,8 +1236,8 @@ github.com/google/pprof v0.0.0-20210226084205-cbba55b83ad5/go.mod h1:kpwsk12EmLe
github.com/google/pprof v0.0.0-20210601050228-01bbb1931b22/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20210609004039-a478d1d731e9/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20210720184732-4bb14d4b1be1/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20250317173921-a4b03ec1a45e h1:ijClszYn+mADRFY17kjQEVQ1XRhq2/JR1M3sGqeJoxs=
github.com/google/pprof v0.0.0-20250317173921-a4b03ec1a45e/go.mod h1:boTsfXsheKC2y+lKOCMpSfarhxDeIzfZG1jqGcPl3cA=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3 h1:LMLX+LgTNWpfvCBdFebv6EsYotImrt/Ppc5cXIriCSo=
github.com/google/pprof v0.0.0-20260802141513-ef3492d7dac3/go.mod h1:jl5iWTm0/hd5PjEYEOuwAJ57L/CibdZfrqZ5XA5GrCk=
github.com/google/renameio v0.1.0/go.mod h1:KWCgfxg9yswjAJkECMjeO8J8rahYeXnNhOm40UhjYkI=
github.com/google/s2a-go v0.1.9 h1:LGD7gtMgezd8a/Xak7mEWL0PjoTQFvpRudN895yqKW0=
github.com/google/s2a-go v0.1.9/go.mod h1:YA0Ei2ZQL3acow2O62kdp9UlnvMmU7kA6Eutn0dXayM=
@@ -1255,8 +1255,8 @@ github.com/googleapis/enterprise-certificate-proxy v0.1.0/go.mod h1:17drOmN3MwGY
github.com/googleapis/enterprise-certificate-proxy v0.2.0/go.mod h1:8C0jb7/mgJe/9KK8Lm7X9ctZC2t60YyIpYEI16jx0Qg=
github.com/googleapis/enterprise-certificate-proxy v0.2.1/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.2.3/go.mod h1:AwSRAtLfXpU5Nm3pW+v7rGDHp09LsPtGY9MduiEsR9k=
github.com/googleapis/enterprise-certificate-proxy v0.3.18 h1:hvVi34VucdrV1IIsiWuqYM8kutw/92MxNEFxCJZEh0k=
github.com/googleapis/enterprise-certificate-proxy v0.3.18/go.mod h1:rSEsBUemEBZEexP2y6jPp16LUmUbjmSbcPMQizR0o4k=
github.com/googleapis/enterprise-certificate-proxy v0.3.20 h1:t/xL64VUoN69MuMRQuJETqYGOw4Z9mSRJK9epIEtwFk=
github.com/googleapis/enterprise-certificate-proxy v0.3.20/go.mod h1:L3D/IQExI6LqEjBdXcZQ1WluSgigQmSwBboFstVPM4w=
github.com/googleapis/gax-go/v2 v2.0.4/go.mod h1:0Wqv26UfaUD9n4G6kQubkQ+KchISgw+vpHVxEJEs9eg=
github.com/googleapis/gax-go/v2 v2.0.5/go.mod h1:DWXyrwAJ9X0FpwwEdw+IPEYBICEFu5mhpdKc/us6bOk=
github.com/googleapis/gax-go/v2 v2.1.0/go.mod h1:Q3nei7sK6ybPYH7twZdmQpAd1MKb7pfu6SK+H1/DsU0=
@@ -1295,8 +1295,8 @@ github.com/grpc-ecosystem/grpc-gateway v1.16.0 h1:gmcG1KaJ57LophUzW0Hy8NmPhnMZb4
github.com/grpc-ecosystem/grpc-gateway v1.16.0/go.mod h1:BDjrQk3hbvj6Nolgz8mAMFbcEtjT1g+wF4CSlocrBnw=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.7.0/go.mod h1:hgWBS7lorOAVIJEQMi4ZsPv9hVvWI6+ch50m39Pf2Ks=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.11.3/go.mod h1:o//XUCC/F+yRGJoPO/VU0GSB0f8Nhgmxx0VIRUvaC0w=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.28.0 h1:HWRh5R2+9EifMyIHV7ZV+MIZqgz+PMpZ14Jynv3O2Zs=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.28.0/go.mod h1:JfhWUomR1baixubs02l85lZYYOm7LV6om4ceouMv45c=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0 h1:5VipnvEpbqr2gA2VbM+nYVbkIF28c5ZQfqCBQ5g2xfk=
github.com/grpc-ecosystem/grpc-gateway/v2 v2.29.0/go.mod h1:Hyl3n6Twe1hvtd9XUXDec4pTvgMSEixRuQKPTMH2bNs=
github.com/hashicorp/errwrap v1.0.0/go.mod h1:YH+1FKiLXxHSkmPseP+kNlulaMuP3n2brvKWEqk/Jc4=
github.com/hashicorp/errwrap v1.1.0 h1:OxrOeh75EUXMY8TBjag2fzXGZ40LB6IKw45YeGUDY2I=
github.com/hashicorp/errwrap v1.1.0/go.mod h1:YH+1FKiLXxHSkmPseP+kNlulaMuP3n2brvKWEqk/Jc4=
@@ -1511,8 +1511,8 @@ github.com/mattn/go-colorable v0.1.15/go.mod h1:6LmQG8QLFO4G5z1gPvYEzlUgJ2wF+stg
github.com/mattn/go-isatty v0.0.12/go.mod h1:cbi8OIDigv2wuxKPP5vlRcQ1OAZbq2CE4Kysco4FUpU=
github.com/mattn/go-isatty v0.0.14/go.mod h1:7GGIvUiUoEMVVmxf/4nioHXj79iQHKdU27kJ6hsGG94=
github.com/mattn/go-isatty v0.0.16/go.mod h1:kYGgaQfpe5nmfYZH+SKPsOc2e4SrIfOl2e/yFXSvRLM=
github.com/mattn/go-isatty v0.0.23 h1:cYwCQTQf3HB6xUC+BtyCLZNr7IzbOmoZbmssVNzSyiQ=
github.com/mattn/go-isatty v0.0.23/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/mattn/go-isatty v0.0.24 h1:tGZZoVgT/KiqK1c8ocVLeDS8BSWMRd47J3Lbz7vsReI=
github.com/mattn/go-isatty v0.0.24/go.mod h1:nMCL3Zebbrt45jsMDgnfIwz6ydEQApk5oEI3HqDio6A=
github.com/mattn/go-runewidth v0.0.3/go.mod h1:LwmH8dsx7+W8Uxz3IHJYH5QSwggIsqBzpuz5H//U1FU=
github.com/mattn/go-runewidth v0.0.13/go.mod h1:Jdepj2loyihRzMpdS35Xk/zdY8IAYHsh153qUoGf23w=
github.com/mattn/go-runewidth v0.0.24 h1:cpokDiIn0MGnhdHwuWnJBITySJ20QyNGnY2kR/ay2DU=
@@ -1636,8 +1636,8 @@ github.com/parquet-go/bitpack v1.0.0 h1:AUqzlKzPPXf2bCdjfj4sTeacrUwsT7NlcYDMUQxP
github.com/parquet-go/bitpack v1.0.0/go.mod h1:XnVk9TH+O40eOOmvpAVZ7K2ocQFrQwysLMnc6M/8lgs=
github.com/parquet-go/jsonlite v1.0.0 h1:87QNdi56wOfsE5bdgas0vRzHPxfJgzrXGml1zZdd7VU=
github.com/parquet-go/jsonlite v1.0.0/go.mod h1:nDjpkpL4EOtqs6NQugUsi0Rleq9sW/OtC1NnZEnxzF0=
github.com/parquet-go/parquet-go v0.30.1 h1:Oy6ganNrAdFiVwy7wNmWagfPTWA2X9Z3tVHBc7JtuX8=
github.com/parquet-go/parquet-go v0.30.1/go.mod h1:navtkAYr2LGoJVp141oXPlO/sxLvaOe3la2JEoD8+rg=
github.com/parquet-go/parquet-go v0.32.0 h1:NWDqTUHfrCS4cJP/Fj2HlxvqsrVedWG3sayMkf+znzM=
github.com/parquet-go/parquet-go v0.32.0/go.mod h1:navtkAYr2LGoJVp141oXPlO/sxLvaOe3la2JEoD8+rg=
github.com/pascaldekloe/goe v0.1.0 h1:cBOtyMzM9HTpWjXfbbunk26uA6nG3a8n06Wieeh0MwY=
github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144TG7ZOy1lc=
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
@@ -1659,8 +1659,8 @@ github.com/phpdave11/gofpdf v1.4.2/go.mod h1:zpO6xFn9yxo3YLyMvW8HcKWVdbNqgIfOOp2
github.com/phpdave11/gofpdi v1.0.12/go.mod h1:vBmVV0Do6hSBHC8uKUQ71JGW+ZGQq74llk/7bXwjDoI=
github.com/phpdave11/gofpdi v1.0.13/go.mod h1:vBmVV0Do6hSBHC8uKUQ71JGW+ZGQq74llk/7bXwjDoI=
github.com/pierrec/lz4/v4 v4.1.15/go.mod h1:gZWDp/Ze/IJXGXf23ltt2EXimqmTUXEy0GFuRQyBid4=
github.com/pierrec/lz4/v4 v4.1.27 h1:+PhzhWDrjRj89TH2sw43nE3+4+W8lSxIuQadEHZyjUk=
github.com/pierrec/lz4/v4 v4.1.27/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pierrec/lz4/v4 v4.1.28 h1:pPEPwRJ4kybBTfGt28q7lQsRJQHhC08axprdLD5Ppio=
github.com/pierrec/lz4/v4 v4.1.28/go.mod h1:EoQMVJgeeEOMsCqCzqFm2O0cJvljX2nGZjcRIPL34O4=
github.com/pierrre/compare v1.0.2 h1:k4IUsHgh+dbcAOIWCfxVa/7G6STjADH2qmhomv+1quc=
github.com/pierrre/compare v1.0.2/go.mod h1:8UvyRHH+9HS8Pczdd2z5x/wvv67krDwVxoOndaIIDVU=
github.com/pierrre/geohash v1.0.0 h1:f/zfjdV4rVofTCz1FhP07T+EMQAvcMM2ioGZVt+zqjI=
@@ -1749,8 +1749,8 @@ github.com/quic-go/qpack v0.6.0 h1:g7W+BMYynC1LbYLSqRt8PBg5Tgwxn214ZZR34VIOjz8=
github.com/quic-go/qpack v0.6.0/go.mod h1:lUpLKChi8njB4ty2bFLX2x4gzDqXwUpaO1DP9qMDZII=
github.com/quic-go/quic-go v0.59.0 h1:OLJkp1Mlm/aS7dpKgTc6cnpynnD2Xg7C1pwL6vy/SAw=
github.com/quic-go/quic-go v0.59.0/go.mod h1:upnsH4Ju1YkqpLXC305eW3yDZ4NfnNbmQRCMWS58IKU=
github.com/rabbitmq/amqp091-go v1.11.0 h1:HxIctVm9Gid/Vtn706necmZ7Wj6pgGI2eqplRbEY8O8=
github.com/rabbitmq/amqp091-go v1.11.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rabbitmq/amqp091-go v1.13.0 h1:L8NA1WtF76C6KA3LAoufjfLgbist/If1UQYcsOjtxXA=
github.com/rabbitmq/amqp091-go v1.13.0/go.mod h1:Hy4jKW5kQART1u+JkDTF9YYOQUHXqMuhrgxOEeS7G4o=
github.com/rclone/Proton-API-Bridge v1.0.4 h1:uGQJRjQC1hVLd5kqLsXc6CWO6oqrVeLoKQYoHapEZDg=
github.com/rclone/Proton-API-Bridge v1.0.4/go.mod h1:VTPBYZotKAeDLlAzxU2O/s14NXk9FxUt9hn1jhH2iY8=
github.com/rclone/go-proton-api v1.0.3 h1:3gBTzR+j0dYiTwtj9yKIdN/aV3W2a8KIPKp0GArojyQ=
@@ -2036,8 +2036,8 @@ github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2 h1:e2nGQPGC5OEPBWlMnLPFpd
github.com/ydb-platform/ydb-go-sdk-auth-environ v0.5.2/go.mod h1:9YzkhlIymWaJGX6KMU3vh5sOf3UKbCXkG/ZdjaI3zNM=
github.com/ydb-platform/ydb-go-sdk/v3 v3.44.0/go.mod h1:oSLwnuilwIpaF5bJJMAofnGgzPJusoI3zWMNb8I+GnM=
github.com/ydb-platform/ydb-go-sdk/v3 v3.47.3/go.mod h1:bWnOIcUHd7+Sl7DN+yhyY1H/I61z53GczvwJgXMgvj0=
github.com/ydb-platform/ydb-go-sdk/v3 v3.146.3 h1:0ybsLjWJ25M8XXGXk7NfjhO+s8bqBcGGWt8iB2DQiHQ=
github.com/ydb-platform/ydb-go-sdk/v3 v3.146.3/go.mod h1:b9NEO6mgaiqsnOMkS003uS82XsKh6GL+ZTFfPqXWz+c=
github.com/ydb-platform/ydb-go-sdk/v3 v3.147.1 h1:WRxyl1UdFD6EmFlujwrHm4sX9ktYb/adZ/SW+WWNOEg=
github.com/ydb-platform/ydb-go-sdk/v3 v3.147.1/go.mod h1:b9NEO6mgaiqsnOMkS003uS82XsKh6GL+ZTFfPqXWz+c=
github.com/ydb-platform/ydb-go-yc v0.12.1 h1:qw3Fa+T81+Kpu5Io2vYHJOwcrYrVjgJlT6t/0dOXJrA=
github.com/ydb-platform/ydb-go-yc v0.12.1/go.mod h1:t/ZA4ECdgPWjAb4jyDe8AzQZB5dhpGbi3iCahFaNwBY=
github.com/ydb-platform/ydb-go-yc-metadata v0.6.1 h1:9E5q8Nsy2RiJMZDNVy0A3KUrIMBPakJ2VgloeWbcI84=
@@ -2073,12 +2073,12 @@ go.einride.tech/aip v0.83.0 h1:TI21IdeOnLTwZEJ3BxtImIZk6bsN2Q+sd0x99SLiQ+M=
go.einride.tech/aip v0.83.0/go.mod h1:E8+wdTApA70odnpFzJgsGogHozC2JCIhFJBKPr8bVig=
go.etcd.io/bbolt v1.5.0 h1:S7GAl7Fxv12yohbwFfIbQCGDWbQbtDGPET4P/bD4lxU=
go.etcd.io/bbolt v1.5.0/go.mod h1:mkltfYE5aUHQxUct9N9V+Kp7aSjFqjgrhcXIS70Lrdk=
go.etcd.io/etcd/api/v3 v3.6.12 h1:OLOZUKEuAA36TR48F0cIaa8FdzrWygjyfrJxXg4iDgs=
go.etcd.io/etcd/api/v3 v3.6.12/go.mod h1:p14EIQXHbuOQbVvL/WEes5uqKnxP9AgKJgpjbMVvzvE=
go.etcd.io/etcd/client/pkg/v3 v3.6.12 h1:36zzB+pQOdHbhN+kH2iJz/K8bJn0ZLtLfPPO7jozTDo=
go.etcd.io/etcd/client/pkg/v3 v3.6.12/go.mod h1:hh2+ZXtfLzs3o6mn92ntgNPBrTJJOvXqICM5g3L3DMY=
go.etcd.io/etcd/client/v3 v3.6.12 h1:kMSP6JcPZMqSJiX+TXdUIBU/4eXEZWBAaui4VihMbIc=
go.etcd.io/etcd/client/v3 v3.6.12/go.mod h1:CMs6fJWYiZQk4ytFjd4lE1diOvvRMmtbbn/alZXd3dQ=
go.etcd.io/etcd/api/v3 v3.7.1 h1:KJG0/DcWGfe3Y1otDf/fsBf0TSSgpxZ5RO/L8SFt73E=
go.etcd.io/etcd/api/v3 v3.7.1/go.mod h1:8bXIpCMeV7E3/XL0Ix123ATn3dB+0V7d9zklHbB0m78=
go.etcd.io/etcd/client/pkg/v3 v3.7.1 h1:rKYsj3pRkR0eK3yjT3XOgrhqfmIfj9pzNgxjh7mfFv4=
go.etcd.io/etcd/client/pkg/v3 v3.7.1/go.mod h1:cnzZGIUzSfjEwLC6UBVsSXlEK1eepS/JUD7wE6PLRT0=
go.etcd.io/etcd/client/v3 v3.7.1 h1:0PEMMC0KuZmVIN+RAbdqfkZ45pYTgKVtmBEbRCvZFUg=
go.etcd.io/etcd/client/v3 v3.7.1/go.mod h1:ffNqALa8tRCYhYo1F9oR489y23K39Gz+BSR3ApAGYq0=
go.mongodb.org/mongo-driver v1.17.9 h1:IexDdCuuNJ3BHrELgBlyaH9p60JXAvdzWR128q+U5tU=
go.mongodb.org/mongo-driver v1.17.9/go.mod h1:LlOhpH5NUEfhxcAwG0UEkMqwYcc4JU18gtCdGudk/tQ=
go.opencensus.io v0.21.0/go.mod h1:mSImk1erAIZhrmZN+AvHh14ztQfjbGwt4TtuofqLduU=
@@ -2182,8 +2182,8 @@ golang.org/x/crypto v0.6.0/go.mod h1:OFC/31mSvZgRz0V1QTNCzfAI1aIRzbiufJtkMIlEp58
golang.org/x/crypto v0.7.0/go.mod h1:pYwdfH91IfpZVANVyUOhSIPZaFoJGxTFbZhFTx+dXZU=
golang.org/x/crypto v0.13.0/go.mod h1:y6Z2r+Rw4iayiXXAIxJIDAJ1zMW4yaTpebo8fPOliYc=
golang.org/x/crypto v0.14.0/go.mod h1:MVFd36DqK4CsrnJYDkBA3VC4m2GkXAM0PvzMCn4JQf4=
golang.org/x/crypto v0.54.0 h1:YLIA59K4fiNzHzjnZt2tUJQjQtUWfWbeHBqKtk3eScw=
golang.org/x/crypto v0.54.0/go.mod h1:KWL8ny2AZdGR2cWmzeHrp2azQPGogOv+HeQaVEXC2dk=
golang.org/x/crypto v0.55.0 h1:+KWHjbgOaAQ66dh/YlkZKHlz9ZUlq61AFirAR9ntP8M=
golang.org/x/crypto v0.55.0/go.mod h1:uq0V9dE/fzQuJtbnL+2EhWOE63vo164FY8xqEnV9xis=
golang.org/x/exp v0.0.0-20180321215751-8460e604b9de/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20180807140117-3d87b88a115f/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20190121172915-509febef88a4/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
@@ -2313,8 +2313,8 @@ golang.org/x/net v0.8.0/go.mod h1:QVkue5JL9kW//ek3r6jTKnTFis1tRmNAW2P1shuFdJc=
golang.org/x/net v0.10.0/go.mod h1:0qNGK6F8kojg2nk9dLZ2mShWaEBan6FAoqfSigmmuDg=
golang.org/x/net v0.15.0/go.mod h1:idbUs1IY1+zTqbi8yxTbhexhEEk5ur9LInksu6HrEpk=
golang.org/x/net v0.16.0/go.mod h1:NxSsAGuq816PNPmqtQdLE42eU2Fs7NoRIZrHJAlaCOE=
golang.org/x/net v0.57.0 h1:K5+3DljvIuDG9/Jv9rvyMywYNFCQ9RSUY6OOTTkT+tE=
golang.org/x/net v0.57.0/go.mod h1:KpXc8iv+r3XplLAG/f7Jsf9RPszJzdR0f58q9vGOuEU=
golang.org/x/net v0.58.0 h1:ynWG7rqYi4ccpTEuPZ2QGWHktVEM9DMCj9yzDE0Q7To=
golang.org/x/net v0.58.0/go.mod h1:YwCddHnFlT7eLQqVprV19OnhLGtc5xOKgE0RyqgfWAU=
golang.org/x/oauth2 v0.0.0-20180821212333-d2e6202438be/go.mod h1:N/0e6XlmueqKjAGxoOufVs8QHGRruUQn6yWY3a++T0U=
golang.org/x/oauth2 v0.0.0-20190226205417-e64efc72b421/go.mod h1:gOpvHmFTYa4IltrdGE7lF6nIHvwfUNPOp7c8zoXwtLw=
golang.org/x/oauth2 v0.0.0-20190604053449-0f29369cfe45/go.mod h1:gOpvHmFTYa4IltrdGE7lF6nIHvwfUNPOp7c8zoXwtLw=
@@ -2500,8 +2500,8 @@ golang.org/x/text v0.8.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.9.0/go.mod h1:e1OnstbJyHTd6l/uOt8jFFHp6TRDWZR/bV3emEE/zU8=
golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.40.0 h1:Ub2Z6/xjgF1WrYQz2nuITOEegKFtiIy+rieRJ5lHZKs=
golang.org/x/text v0.40.0/go.mod h1:hpnzDAfGV753zIKo+wk3u1bVKCGPbrnF7+7LBF/UHVY=
golang.org/x/text v0.41.0 h1:vz/seA0lnX87Othu2f/0L24RcgrXD9/YFTSuGjj3rH8=
golang.org/x/text v0.41.0/go.mod h1:jvf1O8ajNzZqhSrQBPbutR/EB83Cc0CFrezNQIwbb5M=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
@@ -2657,8 +2657,8 @@ google.golang.org/api v0.106.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/
google.golang.org/api v0.107.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.108.0/go.mod h1:2Ts0XTHNVWxypznxWOYUeI4g3WdP9Pk2Qk58+a/O9MY=
google.golang.org/api v0.110.0/go.mod h1:7FC4Vvx1Mooxh8C5HWjzZHcavuS2f6pmJpZx60ca7iI=
google.golang.org/api v0.289.0 h1:DmH0c6NigNFmsvsohM9bxv+MzVhag3aGHnojA5fFQjc=
google.golang.org/api v0.289.0/go.mod h1:weJZ3lldHFYI0DBFNKpJelUDNnusTt5YaOEgxvt8ci8=
google.golang.org/api v0.293.0 h1:p9XIWOf63U4OgYx120ZwVU8+vl4XTPmWfgVPnmOAS9w=
google.golang.org/api v0.293.0/go.mod h1:6n5tjEB1gzwniZTepZ0g5u+wM7Bof5GeULCx/zh8ZE0=
google.golang.org/appengine v1.1.0/go.mod h1:EbEs0AVv82hx2wNQdGPgUI5lhzA/G0D9YwlJXL52JkM=
google.golang.org/appengine v1.4.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
google.golang.org/appengine v1.5.0/go.mod h1:xpcJRLb0r/rnEns0DIKYYv+WjYCduHsrkT7/EB5XEv4=
@@ -2796,8 +2796,8 @@ google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94 h1:YJjbgu+dkp5kUJL
google.golang.org/genproto v0.0.0-20260519071638-aa98bba5eb94/go.mod h1:RRHjglSYABVCWpQ7USCpdfhcd9t4PkajvVwyynZizTc=
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800 h1:admdQBe8jR3VWhBsUrAOaF2Qw6K/+p5pSm1GN8+6Fw4=
google.golang.org/genproto/googleapis/api v0.0.0-20260706201446-f0a921348800/go.mod h1:FPk7EXUKMtImne7AmknoYjT4QXqKIzzRbeQIXzLk6fQ=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260715232425-e75dac1f907d h1:Jkpk39hlTZOIp3RbfvNX9R8Hv+Sw0X89nlU/xFOErsc=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260715232425-e75dac1f907d/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260807164820-c8921c73eeea h1:kVhQEPTpKQahD5+JSBTfBB19wcgQTTjAIn45MBqnyHk=
google.golang.org/genproto/googleapis/rpc v0.0.0-20260807164820-c8921c73eeea/go.mod h1:4Hqkh8ycfw05ld/3BWL7rJOSfebL2Q+DVDeRgYgxUU8=
google.golang.org/grpc v1.19.0/go.mod h1:mqu4LbDTu4XGKhr4mRzUsmM4RtVoemTSY81AxZiDr8c=
google.golang.org/grpc v1.20.1/go.mod h1:10oTOabMzJvdu6/UiuZezV6QK5dSlG84ov/aaiqXj38=
google.golang.org/grpc v1.21.1/go.mod h1:oYelfM1adQP15Ek0mdvEgi9Df8B9CZIaU1084ijfRaM=
@@ -2915,23 +2915,23 @@ modernc.org/b v1.0.0/go.mod h1:uZWcZfRj1BpYzfN9JTerzlNUnnPsV9O2ZA8JsRcubNg=
modernc.org/cc/v3 v3.36.0/go.mod h1:NFUHyPn4ekoC/JHeZFfZurN6ixxawE1BnVonP/oahEI=
modernc.org/cc/v3 v3.36.2/go.mod h1:NFUHyPn4ekoC/JHeZFfZurN6ixxawE1BnVonP/oahEI=
modernc.org/cc/v3 v3.36.3/go.mod h1:NFUHyPn4ekoC/JHeZFfZurN6ixxawE1BnVonP/oahEI=
modernc.org/cc/v4 v4.28.4 h1:Hd/4Es+MBj+/7hSdZaisNyu6bv3V0Dp2MdllyfqaH+c=
modernc.org/cc/v4 v4.28.4/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/cc/v4 v4.29.1 h1:MKgdCV3WykTSPqpVrnxdEDS0HEd2FHpKZDzxzU5LyeI=
modernc.org/cc/v4 v4.29.1/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/ccgo/v3 v3.0.0-20220428102840-41399a37e894/go.mod h1:eI31LL8EwEBKPpNpA4bU1/i+sKOwOrQy8D87zWUcRZc=
modernc.org/ccgo/v3 v3.0.0-20220430103911-bc99d88307be/go.mod h1:bwdAnOoaIt8Ax9YdWGjxWsdkPcZyRPHqrOvJxaKAKGw=
modernc.org/ccgo/v3 v3.16.4/go.mod h1:tGtX0gE9Jn7hdZFeU88slbTh1UtCYKusWOoCJuvkWsQ=
modernc.org/ccgo/v3 v3.16.6/go.mod h1:tGtX0gE9Jn7hdZFeU88slbTh1UtCYKusWOoCJuvkWsQ=
modernc.org/ccgo/v3 v3.16.8/go.mod h1:zNjwkizS+fIFDrDjIAgBSCLkWbJuHF+ar3QRn+Z9aws=
modernc.org/ccgo/v3 v3.16.9/go.mod h1:zNMzC9A9xeNUepy6KuZBbugn3c0Mc9TeiJO4lgvkJDo=
modernc.org/ccgo/v4 v4.34.4 h1:OVnSOWQjVKOYkFxoHYB+qQmSHK5gqMqARM+K9DpR/Ws=
modernc.org/ccgo/v4 v4.34.4/go.mod h1:qdKqE8FNIYyysougB1RX9MxCzp5oJOcQXSobANJ4TuE=
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
modernc.org/ccorpus v1.11.6/go.mod h1:2gEUTrWqdpH2pXsmTM1ZkjeSrUWDpjMu2T6m29L/ErQ=
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
modernc.org/gc/v3 v3.1.3 h1:6QAplYyVO+KdPW3pGnqmJDUxtkec8ooEWvks/hhU3lc=
modernc.org/gc/v3 v3.1.3/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
modernc.org/httpfs v1.0.6/go.mod h1:7dosgurJGp0sPaRanU53W4xZYKh14wfzX420oZADeHM=
@@ -2942,8 +2942,8 @@ modernc.org/libc v1.16.17/go.mod h1:hYIV5VZczAmGZAnG15Vdngn5HSF5cSkbvfz2B7GRuVU=
modernc.org/libc v1.16.19/go.mod h1:p7Mg4+koNjc8jkqwcoFBJx7tXkpj00G77X7A72jXPXA=
modernc.org/libc v1.17.0/go.mod h1:XsgLldpP4aWlPlsjqKRdHPqCxCjISdHfM/yeWC5GyW0=
modernc.org/libc v1.17.1/go.mod h1:FZ23b+8LjxZs7XtFMbSzL/EhPxNbfZbErxEHc7cbD9s=
modernc.org/libc v1.73.4 h1:+ra4Ui8ngyt8HDcO1FTDPWlkAh6yOdaO2yAoh8MddQA=
modernc.org/libc v1.73.4/go.mod h1:DXZ3eO8qMCNn2SnmTNCiC71nJ9Rcq3PsnpU6Vc4rWK8=
modernc.org/libc v1.74.4 h1:fX1Omw4o2/1C2iRkkIsrQTasJQldLhRmuPreXLoWs9k=
modernc.org/libc v1.74.4/go.mod h1:eeQAS9W3sZeKYMFubydxJpII9ybHWshk+7or7bLG9co=
modernc.org/mathutil v1.1.1/go.mod h1:mZW8CKdRPY1v87qxC/wUdX5O1qDzXMP5TH3wjfpga6E=
modernc.org/mathutil v1.2.2/go.mod h1:mZW8CKdRPY1v87qxC/wUdX5O1qDzXMP5TH3wjfpga6E=
modernc.org/mathutil v1.4.1/go.mod h1:mZW8CKdRPY1v87qxC/wUdX5O1qDzXMP5TH3wjfpga6E=
@@ -2962,8 +2962,8 @@ modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
modernc.org/sqlite v1.18.1/go.mod h1:6ho+Gow7oX5V+OiOQ6Tr4xeqbx13UZ6t+Fw9IRUG4d4=
modernc.org/sqlite v1.53.0 h1:20WG8N9q4ji/dEqGk4uiI0c6OPjSeLTNYGFCc3+7c1M=
modernc.org/sqlite v1.53.0/go.mod h1:xoEpOIpGrgT48H5iiyt/YXPCZPEzlfmfFwtk8Lklw8s=
modernc.org/sqlite v1.56.0 h1:/D8e2RfFqoy/Zc6PuC76U28zFwmI/sYx1Kjm4yEn9e0=
modernc.org/sqlite v1.56.0/go.mod h1:yCJ2cmAaIkHQ25oXWrF8H4O1lIfPYPR26yCEDj2P3pQ=
modernc.org/strutil v1.1.0/go.mod h1:lstksw84oURvj9y3tn8lGvRxyRC1S2+g5uuIzNfIOBs=
modernc.org/strutil v1.1.1/go.mod h1:DE+MQQ/hjKBZS2zNInV5hhcipt5rLPWkmpbGeW5mmdw=
modernc.org/strutil v1.1.3/go.mod h1:MEHNA7PdEnEwLvspRMtWTNnp2nnyvMfkimT1NKNAGbw=
+2 -2
View File
@@ -1,6 +1,6 @@
apiVersion: v1
description: SeaweedFS
name: seaweedfs
appVersion: "4.41"
appVersion: "4.44"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.41.0
version: 4.44.0
+28 -2
View File
@@ -22,8 +22,8 @@ helm install --values=values.yaml seaweedfs seaweedfs/seaweedfs
## Info:
* master/filer/volume are stateful sets with anti-affinity on the hostname,
so your deployment will be spread/HA.
* chart is using memsql(mysql) as the filer backend to enable HA (multiple filer instances) and backup/HA memsql can provide.
* mysql user/password are created in a k8s secret (default: `<release>-seaweedfs-db-secret`) and injected to the filer with ENV.
* leveldb2 is the default filer backend; a mysql-compatible database (memsql, ...) enables HA (multiple filer instances) and the backup/HA it can provide.
* with `filer.extraEnvironmentVars.WEED_MYSQL_ENABLED` set to `"true"`, mysql user/password are created in a k8s secret (default: `<release>-seaweedfs-db-secret`) and injected to the filer with ENV. On any other store neither the secret nor the `WEED_MYSQL_*` env, plain or secret-backed, is rendered.
* cert config exists and can be enabled, but not been tested, requires cert-manager to be installed.
## Prerequisites
@@ -289,6 +289,32 @@ stringData:
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"snu8yoP6QAlY0ne4","secretKey":"PNzBcmeLNEdR0oviwm04NQAicOrDH1Km"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"SCigFee6c5lbi04A","secretKey":"kgFhbT38R8WUYVtiFQ1OiSVOrYr3NKku"}],"actions":["Read"]}]}'
```
#### Source S3 credentials from an existing Secret
To keep the keys out of `values.yaml` while still letting the chart generate the
identities file, point an identity at an existing Secret:
```yaml
s3:
enabled: true
enableAuth: true
credentials:
admin:
existingSecret: minio-root
accessKeyKey: root-user
secretKeyKey: root-password
```
`accessKeyKey` and `secretKeyKey` default to the chart's own key names
(`admin_access_key_id`, `admin_secret_access_key`, and the `read_` pair). The
generated `seaweedfs_s3_config` references the keys as `${SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID}`
and the gateway resolves them from the environment, which the chart wires up
from the Secret. Nothing is read from the cluster at render time, so
`helm template`, `--dry-run` and an Argo CD diff all render what an install
applies. Rotating a key in the Secret takes effect on the next pod restart, as
with any other environment variable. The COSI driver parses the config itself
and does not resolve these references.
## Admin Component
The admin component provides a modern web-based administration interface for managing SeaweedFS clusters. It includes:
@@ -69,6 +69,9 @@ spec:
{{- if .Values.admin.priorityClassName }}
priorityClassName: {{ .Values.admin.priorityClassName | quote }}
{{- end }}
{{- if .Values.admin.schedulerName }}
schedulerName: {{ .Values.admin.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.admin.serviceAccountName }}
serviceAccountName: {{ .Values.admin.serviceAccountName | quote }}
@@ -64,6 +64,9 @@ spec:
{{- if .Values.allInOne.priorityClassName }}
priorityClassName: {{ .Values.allInOne.priorityClassName | quote }}
{{- end }}
{{- if .Values.allInOne.schedulerName }}
schedulerName: {{ .Values.allInOne.schedulerName | quote }}
{{- end }}
{{- if .Values.allInOne.serviceAccountName }}
serviceAccountName: {{ .Values.allInOne.serviceAccountName | quote }}
{{- end }}
@@ -81,6 +84,9 @@ spec:
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
env:
{{- include "seaweedfs.licenseEnv" . | nindent 12 }}
{{- if and .Values.allInOne.s3.enabled (or .Values.allInOne.s3.enableAuth .Values.s3.enableAuth .Values.filer.s3.enableAuth) }}
{{- include "seaweedfs.s3.credentialEnv" . | nindent 12 }}
{{- end }}
{{- /* Determine default cluster alias and the corresponding env var keys to avoid conflicts */}}
{{- $mergedExtraEnvironmentVars := dict }}
{{- include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.allInOne "target" $mergedExtraEnvironmentVars) }}
@@ -57,6 +57,9 @@ spec:
{{- if .Values.cosi.priorityClassName }}
priorityClassName: {{ .Values.cosi.priorityClassName | quote }}
{{- end }}
{{- if .Values.cosi.schedulerName }}
schedulerName: {{ .Values.cosi.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
serviceAccountName: {{ include "seaweedfs.componentName" (list . "objectstorage-provisioner") }}
{{- if .Values.cosi.initContainers }}
@@ -10,9 +10,6 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- if .Values.filer.metricsPort }}
monitoring: "true"
{{- end }}
{{- if .Values.filer.annotations }}
annotations:
{{- toYaml .Values.filer.annotations | nindent 4 }}
@@ -10,6 +10,9 @@ metadata:
app.kubernetes.io/managed-by: {{ .Release.Service }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
{{- if .Values.filer.metricsPort }}
monitoring: "true"
{{- end }}
annotations:
service.alpha.kubernetes.io/tolerate-unready-endpoints: "true"
{{- if .Values.filer.annotations }}
@@ -30,6 +30,7 @@ spec:
app.kubernetes.io/name: {{ template "seaweedfs.name" . }}
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: filer
monitoring: "true"
{{- end }}
{{- end }}
{{- end }}
@@ -76,6 +76,9 @@ spec:
{{- if .Values.filer.priorityClassName }}
priorityClassName: {{ .Values.filer.priorityClassName | quote }}
{{- end }}
{{- if .Values.filer.schedulerName }}
schedulerName: {{ .Values.filer.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.filer.initContainers }}
initContainers:
@@ -89,6 +92,7 @@ spec:
image: {{ template "seaweedfs.filer.image" . }}
imagePullPolicy: {{ default "IfNotPresent" .Values.global.seaweedfs.imagePullPolicy }}
env:
{{- $mysqlEnabled := include "seaweedfs.filer.mysqlEnabled" . }}
- name: POD_IP
valueFrom:
fieldRef:
@@ -101,6 +105,7 @@ spec:
valueFrom:
fieldRef:
fieldPath: metadata.namespace
{{- if $mysqlEnabled }}
- name: WEED_MYSQL_USERNAME
valueFrom:
secretKeyRef:
@@ -113,10 +118,21 @@ spec:
name: {{ include "seaweedfs.fullname" . }}-db-secret
key: password
optional: true
{{- end }}
- name: SEAWEEDFS_FULLNAME
value: "{{ include "seaweedfs.fullname" . }}"
{{- if and .Values.filer.s3.enabled .Values.filer.s3.enableAuth }}
{{- include "seaweedfs.s3.credentialEnv" . | nindent 12 }}
{{- end }}
{{- $mergedExtraEnvironmentVars := dict }}
{{- include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.filer "target" $mergedExtraEnvironmentVars) }}
{{- if not $mysqlEnabled }}
{{- range $key := keys $mergedExtraEnvironmentVars }}
{{- if hasPrefix "WEED_MYSQL_" $key }}
{{- $_ := unset $mergedExtraEnvironmentVars $key }}
{{- end }}
{{- end }}
{{- end }}
{{- range $key := keys $mergedExtraEnvironmentVars | sortAlpha }}
{{- $value := index $mergedExtraEnvironmentVars $key }}
- name: {{ $key }}
@@ -129,10 +145,12 @@ spec:
{{- end }}
{{- if .Values.filer.secretExtraEnvironmentVars }}
{{- range $key, $value := .Values.filer.secretExtraEnvironmentVars }}
{{- if or $mysqlEnabled (not (hasPrefix "WEED_MYSQL_" $key)) }}
- name: {{ $key }}
valueFrom: {{ toYaml $value | nindent 16 }}
{{- end }}
{{- end }}
{{- end }}
command:
- "/bin/sh"
- "-ec"
@@ -69,6 +69,9 @@ spec:
{{- if .Values.master.priorityClassName }}
priorityClassName: {{ .Values.master.priorityClassName | quote }}
{{- end }}
{{- if .Values.master.schedulerName }}
schedulerName: {{ .Values.master.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
serviceAccountName: {{ .Values.master.serviceAccountName | default (include "seaweedfs.serviceAccountName" .) | quote }} # for deleting statefulset pods after migration
{{- if .Values.master.initContainers }}
@@ -61,6 +61,9 @@ spec:
{{- if .Values.s3.priorityClassName }}
priorityClassName: {{ .Values.s3.priorityClassName | quote }}
{{- end }}
{{- if .Values.s3.schedulerName }}
schedulerName: {{ .Values.s3.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.s3.serviceAccountName }}
serviceAccountName: {{ .Values.s3.serviceAccountName | quote }}
@@ -91,6 +94,9 @@ spec:
fieldPath: metadata.namespace
- name: SEAWEEDFS_FULLNAME
value: "{{ include "seaweedfs.fullname" . }}"
{{- if .Values.s3.enableAuth }}
{{- include "seaweedfs.s3.credentialEnv" . | nindent 12 }}
{{- end }}
{{- $mergedExtraEnvironmentVars := dict }}
{{- include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.s3 "target" $mergedExtraEnvironmentVars) }}
{{- range $key := keys $mergedExtraEnvironmentVars | sortAlpha }}
@@ -1,4 +1,7 @@
{{- if or (and (or .Values.s3.enabled .Values.allInOne.enabled) .Values.s3.enableAuth (not .Values.s3.existingConfigSecret)) (and .Values.filer.s3.enabled .Values.filer.s3.enableAuth (not .Values.filer.s3.existingConfigSecret)) }}
{{- /* Mirrors the condition the all-in-one deployment mounts this secret under,
so the flags that make it mount are exactly the flags that create it. */}}
{{- $allInOneAuth := and .Values.allInOne.enabled .Values.allInOne.s3.enabled (or .Values.allInOne.s3.enableAuth .Values.s3.enableAuth .Values.filer.s3.enableAuth) (not (or .Values.allInOne.s3.existingConfigSecret .Values.s3.existingConfigSecret .Values.filer.s3.existingConfigSecret)) }}
{{- if or (and (or .Values.s3.enabled .Values.allInOne.enabled) .Values.s3.enableAuth (not .Values.s3.existingConfigSecret)) (and .Values.filer.s3.enabled .Values.filer.s3.enableAuth (not .Values.filer.s3.existingConfigSecret)) $allInOneAuth }}
{{- $secretName := printf "%s-s3-secret" (include "seaweedfs.fullname" .) }}
{{- $legacySecretName := "seaweedfs-s3-secret" }}
{{- $lookupName := $secretName }}
@@ -14,14 +17,20 @@
{{- $adminCreds := $creds.admin | default dict -}}
{{- $access_key_admin := $adminCreds.accessKey -}}
{{- $secret_key_admin := $adminCreds.secretKey -}}
{{- if not (and $access_key_admin $secret_key_admin) -}}
{{- if $adminCreds.existingSecret -}}
{{- $access_key_admin = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "admin" "accessKey")) -}}
{{- $secret_key_admin = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "admin" "secretKey")) -}}
{{- else if not (and $access_key_admin $secret_key_admin) -}}
{{- $access_key_admin = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "admin_access_key_id" "length" 20 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- $secret_key_admin = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "admin_secret_access_key" "length" 40 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- end -}}
{{- $readCreds := $creds.read | default dict -}}
{{- $access_key_read := $readCreds.accessKey -}}
{{- $secret_key_read := $readCreds.secretKey -}}
{{- if not (and $access_key_read $secret_key_read) -}}
{{- if $readCreds.existingSecret -}}
{{- $access_key_read = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "read" "accessKey")) -}}
{{- $secret_key_read = printf "${%s}" (include "seaweedfs.s3.credentialEnvName" (list "read" "secretKey")) -}}
{{- else if not (and $access_key_read $secret_key_read) -}}
{{- $access_key_read = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "read_access_key_id" "length" 20 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- $secret_key_read = include "seaweedfs.getOrGeneratePassword" (dict "namespace" .Release.Namespace "secretName" $secretName "key" "read_secret_access_key" "length" 40 "existingSecret" (ternary $existingSecret nil $reuse)) -}}
{{- end -}}
@@ -41,10 +50,16 @@ metadata:
app.kubernetes.io/instance: {{ .Release.Name }}
app.kubernetes.io/component: s3
stringData:
{{- /* An identity read from an existing Secret keeps its keys there; the
config below names the environment variables carrying them. */}}
{{- if not $adminCreds.existingSecret }}
admin_access_key_id: {{ $access_key_admin }}
admin_secret_access_key: {{ $secret_key_admin }}
{{- end }}
{{- if not $readCreds.existingSecret }}
read_access_key_id: {{ $access_key_read }}
read_secret_access_key: {{ $secret_key_read }}
{{- end }}
seaweedfs_s3_config: '{"identities":[{"name":"anvAdmin","credentials":[{"accessKey":"{{ $access_key_admin }}","secretKey":"{{ $secret_key_admin }}"}],"actions":["Admin","Read","Write"]},{"name":"anvReadOnly","credentials":[{"accessKey":"{{ $access_key_read }}","secretKey":"{{ $secret_key_read }}"}],"actions":["Read"]}]}'
{{- if .Values.filer.s3.auditLogConfig }}
filer_s3_auditLogConfig.json: |
@@ -61,6 +61,9 @@ spec:
{{- if .Values.sftp.priorityClassName }}
priorityClassName: {{ .Values.sftp.priorityClassName | quote }}
{{- end }}
{{- if .Values.sftp.schedulerName }}
schedulerName: {{ .Values.sftp.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.sftp.serviceAccountName }}
serviceAccountName: {{ .Values.sftp.serviceAccountName | quote }}
@@ -76,6 +76,18 @@ Inject extra environment vars in the format key:value, if populated
{{- end }}
{{- end -}}
{{/* Whether the mysql filer store is selected; a flag the chart cannot read counts as selected. */}}
{{- define "seaweedfs.filer.mysqlEnabled" -}}
{{- $merged := dict -}}
{{- $_ := include "seaweedfs.mergeExtraEnvironmentVars" (dict "global" .Values.global.seaweedfs "component" .Values.filer "target" $merged) -}}
{{- $enabled := index $merged "WEED_MYSQL_ENABLED" -}}
{{- if or (kindIs "map" $enabled) (hasKey (.Values.filer.secretExtraEnvironmentVars | default dict) "WEED_MYSQL_ENABLED") -}}
true
{{- else if and $enabled (eq (lower (toString $enabled)) "true") -}}
true
{{- end -}}
{{- end -}}
{{/* Return the proper filer image */}}
{{- define "seaweedfs.filer.image" -}}
{{- if .Values.filer.imageOverride -}}
@@ -496,6 +508,39 @@ true
{{- end }}
{{- end -}}
{{/* Name of the environment variable carrying one generated S3 credential
field, e.g. SEAWEEDFS_S3_ADMIN_ACCESS_KEY_ID. The generated identities file
names it in place of the key when the key lives in an existing Secret.
Usage: include "seaweedfs.s3.credentialEnvName" (list "admin" "accessKey") */}}
{{- define "seaweedfs.s3.credentialEnvName" -}}
{{- $identity := index . 0 -}}
{{- $field := index . 1 -}}
{{- printf "SEAWEEDFS_S3_%s_%s" (upper $identity) (ternary "ACCESS_KEY_ID" "SECRET_ACCESS_KEY" (eq $field "accessKey")) -}}
{{- end -}}
{{/* Environment for the S3 identities the chart generates from an existing
Secret. The gateway resolves the ${VAR} references the identities file
carries, so the keys never enter the rendered manifests and a dry run
renders the same as an install. */}}
{{- define "seaweedfs.s3.credentialEnv" -}}
{{- $creds := $.Values.s3.credentials | default dict -}}
{{- range $identity := list "admin" "read" -}}
{{- $identityCreds := index $creds $identity | default dict -}}
{{- if $identityCreds.existingSecret }}
- name: {{ include "seaweedfs.s3.credentialEnvName" (list $identity "accessKey") }}
valueFrom:
secretKeyRef:
name: {{ $identityCreds.existingSecret | quote }}
key: {{ default (printf "%s_access_key_id" $identity) $identityCreds.accessKeyKey | quote }}
- name: {{ include "seaweedfs.s3.credentialEnvName" (list $identity "secretKey") }}
valueFrom:
secretKeyRef:
name: {{ $identityCreds.existingSecret | quote }}
key: {{ default (printf "%s_secret_access_key" $identity) $identityCreds.secretKeyKey | quote }}
{{- end -}}
{{- end -}}
{{- end -}}
{{/* Generate a compatible trafficDistribution value due to "PreferClose" fast deprecation in k8s v1.35.
Accepts a dict with "value" (the trafficDistribution string) and "Capabilities". */}}
{{- define "seaweedfs.trafficDistribution" -}}
@@ -1,4 +1,4 @@
{{- if .Values.filer.enabled }}
{{- if and .Values.filer.enabled (include "seaweedfs.filer.mysqlEnabled" .) }}
apiVersion: v1
kind: Secret
type: Opaque
@@ -73,6 +73,9 @@ spec:
{{- if $volume.priorityClassName }}
priorityClassName: {{ $volume.priorityClassName | quote }}
{{- end }}
{{- if $volume.schedulerName }}
schedulerName: {{ $volume.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
serviceAccountName: {{ $volume.serviceAccountName | default (include "seaweedfs.serviceAccountName" $) | quote }} # for deleting statefulset pods after migration
{{- $initContainers_exists := include "seaweedfs.volume.initContainers_exists" $ -}}
@@ -64,6 +64,9 @@ spec:
{{- if .Values.worker.priorityClassName }}
priorityClassName: {{ .Values.worker.priorityClassName | quote }}
{{- end }}
{{- if .Values.worker.schedulerName }}
schedulerName: {{ .Values.worker.schedulerName | quote }}
{{- end }}
enableServiceLinks: false
{{- if .Values.worker.serviceAccountName }}
serviceAccountName: {{ .Values.worker.serviceAccountName | quote }}
+39
View File
@@ -242,6 +242,10 @@ master:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to specify a custom scheduler for master pods
# ref: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -535,6 +539,10 @@ volume:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to specify a custom scheduler for volume pods
# ref: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -815,6 +823,10 @@ filer:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to specify a custom scheduler for filer pods
# ref: https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -885,6 +897,7 @@ filer:
# extraEnvVars is a list of extra environment variables to set with the stateful set.
extraEnvironmentVars:
# the WEED_MYSQL_* keys and the db credential secret only render while this is "true"
WEED_MYSQL_ENABLED: "false"
WEED_MYSQL_HOSTNAME: "mysql-db-host"
WEED_MYSQL_PORT: "3306"
@@ -1002,13 +1015,24 @@ s3:
# Optionally provide explicit credentials for the S3 gateway.
# When set, these are used in the generated s3 secret instead of
# auto-generating random credentials.
# An identity may instead name an existing Secret to read its keys from. The
# generated config then references the keys through environment variables, so
# nothing is looked up at render time and a dry run renders what an install
# applies. Note the COSI driver parses the config itself and does not resolve
# those references.
# credentials:
# admin:
# accessKey: ""
# secretKey: ""
# existingSecret: ""
# accessKeyKey: admin_access_key_id
# secretKeyKey: admin_secret_access_key
# read:
# accessKey: ""
# secretKey: ""
# existingSecret: ""
# accessKeyKey: read_access_key_id
# secretKeyKey: read_secret_access_key
auditLogConfig: {}
# You may specify buckets to be created during the install or upgrade process.
# Buckets may be exposed publicly by setting `anonymousRead` to `true`
@@ -1074,6 +1098,10 @@ s3:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# used to assign a custom scheduler to server pods
# ref: https://kubernetes.io/docs/tasks/extend-kubernetes/configure-multiple-schedulers/
schedulerName: ""
# used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -1227,6 +1255,7 @@ sftp:
tolerations: ""
nodeSelector: ""
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
podSecurityContext: {}
containerSecurityContext: {}
@@ -1351,6 +1380,7 @@ admin:
tolerations: ""
nodeSelector: ""
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
podSecurityContext: {}
containerSecurityContext: {}
@@ -1493,6 +1523,7 @@ worker:
tolerations: ""
nodeSelector: ""
priorityClassName: ""
schedulerName: ""
serviceAccountName: ""
podSecurityContext: {}
containerSecurityContext: {}
@@ -1732,6 +1763,10 @@ allInOne:
# ref: https://kubernetes.io/docs/concepts/configuration/pod-priority-preemption/
priorityClassName: ""
# Used to assign a custom scheduler to pods
# ref: https://kubernetes.io/docs/tasks/extend-kubernetes/configure-multiple-schedulers/
schedulerName: ""
# Used to assign a service account.
# ref: https://kubernetes.io/docs/tasks/configure-pod-container/configure-service-account/
serviceAccountName: ""
@@ -1794,6 +1829,10 @@ cosi:
podSecurityContext: {}
containerSecurityContext: {}
# used to assign a custom scheduler to cosi pods
# ref: https://kubernetes.io/docs/tasks/extend-kubernetes/configure-multiple-schedulers/
schedulerName: ""
extraVolumes: ""
extraVolumeMounts: ""
+66
View File
@@ -0,0 +1,66 @@
<?xml version="1.0" encoding="UTF-8"?>
<svg id="Ebene_2" data-name="Ebene 2" xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" viewBox="0 0 512 512">
<defs>
<style>
.cls-1 {
fill: url(#Unbenannter_Verlauf_20);
}
.cls-2 {
fill: url(#Unbenannter_Verlauf_32);
}
.cls-3 {
fill: url(#Unbenannter_Verlauf_42);
}
.cls-4 {
fill: #fff;
}
.cls-5 {
fill: url(#Unbenannter_Verlauf_12);
}
.cls-6 {
fill: url(#Unbenannter_Verlauf_28);
}
</style>
<linearGradient id="Unbenannter_Verlauf_12" data-name="Unbenannter Verlauf 12" x1="19.92" y1="19.92" x2="492.08" y2="492.08" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#0282e9"/>
<stop offset=".48" stop-color="#0162bf"/>
<stop offset=".87" stop-color="#01469b"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_42" data-name="Unbenannter Verlauf 42" x1="180.67" y1="151.78" x2="234.43" y2="384.66" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#23b5e6"/>
<stop offset=".41" stop-color="#138bcf"/>
<stop offset=".86" stop-color="#0158b3"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_28" data-name="Unbenannter Verlauf 28" x1="275.58" y1="46.33" x2="275.58" y2="466.33" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#9fecf6"/>
<stop offset=".49" stop-color="#62dff2"/>
<stop offset="1" stop-color="#0a6cb7"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_20" data-name="Unbenannter Verlauf 20" x1="124.99" y1="226.28" x2="235.06" y2="462.33" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#c9f9fd"/>
<stop offset=".29" stop-color="#9fe7f6"/>
<stop offset=".93" stop-color="#36bbe5"/>
<stop offset="1" stop-color="#2ab6e3"/>
</linearGradient>
<linearGradient id="Unbenannter_Verlauf_32" data-name="Unbenannter Verlauf 32" x1="358.38" y1="251.34" x2="277.45" y2="473.69" gradientUnits="userSpaceOnUse">
<stop offset="0" stop-color="#43cbee"/>
<stop offset=".42" stop-color="#29a1d6"/>
<stop offset="1" stop-color="#0161b2"/>
</linearGradient>
</defs>
<g id="seaweedfs">
<rect id="background" class="cls-5" width="512" height="512" rx="68" ry="68"/>
<circle class="cls-4" cx="366.58" cy="152.44" r="31.33"/>
<circle class="cls-4" cx="386.81" cy="223.33" r="18.67"/>
<circle class="cls-4" cx="383.69" cy="288" r="9.22"/>
<path class="cls-3" d="M177.81,152.44s22.5,8.02,30.17,28.02c7.67,20,3.83,61.31,9.5,81.98s19.33,29,21,45.67-1.67,76-1.67,76c0,0-24-31.89-29.5-64s-21.53-41.06-24.6-65.72c-3.07-24.67.77-32.61,2.1-54.61s-1.84-35.72-7-47.33Z"/>
<path class="cls-6" d="M288.58,46.33c-14.56,7.67-42.22,35.22-42.22,79.11s23.11,54.22,23.11,94.22-38.89,48.11-38.89,95.56,21.56,66,21.56,95.56-11.11,55.56-11.11,55.56c0,0,26.97-17.72,33.33-60s-6.89-54.44,7.22-81.78,39-48.22,39-92.89-36.78-89.22-40.89-108.44-2.33-43.22,8.89-76.89Z"/>
<path class="cls-1" d="M120.58,228.33s31.72,14.27,41.55,36.5c9.83,22.22,3.45,40.17,17.45,61.17s51,41,56,73.33-7.67,66.33-7.67,66.33c0,0-5.67-59-30.33-73.33s-46.33-35.67-54.33-59-1.33-50.67-5.33-69c-4-18.33-7.33-23.67-17.33-36Z"/>
<path class="cls-2" d="M352.25,249.11c-9,5.67-18.87,17.85-21.71,31.2s2.94,19.57-7.06,41.57-22.19,25.28-30.97,49.72.21,35.36-11.33,59.56-23.92,35.17-23.92,35.17c0,0,34.77-8.17,45.69-30.94s5.08-35.22,20.53-52.61,24.67-30.89,27.25-56.72-7.33-26.22,1.53-76.94Z"/>
</g>
</svg>

After

Width:  |  Height:  |  Size: 3.5 KiB

+954 -1
View File
@@ -9263,6 +9263,959 @@
}
]
},
{
"type": "row",
"title": "Plugin Workers",
"collapsed": true,
"id": 229,
"gridPos": {
"h": 1,
"w": 24,
"x": 0,
"y": 27
},
"panels": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Plugin Workers Connected",
"description": "Workers with a live control stream. Unlike the Admin/Maintenance row, this counts plugin workers, which is what a Rust or Go plugin worker actually is.",
"type": "stat",
"id": 219,
"gridPos": {
"h": 4,
"w": 6,
"x": 0,
"y": 28
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "count(SeaweedFS_worker_connected{cluster=~\"$cluster\"} == 1) or vector(0)",
"instant": true,
"range": false,
"refId": "A"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Worker Job Failures (1h)",
"description": "Jobs that ended in failure across all workers in the last hour.",
"type": "stat",
"id": 220,
"gridPos": {
"h": 4,
"w": 6,
"x": 6,
"y": 28
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "thresholds"
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"colorMode": "value",
"graphMode": "area",
"justifyMode": "auto",
"orientation": "auto",
"reduceOptions": {
"calcs": [
"lastNotNull"
],
"fields": "",
"values": false
},
"textMode": "auto"
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(increase(SeaweedFS_worker_jobs_total{cluster=~\"$cluster\",result=\"failed\"}[1h])) or vector(0)",
"instant": true,
"range": false,
"refId": "A"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Objects Seen vs Skipped",
"description": "A sweep that proposes nothing because there was nothing to do and one that proposes nothing because it could read nothing produce the same proposal count. Skips are the difference, and are the thing to alert on.",
"type": "timeseries",
"id": 221,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 32
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type) (rate(SeaweedFS_worker_objects_seen_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "seen {{job_type}}"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type, reason) (rate(SeaweedFS_worker_objects_skipped_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "B",
"legendFormat": "skipped {{job_type}} ({{reason}})"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Detections and Proposals",
"description": "Detection sweeps by outcome, and the work they handed to admin.",
"type": "timeseries",
"id": 222,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 32
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type, result) (rate(SeaweedFS_worker_detections_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "detections {{job_type}} ({{result}})"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type) (rate(SeaweedFS_worker_proposals_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "B",
"legendFormat": "proposals {{job_type}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Jobs Executed",
"description": "Jobs run by these workers, by outcome.",
"type": "timeseries",
"id": 223,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 40
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (job_type, result) (rate(SeaweedFS_worker_jobs_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "{{job_type}} ({{result}})"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Job Duration p99",
"description": "How long a job takes, per job type.",
"type": "timeseries",
"id": 224,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 40
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "s",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "histogram_quantile(0.99, sum(rate(SeaweedFS_worker_job_seconds_bucket{cluster=~\"$cluster\"}[$__rate_interval])) by (le, job_type))",
"range": true,
"refId": "A",
"legendFormat": "{{job_type}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Worker Slots",
"description": "Capacity a worker advertised and how much of it is in use. Held slots are what admin schedules against.",
"type": "timeseries",
"id": 225,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 48
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "short",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (lane) (SeaweedFS_worker_slots_used{cluster=~\"$cluster\"})",
"range": true,
"refId": "A",
"legendFormat": "used {{lane}}"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (lane) (SeaweedFS_worker_slots_total{cluster=~\"$cluster\"})",
"range": true,
"refId": "B",
"legendFormat": "total {{lane}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Control Stream Events",
"description": "Connects, closes and failures. A worker that reconnects steadily is usually two workers sharing one id, evicting each other.",
"type": "timeseries",
"id": 226,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 48
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum by (event) (rate(SeaweedFS_worker_stream_events_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "{{event}}"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Lance Maintenance Reclaimed",
"description": "What the Lance jobs actually removed. A job that runs every minute and reclaims nothing is a different thing from a job that never runs.",
"type": "timeseries",
"id": 227,
"gridPos": {
"h": 8,
"w": 12,
"x": 0,
"y": 56
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "ops",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_fragments_removed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "fragments merged"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_rows_indexed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "B",
"legendFormat": "rows indexed"
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_versions_removed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "C",
"legendFormat": "versions removed"
}
]
},
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"title": "Storage Reclaimed",
"description": "Bytes freed by version cleanup. Its own panel because bytes and counts do not belong on one axis.",
"type": "timeseries",
"id": 228,
"gridPos": {
"h": 8,
"w": 12,
"x": 12,
"y": 56
},
"fieldConfig": {
"defaults": {
"color": {
"mode": "palette-classic"
},
"custom": {
"axisBorderShow": false,
"axisCenteredZero": false,
"axisColorMode": "text",
"axisLabel": "",
"axisPlacement": "auto",
"barAlignment": 0,
"drawStyle": "line",
"fillOpacity": 10,
"gradientMode": "none",
"hideFrom": {
"legend": false,
"tooltip": false,
"viz": false
},
"insertNulls": false,
"lineInterpolation": "linear",
"lineWidth": 1,
"pointSize": 4,
"scaleDistribution": {
"type": "linear"
},
"showPoints": "never",
"spanNulls": false,
"stacking": {
"group": "A",
"mode": "none"
},
"thresholdsStyle": {
"mode": "off"
}
},
"mappings": [],
"unit": "Bps",
"thresholds": {
"mode": "absolute",
"steps": [
{
"color": "green",
"value": null
}
]
}
},
"overrides": []
},
"options": {
"legend": {
"calcs": [
"lastNotNull",
"max"
],
"displayMode": "table",
"placement": "bottom",
"showLegend": true
},
"tooltip": {
"mode": "multi",
"sort": "desc"
}
},
"targets": [
{
"datasource": {
"type": "prometheus",
"uid": "${DS_PROMETHEUS}"
},
"editorMode": "code",
"expr": "sum(rate(SeaweedFS_worker_lance_bytes_reclaimed_total{cluster=~\"$cluster\"}[$__rate_interval]))",
"range": true,
"refId": "A",
"legendFormat": "reclaimed"
}
]
}
]
},
{
"type": "row",
"title": "Go Runtime",
@@ -9272,7 +10225,7 @@
"h": 1,
"w": 24,
"x": 0,
"y": 27
"y": 28
},
"panels": [
{
+65
View File
@@ -2929,6 +2929,70 @@ dependencies = [
"prost 0.13.5",
]
[[package]]
name = "protoc-bin-vendored"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d1c381df33c98266b5f08186583660090a4ffa0889e76c7e9a5e175f645a67fa"
dependencies = [
"protoc-bin-vendored-linux-aarch_64",
"protoc-bin-vendored-linux-ppcle_64",
"protoc-bin-vendored-linux-s390_64",
"protoc-bin-vendored-linux-x86_32",
"protoc-bin-vendored-linux-x86_64",
"protoc-bin-vendored-macos-aarch_64",
"protoc-bin-vendored-macos-x86_64",
"protoc-bin-vendored-win32",
]
[[package]]
name = "protoc-bin-vendored-linux-aarch_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c350df4d49b5b9e3ca79f7e646fde2377b199e13cfa87320308397e1f37e1a4c"
[[package]]
name = "protoc-bin-vendored-linux-ppcle_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a55a63e6c7244f19b5c6393f025017eb5d793fd5467823a099740a7a4222440c"
[[package]]
name = "protoc-bin-vendored-linux-s390_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1dba5565db4288e935d5330a07c264a4ee8e4a5b4a4e6f4e83fad824cc32f3b0"
[[package]]
name = "protoc-bin-vendored-linux-x86_32"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8854774b24ee28b7868cd71dccaae8e02a2365e67a4a87a6cd11ee6cdbdf9cf5"
[[package]]
name = "protoc-bin-vendored-linux-x86_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b38b07546580df720fa464ce124c4b03630a6fb83e05c336fea2a241df7e5d78"
[[package]]
name = "protoc-bin-vendored-macos-aarch_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "89278a9926ce312e51f1d999fee8825d324d603213344a9a706daa009f1d8092"
[[package]]
name = "protoc-bin-vendored-macos-x86_64"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "81745feda7ccfb9471d7a4de888f0652e806d5795b61480605d4943176299756"
[[package]]
name = "protoc-bin-vendored-win32"
version = "3.2.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "95067976aca6421a523e491fce939a3e65249bac4b977adee0ee9771568e8aa3"
[[package]]
name = "pxfm"
version = "0.1.28"
@@ -4529,6 +4593,7 @@ dependencies = [
"prometheus",
"prost 0.13.5",
"prost-types 0.13.5",
"protoc-bin-vendored",
"rand 0.10.2",
"redb",
"reed-solomon-erasure",
+4
View File
@@ -137,6 +137,10 @@ tempfile = "3"
[build-dependencies]
tonic-build = "0.12"
# Ships protoc with the build so neither CI nor a developer needs a system
# install, and so the version is pinned rather than whatever the platform's
# package manager happens to carry.
protoc-bin-vendored = "3"
[patch.crates-io]
reed-solomon-erasure = { path = "vendor/reed-solomon-erasure" }
+10 -2
View File
@@ -1,10 +1,18 @@
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Use the protoc that ships with protoc-bin-vendored rather than a system
// one, so the build needs no package manager and always sees the same
// version. An explicit PROTOC still wins, for packagers supplying their own.
if std::env::var_os("PROTOC").is_none() {
std::env::set_var("PROTOC", protoc_bin_vendored::protoc_bin_path()?);
}
let out_dir = std::path::PathBuf::from(std::env::var("OUT_DIR")?);
tonic_build::configure()
.build_server(true)
.build_client(true)
// filer.proto uses proto3 optional; older protoc (e.g. 3.12 from ubuntu-22.04 apt)
// rejects it without this flag, and newer protoc still accepts the flag
// filer.proto uses proto3 optional, which protoc rejects without this
// flag before 3.15. The vendored protoc is newer, but it still accepts
// the flag, so this keeps a build against an older PROTOC working.
.protoc_arg("--experimental_allow_proto3_optional")
.file_descriptor_set_path(out_dir.join("seaweed_descriptor.bin"))
.compile_protos(
+11
View File
@@ -45,6 +45,8 @@ service VolumeServer {
}
rpc VolumeUnmount (VolumeUnmountRequest) returns (VolumeUnmountResponse) {
}
rpc VolumeConsolidateIndex (VolumeConsolidateIndexRequest) returns (VolumeConsolidateIndexResponse) {
}
rpc VolumeDelete (VolumeDeleteRequest) returns (VolumeDeleteResponse) {
}
rpc VolumeMarkReadonly (VolumeMarkReadonlyRequest) returns (VolumeMarkReadonlyResponse) {
@@ -244,6 +246,12 @@ message VolumeUnmountRequest {
message VolumeUnmountResponse {
}
message VolumeConsolidateIndexRequest {
uint32 volume_id = 1;
}
message VolumeConsolidateIndexResponse {
}
message VolumeDeleteRequest {
uint32 volume_id = 1;
bool only_empty = 2;
@@ -257,6 +265,8 @@ message VolumeDeleteResponse {
message VolumeMarkReadonlyRequest {
uint32 volume_id = 1;
bool persist = 2;
// reject writes but keep accepting deletes, so expiring data can drain the volume
bool can_delete = 3;
}
message VolumeMarkReadonlyResponse {
}
@@ -594,6 +604,7 @@ message VolumeInfo {
uint64 expire_at_sec = 6; // expiration time of ec volume
bool read_only = 7;
EcShardConfig ec_shard_config = 8; // EC shard configuration (optional, null = use default 10+4)
bool read_only_can_delete = 9; // with read_only: writes are rejected but deletes still land
}
// EcShardConfig specifies erasure coding shard configuration
+47
View File
@@ -256,6 +256,8 @@ pub struct VolumeServerConfig {
pub https_client_ca_file: String,
pub grpc_cert_file: String,
pub grpc_key_file: String,
pub grpc_client_cert_file: String,
pub grpc_client_key_file: String,
pub grpc_ca_file: String,
pub grpc_allowed_wildcard_domain: String,
pub grpc_volume_allowed_common_names: Vec<String>,
@@ -805,6 +807,8 @@ fn resolve_config(cli: Cli) -> VolumeServerConfig {
https_client_ca_file: sec.https_client_ca_file,
grpc_cert_file: sec.grpc_cert_file,
grpc_key_file: sec.grpc_key_file,
grpc_client_cert_file: sec.grpc_client_cert_file,
grpc_client_key_file: sec.grpc_client_key_file,
grpc_ca_file: sec.grpc_ca_file,
grpc_allowed_wildcard_domain: sec.grpc_allowed_wildcard_domain,
grpc_volume_allowed_common_names: sec.grpc_volume_allowed_common_names,
@@ -837,6 +841,8 @@ pub struct SecurityConfig {
pub https_client_ca_file: String,
pub grpc_cert_file: String,
pub grpc_key_file: String,
pub grpc_client_cert_file: String,
pub grpc_client_key_file: String,
pub grpc_ca_file: String,
pub grpc_allowed_wildcard_domain: String,
pub grpc_volume_allowed_common_names: Vec<String>,
@@ -882,6 +888,8 @@ const SECURITY_CONFIG_FILE_NAME: &str = "security.toml";
/// [grpc.volume]
/// cert = "/path/to/cert.pem"
/// key = "/path/to/key.pem"
/// client_cert = "/path/to/client-cert.pem"
/// client_key = "/path/to/client-key.pem"
/// allowed_commonNames = "volume-a.internal,volume-b.internal"
/// ```
pub fn parse_security_config(path: &str) -> SecurityConfig {
@@ -1003,6 +1011,8 @@ pub fn parse_security_config(path: &str) -> SecurityConfig {
Section::GrpcVolume => match key {
"cert" => cfg.grpc_cert_file = value.to_string(),
"key" => cfg.grpc_key_file = value.to_string(),
"client_cert" => cfg.grpc_client_cert_file = value.to_string(),
"client_key" => cfg.grpc_client_key_file = value.to_string(),
// Go only reads CA from [grpc], not [grpc.volume]
"allowed_commonNames" => {
cfg.grpc_volume_allowed_common_names =
@@ -1134,6 +1144,12 @@ fn apply_env_overrides(cfg: &mut SecurityConfig) {
if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_KEY") {
cfg.grpc_key_file = v;
}
if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_CLIENT_CERT") {
cfg.grpc_client_cert_file = v;
}
if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_CLIENT_KEY") {
cfg.grpc_client_key_file = v;
}
if let Ok(v) = std::env::var("WEED_GRPC_CA") {
cfg.grpc_ca_file = v;
} else if let Ok(v) = std::env::var("WEED_GRPC_VOLUME_CA") {
@@ -1231,6 +1247,8 @@ mod tests {
"WEED_HTTPS_CLIENT_CA",
"WEED_GRPC_VOLUME_CERT",
"WEED_GRPC_VOLUME_KEY",
"WEED_GRPC_VOLUME_CLIENT_CERT",
"WEED_GRPC_VOLUME_CLIENT_KEY",
"WEED_GRPC_CA",
"WEED_GRPC_VOLUME_CA",
"WEED_GRPC_ALLOWED_WILDCARD_DOMAIN",
@@ -1500,6 +1518,35 @@ key = "/etc/seaweedfs/volume-key.pem"
});
}
#[test]
fn test_parse_security_config_uses_grpc_volume_client_cert() {
let _guard = process_state_lock();
let tmp = tempfile::NamedTempFile::new().unwrap();
std::fs::write(
tmp.path(),
r#"
[grpc.volume]
cert = "/etc/seaweedfs/volume-cert.pem"
key = "/etc/seaweedfs/volume-key.pem"
client_cert = "/etc/seaweedfs/volume-client-cert.pem"
client_key = "/etc/seaweedfs/volume-client-key.pem"
"#,
)
.unwrap();
with_cleared_security_env(|| {
let cfg = parse_security_config(tmp.path().to_str().unwrap());
assert_eq!(
cfg.grpc_client_cert_file,
"/etc/seaweedfs/volume-client-cert.pem"
);
assert_eq!(
cfg.grpc_client_key_file,
"/etc/seaweedfs/volume-client-key.pem"
);
});
}
#[test]
fn test_parse_security_config_uses_grpc_peer_name_policy() {
let _guard = process_state_lock();
+10
View File
@@ -330,6 +330,16 @@ pub fn delete_collection_metrics(collection: &str) {
delete_partial_match_collection(&DISK_SIZE_GAUGE, collection);
}
/// Drop a collection's volume server series once its last volume leaves this
/// server. These gauges are only ever set for collections still present, so the
/// values from the heartbeat that saw the last volume would otherwise stand
/// until the process restarts.
pub fn delete_volume_server_collection_metrics(collection: &str) {
let _ = DISK_SIZE_GAUGE.remove_label_values(&[collection, DISK_SIZE_LABEL_NORMAL]);
let _ = DISK_SIZE_GAUGE.remove_label_values(&[collection, DISK_SIZE_LABEL_DELETED_BYTES]);
delete_partial_match_collection(&READ_ONLY_VOLUME_GAUGE, collection);
}
/// Remove all metric entries from a GaugeVec where the "collection" label matches.
/// This emulates Go's `DeletePartialMatch(prometheus.Labels{"collection": collection})`.
fn delete_partial_match_collection(gauge: &GaugeVec, collection: &str) {
@@ -17,7 +17,7 @@
//! narrow TOCTOU window (a hostname that resolves to a public IP here and then
//! flips to a blocked one when the SDK dials) remains as a follow-up.
use std::net::{IpAddr, Ipv4Addr};
use std::net::{IpAddr, Ipv4Addr, Ipv6Addr};
/// AWS/Azure/GCP IPv4 instance-metadata-service (IMDS) address. It is
/// link-local and thus already covered by [`is_link_local`], but is named
@@ -74,9 +74,52 @@ fn is_cgnat(ip: IpAddr) -> bool {
}
}
/// Returns the IPv4 address carried by an IPv6 transition address -- NAT64
/// 64:ff9b::/96 (RFC 6052), 6to4 2002::/16 (RFC 3056), Teredo 2001:0000::/32
/// (RFC 4380), and the deprecated IPv4-compatible ::/96 (RFC 4291) -- or None
/// when it is not one of those. IPv4-mapped ::ffff:0:0/96 is excluded; it is
/// already normalized via `to_ipv4_mapped`.
fn embedded_transition_ipv4(v6: Ipv6Addr) -> Option<Ipv4Addr> {
let o = v6.octets();
if o[0] == 0x00
&& o[1] == 0x64
&& o[2] == 0xff
&& o[3] == 0x9b
&& o[4..12].iter().all(|&b| b == 0)
{
return Some(Ipv4Addr::new(o[12], o[13], o[14], o[15]));
}
if o[0] == 0x20 && o[1] == 0x02 {
return Some(Ipv4Addr::new(o[2], o[3], o[4], o[5]));
}
if o[0] == 0x20 && o[1] == 0x01 && o[2] == 0x00 && o[3] == 0x00 {
// Teredo obfuscates the client IPv4 as its ones' complement.
return Some(Ipv4Addr::new(
o[12] ^ 0xff,
o[13] ^ 0xff,
o[14] ^ 0xff,
o[15] ^ 0xff,
));
}
if o[..12].iter().all(|&b| b == 0) {
// IPv4-compatible ::a.b.c.d; :: and ::1 are already handled by the
// unspecified / loopback checks before extraction runs.
return Some(Ipv4Addr::new(o[12], o[13], o[14], o[15]));
}
None
}
/// Returns an error if `ip` is not safe to dial from a server that can reach
/// cluster-internal hosts. Mirrors Go's `checkBlockedIP`.
pub fn check_blocked_ip(endpoint: &str, ip: IpAddr) -> Result<(), String> {
check_blocked_ip_policy(endpoint, ip, false)
}
/// Like [`check_blocked_ip`], but `allow_private` keeps RFC 1918 / CGNAT
/// reachable for callers whose target legitimately sits on an internal network
/// (peer volume servers), while still blocking loopback, link-local (IMDS) and
/// unspecified. Mirrors Go's `checkBlockedIPPolicy`.
pub fn check_blocked_ip_policy(endpoint: &str, ip: IpAddr, allow_private: bool) -> Result<(), String> {
// Normalize IPv4-mapped IPv6 (`::ffff:a.b.c.d`) to its IPv4 form so the
// IPv4 deny rules apply. The OS routes these to the embedded IPv4 address,
// so without this `::ffff:127.0.0.1` / `::ffff:169.254.169.254` would slip
@@ -112,17 +155,28 @@ pub fn check_blocked_ip(endpoint: &str, ip: IpAddr) -> Result<(), String> {
endpoint, ip
));
}
if is_private(ip) {
return Err(format!(
"remote endpoint {:?} resolves to private address {}",
endpoint, ip
));
if !allow_private {
if is_private(ip) {
return Err(format!(
"remote endpoint {:?} resolves to private address {}",
endpoint, ip
));
}
if is_cgnat(ip) {
return Err(format!(
"remote endpoint {:?} resolves to CGNAT address {}",
endpoint, ip
));
}
}
if is_cgnat(ip) {
return Err(format!(
"remote endpoint {:?} resolves to CGNAT address {}",
endpoint, ip
));
// IPv6 transition addresses embed an IPv4 destination that routes to the
// same host wherever the matching relay exists (common in IPv6-only cloud).
// to_ipv4_mapped above only covers ::ffff: mapped addresses, so pull the
// embedded IPv4 out of the other forms and re-check it against the rules.
if let IpAddr::V6(v6) = ip {
if let Some(v4) = embedded_transition_ipv4(v6) {
return check_blocked_ip_policy(endpoint, IpAddr::V4(v4), allow_private);
}
}
Ok(())
}
@@ -243,6 +297,59 @@ pub async fn validate_remote_endpoint(endpoint: &str) -> Result<(), String> {
}
}
/// Returns an error if `target` could redirect a replica upload away from a peer
/// volume server. The target must be a bare `host:port` -- a scheme, userinfo,
/// path, query or fragment can smuggle a different destination into the
/// formatted upload URL -- whose host is not loopback, link-local (IMDS) or
/// unspecified. Cluster peers legitimately sit on private networks, so RFC 1918
/// / CGNAT are allowed. Mirrors Go's `validateReplicaTarget`.
pub async fn validate_replica_target(target: &str) -> Result<(), String> {
let trimmed = target.trim();
if trimmed.is_empty() {
return Err("replica target is empty".to_string());
}
if trimmed.contains("://") || trimmed.contains(['/', '?', '#', '@', '\\']) {
return Err(format!("replica target {:?} must be a bare host:port", target));
}
// Require an explicit host:port, handling `[IPv6]:port`. A bracketless IPv6
// literal (which carries its own colons) is rejected; peers are addressed as
// `[ipv6]:port`, matching Go's net.SplitHostPort.
let host = if let Some(rest) = trimmed.strip_prefix('[') {
match rest.split_once(']') {
Some((h, port)) if port.starts_with(':') && port.len() > 1 => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
}
} else {
match trimmed.rsplit_once(':') {
Some((h, port)) if !port.is_empty() && !h.contains(':') => h,
_ => return Err(format!("replica target {:?} must be a bare host:port", target)),
}
};
if host.is_empty() {
return Err(format!("replica target {:?} has no host", target));
}
if is_blocked_imds_host(&host.to_ascii_lowercase()) {
return Err(format!(
"replica target {:?} targets instance metadata service",
target
));
}
if let Ok(ip) = host.parse::<IpAddr>() {
return check_blocked_ip_policy(target, ip, true);
}
let addrs = resolve_host(host).await?;
if addrs.is_empty() {
return Err(format!("resolve replica target host {:?}: no addresses", host));
}
for ip in addrs {
check_blocked_ip_policy(target, ip, true)?;
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
@@ -333,6 +440,49 @@ mod tests {
assert!(check_blocked_ip("e", ip("2606:4700:4700::1111")).is_ok());
}
#[test]
fn rejects_ipv6_transition_addresses() {
// Every transition form encoding an internal IPv4 must be blocked.
assert!(
check_blocked_ip("e", ip("64:ff9b::a9fe:a9fe")) // NAT64 -> IMDS
.unwrap_err()
.contains("metadata")
);
assert!(
check_blocked_ip("e", ip("64:ff9b::7f00:1")) // NAT64 -> loopback
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("2002:a00:1::")) // 6to4 -> 10.0.0.1
.unwrap_err()
.contains("private")
);
assert!(
check_blocked_ip("e", ip("2001:0:4136:e378:8000:63bf:80ff:fffe")) // Teredo -> loopback
.unwrap_err()
.contains("loopback")
);
assert!(
check_blocked_ip("e", ip("::7f00:1")) // IPv4-compatible -> loopback
.unwrap_err()
.contains("loopback")
);
// A NAT64 address outside the 64:ff9b::/96 well-known prefix is not
// decoded (its embedded IPv4 lives elsewhere), so it is left as-is.
assert!(check_blocked_ip("e", ip("64:ff9b:1::a9fe:a9fe")).is_ok());
// Every transition form embedding a public IPv4 (8.8.8.8) still passes:
// NAT64, 6to4, Teredo, IPv4-compatible.
assert!(check_blocked_ip("e", ip("64:ff9b::808:808")).is_ok());
assert!(check_blocked_ip("e", ip("2002:808:808::")).is_ok());
assert!(check_blocked_ip("e", ip("2001::f7f7:f7f7")).is_ok());
assert!(check_blocked_ip("e", ip("::808:808")).is_ok());
// Bracketed transition literal via the full endpoint path.
assert!(precheck_endpoint("http://[64:ff9b::a9fe:a9fe]/")
.unwrap_err()
.contains("metadata"));
}
#[test]
fn rejects_ipv4_mapped_ipv6() {
// IPv4-mapped IPv6 must be unmapped so the IPv4 rules catch it.
@@ -356,4 +506,71 @@ mod tests {
.unwrap_err()
.contains("loopback"));
}
#[test]
fn check_blocked_ip_policy_allows_private_peers() {
// The replica leg targets peer volume servers, which may be private.
assert!(check_blocked_ip_policy("e", ip("10.0.0.7"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("192.168.1.5"), true).is_ok());
assert!(check_blocked_ip_policy("e", ip("100.64.0.42"), true).is_ok());
// Loopback / IMDS / unspecified stay blocked even when private is allowed.
assert!(check_blocked_ip_policy("e", ip("127.0.0.1"), true)
.unwrap_err()
.contains("loopback"));
assert!(check_blocked_ip_policy("e", ip("169.254.169.254"), true)
.unwrap_err()
.contains("metadata"));
assert!(check_blocked_ip_policy("e", ip("0.0.0.0"), true)
.unwrap_err()
.contains("unspecified"));
}
#[tokio::test]
async fn validate_replica_target_rejects_and_allows() {
// A path plus a trailing ?a= would otherwise swallow ?type=replicate.
assert!(validate_replica_target("127.0.0.1:7000/status/x/?a=")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("http://10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("user@10.0.0.7:8080")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("10.0.0.7")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("peer.example.com")
.await
.unwrap_err()
.contains("bare host:port"));
assert!(validate_replica_target("127.0.0.1:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("[::1]:8080")
.await
.unwrap_err()
.contains("loopback"));
assert!(validate_replica_target("169.254.169.254:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("metadata:80")
.await
.unwrap_err()
.contains("metadata"));
assert!(validate_replica_target("")
.await
.unwrap_err()
.contains("empty"));
// Legitimate peer volume servers on private networks pass.
assert!(validate_replica_target("10.0.0.7:8080").await.is_ok());
assert!(validate_replica_target("192.168.1.5:8080").await.is_ok());
assert!(validate_replica_target("[fd00::1]:8080").await.is_ok());
}
}
+33 -1
View File
@@ -7,7 +7,7 @@ pub mod endpoint_guard;
pub mod s3;
pub mod s3_tier;
pub use endpoint_guard::validate_remote_endpoint;
pub use endpoint_guard::{validate_remote_endpoint, validate_replica_target};
use crate::pb::remote_pb::{RemoteConf, RemoteStorageLocation};
@@ -211,4 +211,36 @@ mod tests {
};
assert_eq!(s3_compatible_endpoint(&gcs), None);
}
#[test]
fn azure_endpoint_has_no_ssrf_path() {
// The Go volume server guards the caller-supplied azure endpoint against
// SSRF. This server has no azure backend, so there is nothing to dial:
// azure is not S3-compatible (the endpoint guard does not apply) and
// make_remote_storage_client rejects the type before building a client.
let azure = RemoteConf {
r#type: "azure".to_string(),
azure_endpoint: "https://169.254.169.254/".to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&azure), None);
assert!(make_remote_storage_client(&azure).is_err());
}
#[test]
fn gcs_credentials_have_no_ssrf_path() {
// The Go volume server accepts only static-key gcs credentials and puts
// their token endpoint behind the SSRF guard, because the SDK dials
// whatever url, file or executable the credentials name. This server has
// no gcs backend, so make_remote_storage_client rejects the type before
// any credentials are parsed. Anyone adding one must carry both guards
// over with it.
let gcs = RemoteConf {
r#type: "gcs".to_string(),
gcs_google_application_credentials: r#"{"type":"external_account","credential_source":{"url":"http://169.254.169.254/"}}"#.to_string(),
..Default::default()
};
assert_eq!(s3_compatible_endpoint(&gcs), None);
assert!(make_remote_storage_client(&gcs).is_err());
}
}
+7
View File
@@ -192,6 +192,13 @@ impl Guard {
self.is_write_active = !is_empty_whitelist || !self.signing_key.is_empty();
}
/// Whether any write-side control is configured: a non-empty whitelist or a
/// signing key. Mirrors Go's non-nil `guard` -- when this is false, every
/// caller is allowed and admin RPCs need not carry peer info.
pub fn is_write_active(&self) -> bool {
self.is_write_active
}
/// Check if a remote IP is in the whitelist.
/// Returns true if write security is inactive (no whitelist and no signing key),
/// if the whitelist is empty, or if the IP matches.
+56 -7
View File
@@ -33,23 +33,31 @@ impl Error for GrpcClientError {}
pub fn load_outgoing_grpc_tls(
config: &VolumeServerConfig,
) -> Result<Option<OutgoingGrpcTlsConfig>, GrpcClientError> {
if config.grpc_cert_file.is_empty()
|| config.grpc_key_file.is_empty()
|| config.grpc_ca_file.is_empty()
// prefer a dedicated client certificate: CAs may issue certs with only one of the serverAuth/clientAuth EKUs
let (cert_file, key_file) = if !config.grpc_client_cert_file.is_empty()
&& !config.grpc_client_key_file.is_empty()
{
(&config.grpc_client_cert_file, &config.grpc_client_key_file)
} else {
if !config.grpc_client_cert_file.is_empty() || !config.grpc_client_key_file.is_empty() {
tracing::warn!("grpc.volume.client_cert and grpc.volume.client_key must both be set, falling back to grpc.volume.cert and grpc.volume.key");
}
(&config.grpc_cert_file, &config.grpc_key_file)
};
if cert_file.is_empty() || key_file.is_empty() || config.grpc_ca_file.is_empty() {
return Ok(None);
}
let cert_pem = std::fs::read_to_string(&config.grpc_cert_file).map_err(|e| {
let cert_pem = std::fs::read_to_string(cert_file).map_err(|e| {
GrpcClientError(format!(
"Failed to read outgoing gRPC cert '{}': {}",
config.grpc_cert_file, e
cert_file, e
))
})?;
let key_pem = std::fs::read_to_string(&config.grpc_key_file).map_err(|e| {
let key_pem = std::fs::read_to_string(key_file).map_err(|e| {
GrpcClientError(format!(
"Failed to read outgoing gRPC key '{}': {}",
config.grpc_key_file, e
key_file, e
))
})?;
let ca_pem = std::fs::read_to_string(&config.grpc_ca_file).map_err(|e| {
@@ -236,6 +244,8 @@ mod tests {
https_client_ca_file: String::new(),
grpc_cert_file: String::new(),
grpc_key_file: String::new(),
grpc_client_cert_file: String::new(),
grpc_client_key_file: String::new(),
grpc_ca_file: String::new(),
grpc_allowed_wildcard_domain: String::new(),
grpc_volume_allowed_common_names: vec![],
@@ -265,6 +275,45 @@ mod tests {
assert!(load_outgoing_grpc_tls(&config).unwrap().is_none());
}
fn write_pem_files(dir: &tempfile::TempDir, config: &mut VolumeServerConfig) {
let write = |name: &str, content: &str| {
let path = dir.path().join(name);
std::fs::write(&path, content).unwrap();
path.to_str().unwrap().to_string()
};
config.grpc_cert_file = write("server.pem", "server-cert");
config.grpc_key_file = write("server.key", "server-key");
config.grpc_ca_file = write("ca.pem", "ca");
}
#[test]
fn test_load_outgoing_grpc_tls_prefers_client_cert() {
let dir = tempfile::TempDir::new().unwrap();
let mut config = sample_config();
write_pem_files(&dir, &mut config);
let client_cert = dir.path().join("client.pem");
let client_key = dir.path().join("client.key");
std::fs::write(&client_cert, "client-cert").unwrap();
std::fs::write(&client_key, "client-key").unwrap();
config.grpc_client_cert_file = client_cert.to_str().unwrap().to_string();
config.grpc_client_key_file = client_key.to_str().unwrap().to_string();
let tls = load_outgoing_grpc_tls(&config).unwrap().unwrap();
assert_eq!(tls.cert_pem, "client-cert");
assert_eq!(tls.key_pem, "client-key");
}
#[test]
fn test_load_outgoing_grpc_tls_falls_back_to_server_cert() {
let dir = tempfile::TempDir::new().unwrap();
let mut config = sample_config();
write_pem_files(&dir, &mut config);
let tls = load_outgoing_grpc_tls(&config).unwrap().unwrap();
assert_eq!(tls.cert_pem, "server-cert");
assert_eq!(tls.key_pem, "server-key");
}
#[test]
fn test_build_grpc_endpoint_without_tls_uses_http_scheme() {
let endpoint = build_grpc_endpoint("127.0.0.1:19333", None).unwrap();
+92 -10
View File
@@ -155,6 +155,13 @@ impl VolumeGrpcService {
/// SocketAddr; if it is somehow None we deny, matching "if we don't
/// know who the caller is, refuse."
fn check_grpc_admin_auth<T>(&self, request: &Request<T>) -> Result<(), Status> {
// Mirror Go's `if vs.guard == nil { return nil }`: with no whitelist and
// no signing key, write security is inactive and every caller is allowed.
// Real gRPC connections always carry peer info, so requiring it below only
// affects in-process callers (tests, upgrades) once a control is enabled.
if !self.state.guard.read().unwrap().is_write_active() {
return Ok(());
}
let remote = match request.remote_addr() {
Some(addr) => addr,
None => {
@@ -237,12 +244,17 @@ impl VolumeGrpcService {
Ok(())
}
/// Shared helper matching Go's `makeVolumeReadonly(ctx, v, persist)`.
/// Shared helper matching Go's `makeVolumeReadonly(ctx, v, canDelete, persist)`.
/// 1. Check maintenance mode
/// 2. Notify master (readonly=true)
/// 3. Mark local volume readonly
/// 4. Notify master again (cover heartbeat race)
async fn make_volume_readonly(&self, vid: VolumeId, persist: bool) -> Result<(), Status> {
async fn make_volume_readonly(
&self,
vid: VolumeId,
can_delete: bool,
persist: bool,
) -> Result<(), Status> {
self.state.check_maintenance()?;
let info = {
@@ -268,7 +280,7 @@ impl VolumeGrpcService {
{
let mut store = self.state.store.write().unwrap();
if let Some((_, vol)) = store.find_volume_mut(vid) {
vol.set_read_only_persist(persist)
vol.set_read_only_persist(can_delete, persist)
.map_err(|e| Status::internal(e.to_string()))?;
}
self.state.volume_state_notify.notify_one();
@@ -288,6 +300,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::BatchDeleteRequest>,
) -> Result<Response<volume_server_pb::BatchDeleteResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let mut results = Vec::new();
@@ -939,6 +952,22 @@ impl VolumeServer for VolumeGrpcService {
Ok(Response::new(volume_server_pb::VolumeUnmountResponse {}))
}
async fn volume_consolidate_index(
&self,
request: Request<volume_server_pb::VolumeConsolidateIndexRequest>,
) -> Result<Response<volume_server_pb::VolumeConsolidateIndexResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let vid = VolumeId(request.into_inner().volume_id);
let mut store = self.state.store.write().unwrap();
store
.consolidate_volume_index(vid)
.map_err(|e| Status::internal(e.to_string()))?;
Ok(Response::new(
volume_server_pb::VolumeConsolidateIndexResponse {},
))
}
async fn volume_delete(
&self,
request: Request<volume_server_pb::VolumeDeleteRequest>,
@@ -977,7 +1006,8 @@ impl VolumeServer for VolumeGrpcService {
.find_volume(vid)
.ok_or_else(|| Status::not_found(format!("volume {} not found", vid)))?;
}
self.make_volume_readonly(vid, req.persist).await?;
self.make_volume_readonly(vid, req.can_delete, req.persist)
.await?;
Ok(Response::new(
volume_server_pb::VolumeMarkReadonlyResponse {},
))
@@ -1129,6 +1159,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::SetStateRequest>,
) -> Result<Response<volume_server_pb::SetStateResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
if let Some(new_state) = &req.state {
@@ -1181,6 +1212,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeCopyRequest>,
) -> Result<Response<Self::VolumeCopyStream>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -1982,6 +2014,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::ReadAllNeedlesRequest>,
) -> Result<Response<Self::ReadAllNeedlesStream>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let state = self.state.clone();
@@ -2182,6 +2215,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeTailReceiverRequest>,
) -> Result<Response<volume_server_pb::VolumeTailReceiverResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2292,7 +2326,7 @@ impl VolumeServer for VolumeGrpcService {
// Write needle to local volume
let mut store = state.store.write().unwrap();
store
.write_volume_needle(vid, &mut n)
.write_volume_needle(vid, &mut n, false)
.map_err(|e| Status::internal(format!("write needle: {}", e)))?;
}
@@ -2307,6 +2341,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsGenerateRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsGenerateResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2407,6 +2442,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsRebuildRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsRebuildResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2607,6 +2643,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsCopyRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsCopyResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -2989,6 +3026,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsUnmountRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsUnmountResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3152,6 +3190,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeEcShardsToVolumeRequest>,
) -> Result<Response<volume_server_pb::VolumeEcShardsToVolumeResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3418,6 +3457,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeTierMoveDatToRemoteRequest>,
) -> Result<Response<Self::VolumeTierMoveDatToRemoteStream>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3584,6 +3624,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeTierMoveDatFromRemoteRequest>,
) -> Result<Response<Self::VolumeTierMoveDatFromRemoteStream>, Status> {
self.check_grpc_admin_auth(&request)?;
// Note: Go does NOT check maintenance mode for TierMoveDatFromRemote
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3844,6 +3885,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::FetchAndWriteNeedleRequest>,
) -> Result<Response<volume_server_pb::FetchAndWriteNeedleResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
self.state.check_maintenance()?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
@@ -3919,6 +3961,21 @@ impl VolumeServer for VolumeGrpcService {
.as_secs();
n.set_has_last_modified_date();
// Validate every replica target before writing anything, so a malformed
// or internal target fails the request instead of leaving a local write
// behind. Mirrors the Go volume server. NOTE: like the Rust S3 endpoint
// guard, this validates the up-front DNS answer but does not yet re-check
// at connect time, so a rebinding hostname remains a follow-up.
if !self.state.allow_untrusted_remote_endpoints {
for replica in &req.replicas {
crate::remote_storage::validate_replica_target(&replica.url)
.await
.map_err(|e| {
Status::invalid_argument(format!("reject replica target: {}", e))
})?;
}
}
// Run local write and replica writes concurrently (matches Go's WaitGroup)
let mut handles: Vec<tokio::task::JoinHandle<Result<(), String>>> = Vec::new();
@@ -3930,7 +3987,7 @@ impl VolumeServer for VolumeGrpcService {
let local_handle = tokio::task::spawn_blocking(move || {
let mut store = state_clone.store.write().unwrap();
store
.write_volume_needle(vid, &mut n_clone)
.write_volume_needle(vid, &mut n_clone, false)
.map(|_| ())
.map_err(|e| format!("local write needle {} size {}: {}", needle_id, size, e))
});
@@ -4067,7 +4124,7 @@ impl VolumeServer for VolumeGrpcService {
let mut errs: Vec<String> = Vec::new();
if req.mark_broken_volumes_readonly {
for vid in &broken_vids {
match self.make_volume_readonly(*vid, true).await {
match self.make_volume_readonly(*vid, false, true).await {
Ok(()) => {
details.push(format!("volume {} is now read-only", vid.0));
}
@@ -4099,6 +4156,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::ScrubEcVolumeRequest>,
) -> Result<Response<volume_server_pb::ScrubEcVolumeResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
// Validate mode
@@ -4456,6 +4514,7 @@ impl VolumeServer for VolumeGrpcService {
&self,
request: Request<volume_server_pb::VolumeNeedleStatusRequest>,
) -> Result<Response<volume_server_pb::VolumeNeedleStatusResponse>, Status> {
self.check_grpc_admin_auth(&request)?;
let req = request.into_inner();
let vid = VolumeId(req.volume_id);
let needle_id = NeedleId(req.needle_id);
@@ -5159,7 +5218,7 @@ mod tests {
data_size: "remote-incremental-copy".len() as u32,
..Needle::default()
};
volume.write_needle(&mut needle, true).unwrap();
volume.write_needle(&mut needle, true, false).unwrap();
volume.sync_to_disk().unwrap();
(
std::fs::read(volume.file_name(".dat")).unwrap(),
@@ -5327,7 +5386,7 @@ mod tests {
data_size: b"ec-generate".len() as u32,
..Needle::default()
};
volume.write_needle(&mut needle, true).unwrap();
volume.write_needle(&mut needle, true, false).unwrap();
volume.sync_to_disk().unwrap();
}
@@ -5388,6 +5447,29 @@ mod tests {
(VolumeGrpcService { state }, tmp)
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn test_volume_consolidate_index_rpc() {
let (service, _tmp) = make_local_service_with_volume("consolidate_rpc", None);
// Carry a peer address so the admin gate sees a caller; the test guard
// has an empty whitelist, so any peer is accepted.
let mut request = Request::new(volume_server_pb::VolumeConsolidateIndexRequest {
volume_id: 1,
});
request.extensions_mut().insert(tonic::transport::server::TcpConnectInfo {
local_addr: None,
remote_addr: Some("127.0.0.1:65000".parse().unwrap()),
});
service.volume_consolidate_index(request).await.unwrap();
// Data and index share a directory here, so there is nothing to move;
// the volume stays mounted and keeps its needle.
let store = service.state.store.read().unwrap();
let (_, v) = store.find_volume(VolumeId(1)).unwrap();
assert_eq!(v.file_count(), 1);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn test_volume_incremental_copy_streams_remote_only_volume_data() {
let (service, _tmp, shutdown_tx, dat_bytes, super_block_size, _delete_count) =
@@ -5527,7 +5609,7 @@ mod tests {
data: payload,
..Needle::default()
};
v.write_needle(&mut needle, true).unwrap();
v.write_needle(&mut needle, true, false).unwrap();
v.sync_to_disk().unwrap();
}
let dat_path = {
+28 -2
View File
@@ -2580,11 +2580,17 @@ pub async fn post_handler(
n.set_has_name();
}
// A durable write flushes before it is acked. Read it the way Go's
// r.FormValue does, off the decoded fields, so a percent-encoded value is
// honored here too. ReplicatedWrite forwards the parameter, so a replica
// sees it the same way the primary did.
let fsync = form_value("fsync").as_deref() == Some("true");
let write_result = if let Some(wq) = state.write_queue.get() {
wq.submit(vid, n.clone()).await
wq.submit(vid, n.clone(), fsync).await
} else {
let mut store = state.store.write().unwrap();
store.write_volume_needle(vid, &mut n)
store.write_volume_needle(vid, &mut n, fsync)
};
// Replicate to remote volume servers if this volume has replicas.
@@ -3852,6 +3858,26 @@ fn parse_content_disposition_filename(value: &str) -> Option<String> {
mod tests {
use super::*;
/// The upload handler reads fsync off the decoded query fields rather than
/// matching the raw string, because Go's r.FormValue decodes and a raw
/// match would silently drop a percent-encoded value.
#[test]
fn test_encoded_query_field_decodes() {
let raw = "fsync=%74rue";
assert!(
!raw.split('&').any(|p| p == "fsync=true"),
"a raw match is exactly what misses this"
);
let fields: Vec<(String, String)> = serde_urlencoded::from_str(raw).unwrap();
assert_eq!(
fields
.iter()
.find(|(k, _)| k == "fsync")
.map(|(_, v)| v.as_str()),
Some("true")
);
}
#[test]
fn test_parse_url_path_comma() {
let (vid, nid, cookie) = parse_url_path("/3,01637037d6").unwrap();
+130 -20
View File
@@ -771,6 +771,10 @@ fn collect_volume_snapshot(
build_heartbeat_with_ec_status(config, &mut store, Vec::new(), true, false).1
}
/// The heartbeat alone, without the volume snapshot the send loop pairs it
/// with. Only the tests want it that way; the loop calls
/// collect_heartbeat_with_snapshot directly.
#[cfg(test)]
fn collect_heartbeat(
config: &HeartbeatConfig,
state: &Arc<VolumeServerState>,
@@ -838,8 +842,7 @@ fn build_heartbeat_with_ec_status(
// master can tell whether applying what it was sent leaves it current.
// Volumes skipped below -- quarantined, phantom, expired -- are in neither.
let mut volume_digest: u64 = 0;
let (send_full_list, report_generation) = store.volume_report.begin();
let mut reported_hashes: HashMap<VolumeReportKey, u64> = HashMap::new();
let (send_full_list, report_generation, report_pass) = store.volume_report.begin();
let mut changed_volumes = Vec::new();
let mut max_file_key = NeedleId(0);
let mut max_volume_counts: HashMap<String, u32> = HashMap::new();
@@ -937,8 +940,14 @@ fn build_heartbeat_with_ec_status(
let hash = report_hash(&volume_message);
volume_digest ^= hash;
let key: VolumeReportKey = (volume_message.disk_id, volume_message.id);
reported_hashes.insert(key, hash);
if send_full_list || store.volume_report.changed(key, hash) {
// A snapshot must leave the reporting state as it found it, so
// it asks rather than marks.
let is_news = if commit_report {
store.volume_report.record(key, hash, report_pass)
} else {
store.volume_report.changed(key, hash)
};
if send_full_list || is_news {
changed_volumes.push(volume_message.clone());
}
volumes.push(volume_message);
@@ -947,24 +956,30 @@ fn build_heartbeat_with_ec_status(
should_delete_volume = true;
}
// Track disk size by collection
let entry = disk_sizes.entry(vol.collection.clone()).or_insert((0, 0));
// Track disk size by collection. A volume on its way out is left
// out: an entry here is also what says the collection is still on
// this server.
if !should_delete_volume {
let entry = disk_sizes.entry(vol.collection.clone()).or_insert((0, 0));
entry.0 += volume_size;
entry.1 += vol.deleted_size();
}
let read_only = ro_counts.entry(vol.collection.clone()).or_default();
if !should_delete_volume && vol.is_read_only() {
read_only.is_read_only += 1;
if vol.is_no_write_or_delete() {
read_only.no_write_or_delete += 1;
}
if vol.is_no_write_can_delete() {
read_only.no_write_can_delete += 1;
}
if loc.is_disk_space_low.load(Ordering::Relaxed) {
read_only.is_disk_space_low += 1;
// An entry here is what says the collection is still on this
// server, so a volume on its way out must not make one.
if !should_delete_volume {
let read_only = ro_counts.entry(vol.collection.clone()).or_default();
if vol.is_read_only() {
read_only.is_read_only += 1;
if vol.is_no_write_or_delete() {
read_only.no_write_or_delete += 1;
}
if vol.is_no_write_can_delete() {
read_only.no_write_can_delete += 1;
}
if loc.is_disk_space_low.load(Ordering::Relaxed) {
read_only.is_disk_space_low += 1;
}
}
}
@@ -998,6 +1013,17 @@ fn build_heartbeat_with_ec_status(
.with_label_values(&[col, crate::metrics::READ_ONLY_LABEL_IS_DISK_SPACE_LOW])
.set(counts.is_disk_space_low as f64);
}
// ro_counts has an entry for every collection that kept a volume through
// this pass, including the ones counting zero read-only volumes.
{
let mut reported = store.reported_collections.lock().unwrap();
for col in reported.iter() {
if !ro_counts.contains_key(col) {
crate::metrics::delete_volume_server_collection_metrics(col);
}
}
*reported = ro_counts.keys().cloned().collect();
}
// Update max volumes gauge
let total_max: i64 = max_volume_counts.values().map(|v| *v as i64).sum();
crate::metrics::MAX_VOLUMES.set(total_max);
@@ -1005,7 +1031,7 @@ fn build_heartbeat_with_ec_status(
// Only when this heartbeat is going to be sent: marking volumes reported
// and then discarding the message would leave the master never told.
if commit_report {
store.volume_report.commit(reported_hashes, report_generation);
store.volume_report.commit(report_pass, report_generation);
}
// has_no_volumes says the server holds nothing, so it may only be derived
@@ -1073,6 +1099,14 @@ fn collect_live_ec_shards(
.with_label_values(&[col, crate::metrics::DISK_SIZE_LABEL_EC])
.set(*size as f64);
}
let mut reported = store.reported_ec_collections.lock().unwrap();
for col in reported.iter() {
if !ec_sizes.contains_key(col) {
let _ = crate::metrics::DISK_SIZE_GAUGE
.remove_label_values(&[col, crate::metrics::DISK_SIZE_LABEL_EC]);
}
}
*reported = ec_sizes.keys().cloned().collect();
}
ec_shards
@@ -1424,10 +1458,10 @@ mod tests {
store.volume_report.accept_deltas();
build_heartbeat(&test_config(), &mut store);
let (full, generation) = store.volume_report.begin();
let (full, generation, pass) = store.volume_report.begin();
assert!(!full);
store.volume_report.request_full_list();
store.volume_report.commit(HashMap::new(), generation);
store.volume_report.commit(pass, generation);
let heartbeat = build_heartbeat(&test_config(), &mut store);
assert_eq!(heartbeat.volumes.len(), 2);
@@ -1592,6 +1626,82 @@ mod tests {
);
}
fn collection_series(gauge: &prometheus::GaugeVec, collection: &str) -> usize {
use prometheus::core::Collector;
gauge
.collect()
.iter()
.flat_map(|family| family.get_metric().to_vec())
.filter(|metric| {
metric
.get_label()
.iter()
.any(|label| label.get_name() == "collection" && label.get_value() == collection)
})
.count()
}
// The per-collection gauges are only ever set for collections the heartbeat
// still finds on this server. A volume.balance that moves a collection's
// last volume off a server used to leave its read-only count - marked
// read-only for the move, moments before it went - standing on that server
// until a restart, with nothing in volume.list to match it.
#[test]
fn test_build_heartbeat_clears_metrics_of_departed_collection() {
let temp_dir = tempfile::tempdir().unwrap();
let dir = temp_dir.path().to_str().unwrap();
let collection = "heartbeat_departed_case";
let mut store = Store::new(NeedleMapKind::InMemory);
store
.add_location(
dir,
dir,
8,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
store
.add_volume(
VolumeId(21),
collection,
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
{
let (_, volume) = store.find_volume_mut(VolumeId(21)).unwrap();
volume.set_read_only().unwrap();
}
build_heartbeat(&test_config(), &mut store);
assert_eq!(
READ_ONLY_VOLUME_GAUGE
.with_label_values(&[collection, READ_ONLY_LABEL_IS_READ_ONLY])
.get(),
1.0
);
assert!(store.unmount_volume(VolumeId(21)));
build_heartbeat(&test_config(), &mut store);
assert_eq!(
collection_series(&READ_ONLY_VOLUME_GAUGE, collection),
0,
"read-only series left after the collection left the server"
);
assert_eq!(
collection_series(&DISK_SIZE_GAUGE, collection),
0,
"disk size series left after the collection left the server"
);
}
#[test]
fn test_build_heartbeat_reports_disk_bytes() {
let temp_dir = tempfile::tempdir().unwrap();
+2
View File
@@ -163,6 +163,8 @@ mod tests {
https_client_ca_file: String::new(),
grpc_cert_file: String::new(),
grpc_key_file: String::new(),
grpc_client_cert_file: String::new(),
grpc_client_key_file: String::new(),
grpc_ca_file: String::new(),
grpc_allowed_wildcard_domain: String::new(),
grpc_volume_allowed_common_names: vec![],
+20 -14
View File
@@ -1,9 +1,10 @@
//! Async batched write processing for the volume server.
//!
//! Instead of each upload handler directly calling `write_needle` and syncing,
//! writes are submitted to a queue. A background worker drains the queue in
//! batches (up to 128 entries), groups them by volume ID, processes them
//! together, and syncs once per volume for the entire batch.
//! Instead of each upload handler directly calling `write_needle`, writes are
//! submitted to a queue. A background worker drains the queue in batches (up to
//! 128 entries), groups them by volume ID, and processes them together under a
//! single store lock. Requests that asked for `fsync` are flushed by
//! `write_needle` itself, one flush per durable write.
use std::sync::Arc;
@@ -19,10 +20,14 @@ use super::volume_server::VolumeServerState;
/// Result of a single write operation: (offset, size, is_unchanged).
pub type WriteResult = Result<(u64, Size, bool), VolumeError>;
/// The needles queued for one volume, each with the durability it asked for.
type VolumeBatch = Vec<(Needle, bool, oneshot::Sender<WriteResult>)>;
/// A request to write a needle, submitted to the write queue.
pub struct WriteRequest {
pub volume_id: VolumeId,
pub needle: Needle,
pub fsync: bool,
pub response_tx: oneshot::Sender<WriteResult>,
}
@@ -54,11 +59,12 @@ impl WriteQueue {
/// Submit a write request and wait for the result.
///
/// Returns `Err` if the worker has shut down or the response channel was dropped.
pub async fn submit(&self, volume_id: VolumeId, needle: Needle) -> WriteResult {
pub async fn submit(&self, volume_id: VolumeId, needle: Needle, fsync: bool) -> WriteResult {
let (response_tx, response_rx) = oneshot::channel();
let request = WriteRequest {
volume_id,
needle,
fsync,
response_tx,
};
@@ -138,14 +144,14 @@ fn process_batch(state: Arc<VolumeServerState>, batch: Vec<WriteRequest>) {
// Group requests by volume ID for efficient processing.
// We use a Vec of (VolumeId, Vec<(Needle, Sender)>) to preserve order
// and avoid requiring Hash on VolumeId.
let mut groups: Vec<(VolumeId, Vec<(Needle, oneshot::Sender<WriteResult>)>)> = Vec::new();
let mut groups: Vec<(VolumeId, VolumeBatch)> = Vec::new();
for req in batch {
let vid = req.volume_id;
if let Some(group) = groups.iter_mut().find(|(v, _)| *v == vid) {
group.1.push((req.needle, req.response_tx));
group.1.push((req.needle, req.fsync, req.response_tx));
} else {
groups.push((vid, vec![(req.needle, req.response_tx)]));
groups.push((vid, vec![(req.needle, req.fsync, req.response_tx)]));
}
}
@@ -153,8 +159,8 @@ fn process_batch(state: Arc<VolumeServerState>, batch: Vec<WriteRequest>) {
let mut store = state.store.write().unwrap();
for (vid, entries) in groups {
for (mut needle, response_tx) in entries {
let result = store.write_volume_needle(vid, &mut needle);
for (mut needle, fsync, response_tx) in entries {
let result = store.write_volume_needle(vid, &mut needle, fsync);
// Send result back; ignore error if receiver dropped.
let _ = response_tx.send(result);
}
@@ -239,7 +245,7 @@ mod tests {
..Needle::default()
};
let result = queue.submit(VolumeId(999), needle).await;
let result = queue.submit(VolumeId(999), needle, false).await;
assert!(result.is_err());
match result {
Err(VolumeError::NotFound) => {} // expected
@@ -264,7 +270,7 @@ mod tests {
data_size: 10,
..Needle::default()
};
q.submit(VolumeId(1), needle).await
q.submit(VolumeId(1), needle, false).await
}));
}
@@ -293,7 +299,7 @@ mod tests {
data_size: 4,
..Needle::default()
};
q.submit(VolumeId(42), needle).await
q.submit(VolumeId(42), needle, false).await
}));
}
@@ -327,7 +333,7 @@ mod tests {
data_size: 0,
..Needle::default()
};
let result = queue2.submit(VolumeId(1), needle).await;
let result = queue2.submit(VolumeId(1), needle, false).await;
assert!(result.is_err()); // NotFound is fine -- the point is it doesn't panic
}
}
@@ -400,7 +400,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
let original_dat_size = v.dat_file_size().unwrap();
@@ -746,7 +746,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -798,7 +798,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut needle, true).unwrap();
v.write_needle(&mut needle, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -875,7 +875,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1116,7 +1116,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1198,7 +1198,7 @@ mod tests {
data_size: 5,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
v.close();
@@ -1509,7 +1509,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1585,7 +1585,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
@@ -1632,7 +1632,7 @@ mod tests {
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
}
v.sync_to_disk().unwrap();
v.close();
+106 -3
View File
@@ -4,8 +4,10 @@
//! It coordinates volume placement, lookup, and lifecycle operations.
//! Matches Go's storage/store.go.
use std::collections::HashSet;
use std::io;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::Mutex;
use crate::config::MinFreeSpace;
use crate::pb::master_pb;
@@ -32,6 +34,12 @@ pub struct Store {
pub data_center: String,
pub rack: String,
pub volume_report: crate::storage::volume_report::VolumeReportState,
/// Collections the last heartbeat set per-collection gauges for. Those
/// gauges are only ever set for collections still held here, so one whose
/// last volume leaves - moved away by volume.balance, say - would keep
/// reporting the heartbeat that saw it.
pub reported_collections: Mutex<HashSet<String>>,
pub reported_ec_collections: Mutex<HashSet<String>>,
}
impl Store {
@@ -49,6 +57,8 @@ impl Store {
volume_report: Default::default(),
data_center: String::new(),
rack: String::new(),
reported_collections: Mutex::new(HashSet::new()),
reported_ec_collections: Mutex::new(HashSet::new()),
}
}
@@ -156,6 +166,26 @@ impl Store {
self.find_volume(vid).is_some()
}
/// Move a volume's index into the configured `-dir.idx` directory, reloading
/// the volume in place. A no-op when the location has no separate index
/// directory. Mirrors Go's `Store::ConsolidateVolumeIndex`.
pub fn consolidate_volume_index(&mut self, vid: VolumeId) -> Result<(), VolumeError> {
for loc in self.locations.iter_mut() {
let idx_dir = loc.idx_directory.clone();
let data_dir = loc.directory.clone();
if let Some(v) = loc.find_volume_mut(vid) {
if idx_dir == data_dir {
return Ok(());
}
return v.relocate_index_to(&idx_dir);
}
}
Err(VolumeError::Io(io::Error::new(
io::ErrorKind::NotFound,
format!("volume {} not found on disk", vid),
)))
}
// ---- Volume lifecycle ----
/// Find the location with fewest volumes (load-balance) of the given disk type.
@@ -499,11 +529,13 @@ impl Store {
vol.re_lookup_needle_data_offset(needle_id)
}
/// Write a needle to a volume.
/// Write a needle to a volume. With `fsync` the volume flushes its .dat
/// before returning, so the caller can ack a durable write.
pub fn write_volume_needle(
&mut self,
vid: VolumeId,
n: &mut Needle,
fsync: bool,
) -> Result<(u64, Size, bool), VolumeError> {
// Check disk space on the location containing this volume.
// We do this before the mutable borrow to avoid borrow conflicts.
@@ -519,7 +551,7 @@ impl Store {
}
let (_, vol) = self.find_volume_mut(vid).ok_or(VolumeError::NotFound)?;
vol.write_needle(n, true)
vol.write_needle(n, true, fsync)
}
/// Delete a needle from a volume.
@@ -1365,7 +1397,9 @@ mod tests {
data_size: 11,
..Needle::default()
};
let (offset, _size, unchanged) = store.write_volume_needle(VolumeId(1), &mut n).unwrap();
let (offset, _size, unchanged) = store
.write_volume_needle(VolumeId(1), &mut n, false)
.unwrap();
assert!(!unchanged);
assert!(offset > 0);
@@ -1388,6 +1422,75 @@ mod tests {
assert!(deleted.0 > 0);
}
#[test]
fn test_consolidate_volume_index_not_found() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut store = make_test_store(&[dir]);
let err = store.consolidate_volume_index(VolumeId(9)).unwrap_err();
assert!(matches!(err, VolumeError::Io(ref e)
if e.kind() == std::io::ErrorKind::NotFound));
}
#[test]
fn test_consolidate_volume_index_noop_with_separate_idx_dir() {
let root = TempDir::new().unwrap();
let data_dir = root.path().join("data");
let idx_dir = root.path().join("idx");
std::fs::create_dir_all(&data_dir).unwrap();
std::fs::create_dir_all(&idx_dir).unwrap();
let mut store = Store::new(NeedleMapKind::InMemory);
store
.add_location(
data_dir.to_str().unwrap(),
idx_dir.to_str().unwrap(),
10,
DiskType::HardDrive,
MinFreeSpace::Percent(1.0),
Vec::new(),
)
.unwrap();
store
.add_volume(
VolumeId(1),
"",
None,
None,
0,
DiskType::HardDrive,
Version::current(),
)
.unwrap();
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(0xaa),
data: b"co-located".to_vec(),
data_size: 10,
..Needle::default()
};
store
.write_volume_needle(VolumeId(1), &mut n, false)
.unwrap();
// add_volume already placed the index in the -dir.idx directory, so
// consolidation has nothing to move and leaves the volume readable.
store.consolidate_volume_index(VolumeId(1)).unwrap();
let idx_file = format!(
"{}.idx",
volume_file_name(idx_dir.to_str().unwrap(), "", VolumeId(1))
);
assert!(std::path::Path::new(&idx_file).exists());
let mut got = Needle {
id: NeedleId(1),
..Needle::default()
};
let count = store.read_volume_needle(VolumeId(1), &mut got).unwrap();
assert_eq!(count, 10);
assert_eq!(got.data, b"co-located");
}
#[test]
fn test_store_multi_location() {
let tmp1 = TempDir::new().unwrap();
File diff suppressed because it is too large Load Diff
@@ -246,7 +246,7 @@ mod tests {
data: data.clone(),
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
written.push(data);
}
written
+87 -30
View File
@@ -1,7 +1,6 @@
//! Mirror of `weed/storage/store_volume_report.go`.
use std::collections::HashMap;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::Mutex;
/// Identifies one reported copy. Keyed by disk as well as id because a volume
@@ -9,6 +8,34 @@ use std::sync::Mutex;
/// other's changes untold.
pub type VolumeReportKey = (u32, u32);
/// What the master was told about one volume copy: the hash that detects
/// change, and the heartbeat pass that last found the copy held.
#[derive(Clone, Copy)]
struct ReportedVolume {
hash: u64,
pass: u64,
}
/// Everything a heartbeat reads and writes about what the master was told,
/// under one lock. The full-list flag and the generation that answers it have
/// to move together: split across atomics, a request landing between two of
/// them is answered by a heartbeat that never carried a list. Go holds a single
/// mutex over the same fields.
#[derive(Default)]
struct ReportState {
/// Set once the master says it compares digests. Until then the whole list
/// goes every time, which is what an older master needs.
deltas_accepted: bool,
full_list_needed: bool,
/// Counts requests for the whole list, so one arriving while a heartbeat is
/// being built is not marked satisfied by it.
full_list_generation: u64,
/// Numbers heartbeats, so one can mark the copies it finds held without
/// building a second map of them.
pass: u64,
last_reported: HashMap<VolumeReportKey, ReportedVolume>,
}
/// Remembers what the master was last told about each volume, so a heartbeat
/// can carry only what moved since.
///
@@ -18,56 +45,86 @@ pub type VolumeReportKey = (u32, u32);
/// until one accepts changes.
#[derive(Default)]
pub struct VolumeReportState {
/// Set once the master says it compares digests. Until then the whole list
/// goes every time, which is what an older master needs.
deltas_accepted: AtomicBool,
full_list_needed: AtomicBool,
/// Counts requests for the whole list, so one arriving while a heartbeat is
/// being built is not marked satisfied by it.
full_list_generation: AtomicU64,
last_reported: Mutex<HashMap<VolumeReportKey, u64>>,
state: Mutex<ReportState>,
}
impl VolumeReportState {
/// Drops everything known about the master's view.
pub fn reset(&self) {
self.deltas_accepted.store(false, Ordering::Relaxed);
self.full_list_needed.store(true, Ordering::Relaxed);
self.full_list_generation.fetch_add(1, Ordering::Relaxed);
self.last_reported.lock().unwrap().clear();
let mut state = self.state.lock().unwrap();
state.deltas_accepted = false;
state.full_list_needed = true;
state.full_list_generation += 1;
state.last_reported.clear();
}
pub fn accept_deltas(&self) {
self.deltas_accepted.store(true, Ordering::Relaxed);
self.state.lock().unwrap().deltas_accepted = true;
}
pub fn request_full_list(&self) {
self.full_list_needed.store(true, Ordering::Relaxed);
self.full_list_generation.fetch_add(1, Ordering::Relaxed);
let mut state = self.state.lock().unwrap();
state.full_list_needed = true;
state.full_list_generation += 1;
}
/// Reports whether this heartbeat must carry the whole list, and the
/// request it answers.
pub fn begin(&self) -> (bool, u64) {
let full = self.full_list_needed.load(Ordering::Relaxed)
|| !self.deltas_accepted.load(Ordering::Relaxed);
(full, self.full_list_generation.load(Ordering::Relaxed))
/// Opens a heartbeat: whether it must carry the whole list, the request it
/// answers, and the pass number that marks the copies it finds still held.
pub fn begin(&self) -> (bool, u64, u64) {
let mut state = self.state.lock().unwrap();
state.pass += 1;
(
state.full_list_needed || !state.deltas_accepted,
state.full_list_generation,
state.pass,
)
}
/// Reports whether the master needs telling about this volume, given what
/// it was last told.
/// it was last told. For a caller that is only taking a snapshot and so
/// must leave the reporting state alone; a heartbeat calls `record`.
pub fn changed(&self, key: VolumeReportKey, hash: u64) -> bool {
self.last_reported.lock().unwrap().get(&key) != Some(&hash)
self.state
.lock()
.unwrap()
.last_reported
.get(&key)
.is_none_or(|previous| previous.hash != hash)
}
/// Records what this heartbeat told the master. Volumes absent from
/// `reported` are forgotten, so one that comes back is reported again.
pub fn commit(&self, reported: HashMap<VolumeReportKey, u64>, generation: u64) {
*self.last_reported.lock().unwrap() = reported;
/// Marks one volume copy as held by the heartbeat being built, and reports
/// whether the master needs telling about it. It updates the entry already
/// held rather than build a second map beside it, so a server whose volumes
/// are quiet allocates nothing per volume per heartbeat.
pub fn record(&self, key: VolumeReportKey, hash: u64, pass: u64) -> bool {
let mut state = self.state.lock().unwrap();
match state.last_reported.get_mut(&key) {
Some(previous) => {
let changed = previous.hash != hash;
previous.hash = hash;
previous.pass = pass;
changed
}
None => {
state
.last_reported
.insert(key, ReportedVolume { hash, pass });
true
}
}
}
/// Closes the heartbeat. Copies this pass did not find are forgotten, so one
/// that comes back is reported again.
pub fn commit(&self, pass: u64, generation: u64) {
let mut state = self.state.lock().unwrap();
state
.last_reported
.retain(|_, reported| reported.pass == pass);
// A request that arrived while this heartbeat was being built asked
// about a later state than it carries, so it stands.
if self.full_list_generation.load(Ordering::Relaxed) == generation {
self.full_list_needed.store(false, Ordering::Relaxed);
if state.full_list_generation == generation {
state.full_list_needed = false;
}
}
}
@@ -35,8 +35,9 @@ pub fn report_hash(m: &master_pb::VolumeInformationMessage) -> u64 {
h = fold(h, xxh64(&(m.modified_at_second as u64).to_le_bytes(), 0));
h = fold(h, xxh64(m.collection.as_bytes(), 0));
h = fold(h, xxh64(m.disk_type.as_bytes(), 0));
// Not the remote storage key: the master does not keep it, so a change to
// it alters nothing its copy holds.
h = fold(h, xxh64(m.remote_storage_name.as_bytes(), 0));
h = fold(h, xxh64(m.remote_storage_key.as_bytes(), 0));
h
}
@@ -67,11 +68,11 @@ mod tests {
#[test]
fn report_hash_vectors() {
let empty = master_pb::VolumeInformationMessage::default();
assert_eq!(report_hash(&empty), 17122085700329870549);
assert_eq!(report_hash(&empty), 10988706248825469653);
let mut one = master_pb::VolumeInformationMessage::default();
one.id = 1;
assert_eq!(report_hash(&one), 12867601919960834066);
assert_eq!(report_hash(&one), 2035849960016744285);
let full = master_pb::VolumeInformationMessage {
id: 42,
@@ -91,7 +92,7 @@ mod tests {
disk_type: "ssd".to_string(),
disk_id: 2,
};
assert_eq!(report_hash(&full), 12500327696413250175);
assert_eq!(report_hash(&full), 2748844479819636032);
}
#[test]
+133
View File
@@ -0,0 +1,133 @@
//! Parity guard for the gRPC admin-auth gate.
//!
//! Every `VolumeServer` handler must either call `check_grpc_admin_auth` or be
//! listed here as intentionally ungated, with the reason it stays open. Mirrors
//! the Go side's `TestVolumeServerAdminAuthCoverage`
//! (`weed/server/volume_grpc_admin_auth_coverage_test.go`). The gate had
//! silently dropped off 15 handlers on the Rust side; this keeps the two
//! implementations from drifting apart again.
use std::collections::{HashMap, HashSet};
/// Handlers that intentionally run without `check_grpc_admin_auth`, each with
/// the reason. Kept in sync with `ungatedVolumeServerRPCs` on the Go side.
///
/// The split is by caller, not by how destructive the call is: the guard checks
/// the peer IP against `-whiteList`, which holds masters, shell hosts and
/// workers -- not every peer volume server. Gating a volume-server ->
/// volume-server call therefore breaks replication, EC and tiering, so those
/// calls stay open and need a cluster-peer identity before they can be closed.
fn ungated_handlers() -> HashMap<&'static str, &'static str> {
[
// Cluster-internal: issued volume server -> volume server.
("copy_file", "replica sync and EC task pull whole files from a peer"),
("read_needle_blob", "replica sync, vacuum and EC rebuild read needles from a peer"),
("read_needle_meta", "replica sync compares needle metadata across peers"),
("write_needle_blob", "replica sync repairs a peer's needle"),
("receive_file", "EC shard distribution pushes shards to a peer"),
("read_volume_file_status", "the copy path queries the source volume server"),
("volume_ec_shard_read", "a volume server reads EC shards held by a peer"),
("volume_ec_blob_delete", "EC delete is fanned out to the shard holders"),
("volume_ec_shards_info", "EC verification polls shard holders"),
("volume_ec_shards_mount", "EC shard distribution mounts on the receiving peer"),
("volume_incremental_copy", "volume backup pulls increments from a peer"),
("volume_sync_status", "sync compares volume state across peers"),
("volume_tail_sender", "the tail source streams to the receiving peer"),
("volume_status", "replica sync and the master's vacuum loop poll volume status"),
// Read-only or liveness: no state change.
("ping", "liveness probe"),
("get_state", "read-only volume server state"),
("query", "read-only data query"),
("vacuum_volume_check", "read-only garbage ratio; the vacuum steps that act on it are gated"),
("volume_server_status", "read-only status, the gRPC counterpart of the /status page"),
]
.into_iter()
.collect()
}
/// Whether a handler body actually invokes the gate. We can't run a real AST
/// pass like the Go side without pulling in a parser, so approximate it: drop
/// `//` line comments (the realistic way a stray mention would sneak in) and
/// look for the call form `check_grpc_admin_auth(`, not the bare identifier, so
/// a comment or doc reference can't make an ungated handler pass.
fn calls_gate(body: &str) -> bool {
body.lines()
.map(|line| line.split_once("//").map_or(line, |(code, _)| code))
.any(|code| code.contains("check_grpc_admin_auth("))
}
/// Enumerate the `VolumeServer` trait handlers and whether each gates on
/// `check_grpc_admin_auth`, by scanning the source of the trait impl block.
fn handler_gating(src: &str) -> Vec<(String, bool)> {
let impl_start = src
.find("impl VolumeServer for VolumeGrpcService {")
.expect("locate the VolumeServer trait impl");
let after = &src[impl_start..];
// rustfmt keeps every method body indented, so the impl's own closing brace
// is the first line that begins at column 0.
let impl_end = after
.match_indices("\n}")
.next()
.map(|(i, _)| i + 1)
.expect("locate the end of the trait impl");
let block = &after[..impl_end];
let marker = "\n async fn ";
let starts: Vec<usize> = block.match_indices(marker).map(|(i, _)| i).collect();
let mut handlers = Vec::with_capacity(starts.len());
for (i, &start) in starts.iter().enumerate() {
let name_start = start + marker.len();
let name_end = name_start
+ block[name_start..]
.find('(')
.expect("handler signature has an argument list");
let name = block[name_start..name_end].to_string();
let body_end = starts.get(i + 1).copied().unwrap_or(block.len());
let gated = calls_gate(&block[start..body_end]);
handlers.push((name, gated));
}
handlers
}
#[test]
fn volume_server_admin_auth_coverage() {
let src = std::fs::read_to_string(concat!(
env!("CARGO_MANIFEST_DIR"),
"/src/server/grpc_server.rs"
))
.expect("read grpc_server.rs");
let handlers = handler_gating(&src);
assert!(
handlers.len() >= 40,
"parsed only {} handlers from the trait impl, expected the full service",
handlers.len()
);
let ungated = ungated_handlers();
let mut problems = Vec::new();
for (name, gated) in &handlers {
match (gated, ungated.contains_key(name.as_str())) {
(true, true) => problems.push(format!(
"{name} calls check_grpc_admin_auth but is also listed as intentionally ungated; drop it from ungated_handlers"
)),
(false, false) => problems.push(format!(
"{name} does not call check_grpc_admin_auth and is not listed as intentionally ungated; gate it, or add it with the reason it must stay open"
)),
_ => {}
}
}
// Keep the exemption list honest: an entry for a handler that no longer
// exists hides the fact that nothing is being exempted.
let names: HashSet<&str> = handlers.iter().map(|(n, _)| n.as_str()).collect();
for name in ungated.keys() {
if !names.contains(name) {
problems.push(format!(
"ungated_handlers lists \"{name}\", which is not a VolumeServer handler"
));
}
}
assert!(problems.is_empty(), "admin-auth coverage gaps:\n{}", problems.join("\n"));
}
+114 -2
View File
@@ -385,6 +385,46 @@ async fn write_then_read_needle() {
assert_eq!(body, payload, "GET body should match written data");
}
// A durable upload takes the same route through the handler; the flush is not
// observable from here, but a broken wiring would show up as a failed write.
#[tokio::test]
async fn write_with_fsync_then_read_needle() {
let (state, _tmp) = test_state();
let uri = "/1,01637037d6?fsync=true";
let payload = b"durable through the handler";
let app = build_admin_router(state.clone());
let response = app
.oneshot(
Request::builder()
.method("POST")
.uri(uri)
.body(Body::from(payload.to_vec()))
.unwrap(),
)
.await
.unwrap();
assert_eq!(
response.status(),
StatusCode::CREATED,
"a durable POST should return 201 Created"
);
let app = build_admin_router(state.clone());
let response = app
.oneshot(
Request::builder()
.uri("/1,01637037d6")
.body(Body::empty())
.unwrap(),
)
.await
.unwrap();
assert_eq!(response.status(), StatusCode::OK);
assert_eq!(body_bytes(response).await, payload);
}
// ============================================================================
// 5. DELETE deletes a needle, subsequent GET returns 404
// ============================================================================
@@ -739,6 +779,76 @@ async fn replicate_write_raw_body_is_stored() {
assert_eq!(body_bytes(response).await, payload);
}
/// Go reads fsync through r.FormValue, which decodes the query, so a
/// percent-encoded value has to reach the write path here too.
#[tokio::test]
async fn write_with_percent_encoded_fsync_is_accepted() {
let (state, _tmp) = test_state();
let payload = b"encoded durable payload";
let app = build_admin_router(state.clone());
let response = app
.oneshot(
Request::builder()
.method("POST")
.uri("/1,01637037d6?fsync=%74rue")
.body(Body::from(payload.to_vec()))
.unwrap(),
)
.await
.unwrap();
assert_eq!(response.status(), StatusCode::CREATED);
let app = build_admin_router(state.clone());
let response = app
.oneshot(
Request::builder()
.uri("/1,01637037d6")
.body(Body::empty())
.unwrap(),
)
.await
.unwrap();
assert_eq!(response.status(), StatusCode::OK);
assert_eq!(body_bytes(response).await, payload);
}
/// The fan-out query a Go primary sends for a durable write: `fsync=true` rides
/// along with `type=replicate`, and the replica has to honor it rather than ack
/// out of the page cache.
#[tokio::test]
async fn replicate_write_with_fsync_is_stored() {
let (state, _tmp) = test_state();
let uri = "/1,01637037d6?fsync=true&type=replicate";
let payload = b"durable replica bytes";
let app = build_admin_router(state.clone());
let response = app
.oneshot(
Request::builder()
.method("POST")
.uri(uri)
.body(Body::from(payload.to_vec()))
.unwrap(),
)
.await
.unwrap();
assert_eq!(response.status(), StatusCode::CREATED);
let app = build_admin_router(state.clone());
let response = app
.oneshot(
Request::builder()
.uri("/1,01637037d6")
.body(Body::empty())
.unwrap(),
)
.await
.unwrap();
assert_eq!(response.status(), StatusCode::OK);
assert_eq!(body_bytes(response).await, payload);
}
/// Multipart `type=replicate` write (the shape the Go gateway uploader sends)
/// is stored and reads back.
#[tokio::test]
@@ -882,7 +992,7 @@ async fn chunk_manifest_expands_chunk_stored_on_ec_volume() {
data_size: chunk_data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.write_needle(&mut n, true, false).unwrap();
v.sync_to_disk().unwrap();
v.close();
}
@@ -916,7 +1026,9 @@ async fn chunk_manifest_expands_chunk_stored_on_ec_volume() {
};
n.data_size = n.data.len() as u32;
n.set_is_chunk_manifest();
store.write_volume_needle(VolumeId(1), &mut n).unwrap();
store
.write_volume_needle(VolumeId(1), &mut n, false)
.unwrap();
}
// GET the manifest object; expect the reconstructed chunk bytes.
+1
View File
@@ -0,0 +1 @@
target/
+6812
View File
File diff suppressed because it is too large Load Diff
+27
View File
@@ -0,0 +1,27 @@
# Rust plugin workers for SeaweedFS.
#
# `core` is the plugin.proto contract and nothing else; a worker crate beside it
# supplies job handlers and a binary. Adding a worker means adding a member here,
# not touching the protocol.
[workspace]
resolver = "2"
members = ["crates/core", "crates/lance"]
[workspace.package]
version = "0.1.0"
edition = "2021"
[workspace.dependencies]
anyhow = "1"
async-trait = "0.1"
prost = "0.13"
prost-types = "0.13"
tokio = { version = "1", features = ["full"] }
tokio-stream = "0.1"
tonic = { version = "0.12", features = ["tls"] }
# Already in the tree via tonic; named here so the metrics server can use them.
axum = "0.7"
prometheus = { version = "0.13", default-features = false }
tonic-build = "0.12"
tracing = "0.1"
tracing-subscriber = { version = "0.3", features = ["env-filter"] }
+77
View File
@@ -0,0 +1,77 @@
# SeaweedFS Rust workers
`weed/pb/plugin.proto` is a language-agnostic contract: a maintenance worker
connects out to admin, announces the job types it can detect and execute, and
answers requests on that one stream. `weed worker -admin=host:23646` is the Go
implementation of it from outside the admin process. This workspace is the Rust
one.
crates/core the contract: stream, handshake, heartbeat, registry, config forms
crates/lance maintenance jobs for Lance tables, and a binary
`core` knows nothing about any job. A second worker is a new crate beside
`lance` that depends on it, not a fork of the protocol.
## Running
cargo run -p weed-lance-worker -- --admin 127.0.0.1:23646
The admin's *HTTP* address is what an operator has; the gRPC port is derived
from it the way the Go side does. Dialling the HTTP port fails as "frame with
invalid size", which reads like a protocol bug rather than a wrong port.
## Metrics
cargo run -p weed-lance-worker -- --admin 127.0.0.1:23646 --metrics-port 9328
Serves `/health`, `/ready` and `/metrics` on that port, the same three the Go
worker serves under `weed worker -metricsPort`, so one scrape config covers
workers in either language. Off by default, and bound to loopback unless
`--metrics-ip` says otherwise, because the endpoint is unauthenticated. 9328
continues the series the other components use (master 9324, volume 9325, filer
9326, s3 9327); an IPv6 address works with or without brackets.
Grafana: the "Plugin Workers" row of `other/metrics/grafana_seaweedfs.json`
graphs these. Its panels filter on `$cluster`, which comes from the scrape job's
labels, so scrape the worker the way the rest of the cluster is scraped or the
row stays empty.
Names are `SeaweedFS_worker_*`, matching the Go side's convention. The pair
worth alerting on is `objects_seen_total` and `objects_skipped_total`: a sweep
that proposes nothing and a sweep that could read nothing look identical from
`proposals_total` alone.
SeaweedFS_worker_connected 1
SeaweedFS_worker_objects_seen_total{job_type="lance_compact"} 7
SeaweedFS_worker_proposals_total{job_type="lance_compact"} 2
SeaweedFS_worker_jobs_total{job_type="lance_compact",result="ok"} 2
SeaweedFS_worker_lance_fragments_removed_total 25
`/ready` follows the control stream: a worker whose admin has gone away is
running but is not going to do anything.
## Credentials
The worker holds none. It asks the namespace to describe a table with
`vend_credentials` and hands the `storage_options` that come back to lance. A
gateway without STS configured vends no credentials at all, so `--access-key`
and `--secret-key` supply a fallback; anything the namespace does vend wins over
them.
## State
All three jobs are implemented and tested end to end against a live gateway:
compaction result: 12 fragments became 1
reindex result: 512 uncovered rows became 0
cleanup result: removed 14 versions and 24272 bytes
`cargo test -p weed-lance-worker` runs them when `WEED_LANCE_NAMESPACE` names a
live namespace and skips otherwise, the way the Go integration tests skip
without Docker. Each test seeds the table it needs, including building a vector
index and then appending rows outside it, so a run does not depend on what the
previous one left behind — the first version of these did, and quietly stopped
testing anything once it had done its job.
The handshake, descriptor exchange and heartbeat work against a live admin,
which logs the worker connecting and prefetches all three descriptors.
+23
View File
@@ -0,0 +1,23 @@
[package]
name = "seaweed-worker-core"
version.workspace = true
edition.workspace = true
description = "SeaweedFS plugin.proto worker contract"
[lib]
name = "seaweed_worker_core"
[dependencies]
anyhow.workspace = true
async-trait.workspace = true
prost.workspace = true
prost-types.workspace = true
tokio.workspace = true
tokio-stream.workspace = true
tonic.workspace = true
axum.workspace = true
prometheus.workspace = true
tracing.workspace = true
[build-dependencies]
tonic-build.workspace = true
+12
View File
@@ -0,0 +1,12 @@
fn main() -> Result<(), Box<dyn std::error::Error>> {
// Compiled straight out of the Go tree, the way seaweed-volume already reads
// filer.proto, so the contract cannot drift from a vendored copy.
tonic_build::configure()
// The server half is only for tests, which stand up a fake admin.
.build_server(true)
.build_client(true)
.protoc_arg("--experimental_allow_proto3_optional")
.compile_protos(&["../../../weed/pb/plugin.proto"], &["../../../weed/pb/"])?;
println!("cargo:rerun-if-changed=../../../weed/pb/plugin.proto");
Ok(())
}
+73
View File
@@ -0,0 +1,73 @@
//! SeaweedFS addresses the way the Go tree does.
//!
//! An operator gives a worker the admin's HTTP address, and the gRPC port is
//! derived from it rather than asked for separately. Dialling the HTTP port
//! instead fails as "frame with invalid size", which reads like a protocol bug
//! rather than a wrong port, so getting this right is worth its own module.
//! Mirrors pb.ServerToGrpcAddress in weed/pb/grpc_client_server.go.
const GRPC_PORT_OFFSET: u16 = 10000;
/// Converts `host:port` to the gRPC address, and accepts the explicit
/// `host:port.grpcPort` form the Go side also understands.
pub fn server_to_grpc_address(server: &str) -> Option<String> {
let (host, port_part) = server.rsplit_once(':')?;
// "port.grpcPort" states the gRPC port outright.
if let Some((_, grpc_port)) = port_part.split_once('.') {
if let Ok(port) = grpc_port.parse::<u16>() {
return Some(join_host_port(host, port));
}
}
let port: u16 = port_part.parse().ok()?;
Some(join_host_port(host, port.checked_add(GRPC_PORT_OFFSET)?))
}
fn join_host_port(host: &str, port: u16) -> String {
// An IPv6 literal has to keep its brackets or the port reads as part of it.
if host.contains(':') && !host.starts_with('[') {
format!("[{host}]:{port}")
} else {
format!("{host}:{port}")
}
}
#[cfg(test)]
mod tests {
use super::server_to_grpc_address;
#[test]
fn derives_the_grpc_port() {
assert_eq!(
server_to_grpc_address("localhost:23646").as_deref(),
Some("localhost:33646")
);
assert_eq!(
server_to_grpc_address("127.0.0.1:9333").as_deref(),
Some("127.0.0.1:19333")
);
}
#[test]
fn honours_an_explicit_grpc_port() {
assert_eq!(
server_to_grpc_address("localhost:23646.33999").as_deref(),
Some("localhost:33999")
);
}
#[test]
fn brackets_ipv6_literals() {
assert_eq!(
server_to_grpc_address("::1:23646").as_deref(),
Some("[::1]:33646")
);
}
#[test]
fn rejects_what_it_cannot_parse() {
assert!(server_to_grpc_address("localhost").is_none());
assert!(server_to_grpc_address("localhost:notaport").is_none());
}
}
+49
View File
@@ -0,0 +1,49 @@
use std::time::Duration;
/// How one worker process connects and how much work it will take on.
#[derive(Clone, Debug)]
pub struct WorkerOptions {
/// Admin gRPC address, e.g. "localhost:23646".
pub admin_address: String,
pub worker_id: String,
pub worker_version: String,
/// Advertised address; empty when the worker takes no inbound connections.
pub worker_address: String,
pub heartbeat_interval: Duration,
pub reconnect_delay: Duration,
pub max_detection_concurrency: i32,
pub max_execution_concurrency: i32,
/// mTLS for the control stream, mirroring the Go worker's `[grpc.worker]`
/// section of security.toml. None means plaintext, which is the default the
/// Go worker also takes when no certificates are configured.
pub tls: Option<TlsOptions>,
}
/// Certificates for the control stream. All three are required together: the
/// cluster's gRPC TLS is mutual, so a CA without a client identity gets refused
/// by admin rather than falling back to one-way TLS.
#[derive(Clone, Debug)]
pub struct TlsOptions {
pub ca_path: String,
pub client_cert_path: String,
pub client_key_path: String,
/// Name to verify the server certificate against, when the address a worker
/// dials is not the name the certificate carries.
pub server_name: Option<String>,
}
impl Default for WorkerOptions {
fn default() -> Self {
Self {
admin_address: "localhost:23646".to_string(),
worker_id: String::new(),
worker_version: env!("CARGO_PKG_VERSION").to_string(),
worker_address: String::new(),
heartbeat_interval: Duration::from_secs(10),
reconnect_delay: Duration::from_secs(5),
max_detection_concurrency: 1,
max_execution_concurrency: 1,
tls: None,
}
}
}
@@ -0,0 +1,95 @@
//! Builders for the config forms a worker returns in its JobTypeDescriptor.
//!
//! Admin renders these into the job's settings page, so a worker written in any
//! language gets a UI without touching Go or templ. That only holds if the
//! field types and defaults are right, which is why they are built here rather
//! than spelled out at each call site.
use std::collections::HashMap;
use crate::pb::{
config_value::Kind, ConfigField, ConfigFieldType, ConfigForm, ConfigSection, ConfigValue,
};
pub fn int_value(value: i64) -> ConfigValue {
ConfigValue {
kind: Some(Kind::Int64Value(value)),
}
}
pub fn bool_value(value: bool) -> ConfigValue {
ConfigValue {
kind: Some(Kind::BoolValue(value)),
}
}
pub fn string_value(value: impl Into<String>) -> ConfigValue {
ConfigValue {
kind: Some(Kind::StringValue(value.into())),
}
}
/// Reads an integer a request carried, falling back when admin sent nothing.
pub fn int_or(values: &HashMap<String, ConfigValue>, name: &str, fallback: i64) -> i64 {
match values.get(name).and_then(|v| v.kind.as_ref()) {
Some(Kind::Int64Value(value)) => *value,
Some(Kind::DoubleValue(value)) => *value as i64,
_ => fallback,
}
}
pub fn bool_or(values: &HashMap<String, ConfigValue>, name: &str, fallback: bool) -> bool {
match values.get(name).and_then(|v| v.kind.as_ref()) {
Some(Kind::BoolValue(value)) => *value,
_ => fallback,
}
}
pub fn string_or(values: &HashMap<String, ConfigValue>, name: &str, fallback: &str) -> String {
match values.get(name).and_then(|v| v.kind.as_ref()) {
Some(Kind::StringValue(value)) if !value.is_empty() => value.clone(),
_ => fallback.to_string(),
}
}
pub fn number_field(name: &str, label: &str, description: &str, min: i64, max: i64) -> ConfigField {
ConfigField {
name: name.to_string(),
label: label.to_string(),
description: description.to_string(),
field_type: ConfigFieldType::Int64 as i32,
min_value: Some(int_value(min)),
max_value: Some(int_value(max)),
..Default::default()
}
}
pub fn bool_field(name: &str, label: &str, description: &str) -> ConfigField {
ConfigField {
name: name.to_string(),
label: label.to_string(),
description: description.to_string(),
field_type: ConfigFieldType::Bool as i32,
..Default::default()
}
}
pub fn form(
form_id: &str,
title: &str,
fields: Vec<ConfigField>,
defaults: HashMap<String, ConfigValue>,
) -> ConfigForm {
ConfigForm {
form_id: form_id.to_string(),
title: title.to_string(),
description: String::new(),
sections: vec![ConfigSection {
section_id: format!("{form_id}-main"),
title: title.to_string(),
description: String::new(),
fields,
}],
default_values: defaults,
}
}
+26
View File
@@ -0,0 +1,26 @@
//! The SeaweedFS plugin worker contract, in Rust.
//!
//! `weed/pb/plugin.proto` is a language-agnostic gRPC stream: a worker connects
//! out to admin, announces the job types it can detect and execute, and then
//! answers requests on that one stream. `weed worker -admin=host:23646` is the
//! Go implementation of the same contract from outside the admin process; this
//! is the Rust one, and it knows nothing about any particular job.
pub mod address;
pub mod config;
pub mod config_form;
pub mod metrics;
pub mod registry;
pub mod senders;
pub mod stream;
/// Generated plugin.proto types.
pub mod pb {
tonic::include_proto!("plugin");
}
pub use config::{TlsOptions, WorkerOptions};
pub use metrics::Metrics;
pub use registry::{JobHandler, Preview, PreviewProvider, Registry};
pub use senders::{DetectionSender, ExecutionSender};
pub use stream::run;
+395
View File
@@ -0,0 +1,395 @@
//! What a worker will say about itself when nobody is watching the logs.
//!
//! The Go plugin worker serves `/health`, `/ready` and `/metrics` on an optional
//! port (`weed worker -metricsPort`); this is the same contract, so one scrape
//! config covers workers in either language. Names follow the Go side's
//! convention, `SeaweedFS_<subsystem>_<name>`, under a `worker` subsystem that
//! nothing else uses.
//!
//! Nothing here is Lance-specific. A worker for another format registers its own
//! collectors on the same registry and gets the same endpoint.
use std::net::SocketAddr;
use std::sync::Arc;
use anyhow::{Context, Result};
use prometheus::{
Encoder, HistogramOpts, HistogramVec, IntCounter, IntCounterVec, IntGauge, IntGaugeVec, Opts,
Registry, TextEncoder,
};
use std::sync::atomic::{AtomicBool, Ordering};
use tracing::{info, warn};
const NAMESPACE: &str = "SeaweedFS";
const SUBSYSTEM: &str = "worker";
/// The metrics one worker process publishes, and the registry they live on.
///
/// Cloneable because the stream loop, the job handlers and the HTTP server all
/// hold it; everything inside is already shared.
#[derive(Clone)]
pub struct Metrics {
registry: Registry,
connected: IntGauge,
connects: IntCounterVec,
slots_used: IntGaugeVec,
slots_total: IntGaugeVec,
detections: IntCounterVec,
detection_seconds: HistogramVec,
proposals: IntCounterVec,
objects_seen: IntCounterVec,
objects_skipped: IntCounterVec,
jobs: IntCounterVec,
job_seconds: HistogramVec,
previews: IntCounterVec,
ready: Arc<AtomicBool>,
}
impl Metrics {
pub fn new(worker_id: &str, worker_version: &str) -> Result<Self> {
let registry = Registry::new();
// Build info as a constant 1, the way the Go side does it, so a scrape
// can tell which worker and which build answered.
let build = IntGaugeVec::new(
Opts::new("build_info", "Worker build information.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["worker_id", "version"],
)?;
build.with_label_values(&[worker_id, worker_version]).set(1);
registry.register(Box::new(build))?;
let connected = IntGauge::with_opts(
Opts::new("connected", "1 while the admin control stream is up.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
)?;
let connects = IntCounterVec::new(
Opts::new(
"stream_events_total",
"Control stream lifecycle events by outcome (connected, closed, failed, shutdown).",
)
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["event"],
)?;
let slots_used = IntGaugeVec::new(
Opts::new("slots_used", "Slots currently held, by lane.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["lane"],
)?;
let slots_total = IntGaugeVec::new(
Opts::new("slots_total", "Slots this worker advertises, by lane.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["lane"],
)?;
let detections = IntCounterVec::new(
Opts::new(
"detections_total",
"Detection sweeps by job type and outcome.",
)
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["job_type", "result"],
)?;
let detection_seconds = HistogramVec::new(
HistogramOpts::new("detection_seconds", "How long a detection sweep took.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM)
.buckets(prometheus::exponential_buckets(0.01, 2.0, 14)?),
&["job_type"],
)?;
let proposals = IntCounterVec::new(
Opts::new("proposals_total", "Jobs proposed to admin, by job type.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["job_type"],
)?;
let objects_seen = IntCounterVec::new(
Opts::new(
"objects_seen_total",
"Objects a detection sweep read, by job type.",
)
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["job_type"],
)?;
let objects_skipped = IntCounterVec::new(
Opts::new(
"objects_skipped_total",
"Objects a sweep could not read, by job type and reason. A sweep that \
proposes nothing looks the same as one that could read nothing; this is \
what tells them apart.",
)
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["job_type", "reason"],
)?;
let jobs = IntCounterVec::new(
Opts::new("jobs_total", "Jobs executed by job type and outcome.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["job_type", "result"],
)?;
let job_seconds = HistogramVec::new(
HistogramOpts::new("job_seconds", "How long a job took to run.")
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM)
.buckets(prometheus::exponential_buckets(0.05, 2.0, 16)?),
&["job_type"],
)?;
let previews = IntCounterVec::new(
Opts::new(
"previews_total",
"Object previews served to admin, by outcome.",
)
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
&["result"],
)?;
for collector in [
Box::new(connected.clone()) as Box<dyn prometheus::core::Collector>,
Box::new(connects.clone()),
Box::new(slots_used.clone()),
Box::new(slots_total.clone()),
Box::new(detections.clone()),
Box::new(detection_seconds.clone()),
Box::new(proposals.clone()),
Box::new(objects_seen.clone()),
Box::new(objects_skipped.clone()),
Box::new(jobs.clone()),
Box::new(job_seconds.clone()),
Box::new(previews.clone()),
] {
registry.register(collector)?;
}
Ok(Self {
registry,
connected,
connects,
slots_used,
slots_total,
detections,
detection_seconds,
proposals,
objects_seen,
objects_skipped,
jobs,
job_seconds,
previews,
ready: Arc::new(AtomicBool::new(false)),
})
}
/// The registry, so a worker can add collectors of its own.
pub fn registry(&self) -> &Registry {
&self.registry
}
pub fn stream_connected(&self) {
self.connected.set(1);
self.ready.store(true, Ordering::Relaxed);
self.connects.with_label_values(&["connected"]).inc();
}
/// `event` is why the stream ended: closed, failed, or shutdown.
pub fn stream_ended(&self, event: &str) {
self.connected.set(0);
self.ready.store(false, Ordering::Relaxed);
self.connects.with_label_values(&[event]).inc();
}
pub fn set_slots(&self, lane: &str, used: i64, total: i64) {
self.slots_used.with_label_values(&[lane]).set(used);
self.slots_total.with_label_values(&[lane]).set(total);
}
pub fn detection_finished(&self, job_type: &str, result: &str, seconds: f64, proposals: usize) {
self.detections
.with_label_values(&[job_type, result])
.inc_by(1);
self.detection_seconds
.with_label_values(&[job_type])
.observe(seconds);
self.proposals
.with_label_values(&[job_type])
.inc_by(proposals as u64);
}
pub fn object_seen(&self, job_type: &str) {
self.objects_seen.with_label_values(&[job_type]).inc();
}
pub fn object_skipped(&self, job_type: &str, reason: &str) {
self.objects_skipped
.with_label_values(&[job_type, reason])
.inc();
}
pub fn job_finished(&self, job_type: &str, result: &str, seconds: f64) {
self.jobs.with_label_values(&[job_type, result]).inc();
self.job_seconds
.with_label_values(&[job_type])
.observe(seconds);
}
pub fn preview_finished(&self, result: &str) {
self.previews.with_label_values(&[result]).inc();
}
/// A counter this worker's own jobs can raise, e.g. fragments removed.
/// Registered lazily so a format's numbers live beside the generic ones
/// without core having to know what they are.
pub fn counter(&self, name: &str, help: &str) -> Result<IntCounter> {
let counter = IntCounter::with_opts(
Opts::new(name, help)
.namespace(NAMESPACE)
.subsystem(SUBSYSTEM),
)?;
self.registry.register(Box::new(counter.clone()))?;
Ok(counter)
}
fn gather(&self) -> Result<String> {
let mut buffer = Vec::new();
TextEncoder::new().encode(&self.registry.gather(), &mut buffer)?;
Ok(String::from_utf8(buffer)?)
}
fn is_ready(&self) -> bool {
self.ready.load(Ordering::Relaxed)
}
}
/// Serves the metrics endpoints until the process stops. Failing to bind is
/// logged rather than fatal: a worker that cannot publish metrics should still
/// do its work.
pub async fn serve(metrics: Metrics, addr: SocketAddr) -> Result<()> {
use axum::extract::State;
use axum::http::StatusCode;
use axum::routing::get;
use axum::Router;
let app = Router::new()
.route("/health", get(|| async { StatusCode::OK }))
.route(
"/ready",
get(|State(metrics): State<Metrics>| async move {
if metrics.is_ready() {
StatusCode::OK
} else {
StatusCode::SERVICE_UNAVAILABLE
}
}),
)
.route(
"/metrics",
get(|State(metrics): State<Metrics>| async move {
match metrics.gather() {
Ok(body) => (StatusCode::OK, body),
Err(err) => (StatusCode::INTERNAL_SERVER_ERROR, format!("{err:#}")),
}
}),
)
.with_state(metrics);
let listener = tokio::net::TcpListener::bind(addr)
.await
.with_context(|| format!("bind the metrics server to {addr}"))?;
info!("serving worker metrics on http://{addr}/metrics");
axum::serve(listener, app)
.await
.context("metrics server stopped")
}
/// Starts the metrics server in the background, warning rather than failing when
/// it cannot start.
pub fn spawn(metrics: Metrics, addr: SocketAddr) {
tokio::spawn(async move {
if let Err(err) = serve(metrics, addr).await {
warn!("worker metrics server: {err:#}");
}
});
}
#[cfg(test)]
mod tests {
use super::*;
fn metrics() -> Metrics {
Metrics::new("worker-1", "0.1.0").expect("build metrics")
}
// A scrape has to be able to tell these apart: a worker that swept and found
// nothing to do, and one that could not read anything it swept.
#[test]
fn a_sweep_reports_what_it_read_and_what_it_could_not() {
let metrics = metrics();
metrics.object_seen("lance_compact");
metrics.object_seen("lance_compact");
metrics.object_skipped("lance_compact", "open");
metrics.detection_finished("lance_compact", "ok", 0.5, 0);
let text = metrics.gather().expect("gather");
assert!(text.contains("SeaweedFS_worker_objects_seen_total{job_type=\"lance_compact\"} 2"));
assert!(text.contains(
"SeaweedFS_worker_objects_skipped_total{job_type=\"lance_compact\",reason=\"open\"} 1"
));
assert!(text.contains(
"SeaweedFS_worker_detections_total{job_type=\"lance_compact\",result=\"ok\"} 1"
));
assert!(text.contains("SeaweedFS_worker_proposals_total{job_type=\"lance_compact\"} 0"));
}
// /ready follows the stream, because a worker with no admin behind it is
// running but not doing anything.
#[test]
fn readiness_follows_the_stream() {
let metrics = metrics();
assert!(!metrics.is_ready(), "not ready before the stream is up");
metrics.stream_connected();
assert!(metrics.is_ready());
assert!(metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_connected 1"));
metrics.stream_ended("closed");
assert!(!metrics.is_ready());
assert!(metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_connected 0"));
assert!(metrics
.gather()
.unwrap()
.contains("SeaweedFS_worker_stream_events_total{event=\"closed\"} 1"));
}
#[test]
fn build_info_names_the_worker() {
let text = metrics().gather().expect("gather");
assert!(text
.contains("SeaweedFS_worker_build_info{version=\"0.1.0\",worker_id=\"worker-1\"} 1"));
}
// A format's own numbers land on the same registry, so one endpoint serves
// both and a second worker implementation needs no new plumbing.
#[test]
fn a_worker_can_add_counters_of_its_own() {
let metrics = metrics();
let counter = metrics
.counter("lance_fragments_removed_total", "Fragments merged away.")
.expect("register");
counter.inc_by(16);
let text = metrics.gather().expect("gather");
assert!(text.contains("SeaweedFS_worker_lance_fragments_removed_total 16"));
}
}
@@ -0,0 +1,89 @@
use std::collections::HashMap;
use std::sync::Arc;
use anyhow::Result;
use async_trait::async_trait;
use crate::pb::{ExecuteJobRequest, JobTypeCapability, JobTypeDescriptor, RunDetectionRequest};
use crate::senders::{DetectionSender, ExecutionSender};
/// One job type, worker side. Mirrors the Go JobHandler interface in
/// weed/plugin/worker/worker.go so the two stay readable against each other.
#[async_trait]
pub trait JobHandler: Send + Sync {
fn capability(&self) -> JobTypeCapability;
/// The descriptor admin renders as this job's settings page.
fn descriptor(&self) -> JobTypeDescriptor;
async fn detect(
&self,
request: &RunDetectionRequest,
sender: &dyn DetectionSender,
) -> Result<()>;
async fn execute(
&self,
request: &ExecuteJobRequest,
sender: &dyn ExecutionSender,
) -> Result<()>;
}
/// Sample rows of one object, already rendered as text. The worker formats
/// them because it is the only side that knows the object's types.
pub struct Preview {
pub columns: Vec<String>,
pub rows: Vec<Vec<String>>,
/// Rows in the object, which is not the number sampled.
pub total_rows: i64,
}
/// Reads sample rows of a format admin cannot parse itself.
///
/// This is deliberately not a JobHandler: a preview is answered while someone
/// waits on a page, so it neither schedules nor reports progress.
#[async_trait]
pub trait PreviewProvider: Send + Sync {
/// The format this provider reads, matched case-insensitively against what
/// the catalog recorded.
fn format(&self) -> &str;
async fn preview(&self, object_id: &[String], row_limit: usize) -> Result<Preview>;
}
/// The handlers one worker process serves. A process may serve several job
/// types, which is why WorkerHello carries a list of capabilities.
#[derive(Default, Clone)]
pub struct Registry {
handlers: HashMap<String, Arc<dyn JobHandler>>,
previews: HashMap<String, Arc<dyn PreviewProvider>>,
}
impl Registry {
pub fn new() -> Self {
Self::default()
}
pub fn register(mut self, handler: Arc<dyn JobHandler>) -> Self {
self.handlers.insert(handler.capability().job_type, handler);
self
}
pub fn with_preview(mut self, provider: Arc<dyn PreviewProvider>) -> Self {
self.previews
.insert(provider.format().to_ascii_uppercase(), provider);
self
}
pub fn get(&self, job_type: &str) -> Option<Arc<dyn JobHandler>> {
self.handlers.get(job_type).cloned()
}
pub fn preview_provider(&self, format: &str) -> Option<Arc<dyn PreviewProvider>> {
self.previews.get(&format.to_ascii_uppercase()).cloned()
}
pub fn capabilities(&self) -> Vec<JobTypeCapability> {
self.handlers.values().map(|h| h.capability()).collect()
}
pub fn is_empty(&self) -> bool {
self.handlers.is_empty()
}
}
+203
View File
@@ -0,0 +1,203 @@
use anyhow::Result;
use tokio::sync::mpsc;
use crate::pb::{
worker_to_admin_message::Body, ActivityEvent, DetectionComplete, DetectionProposals,
JobCompleted, JobProgressUpdate, WorkerObservations, WorkerToAdminMessage,
};
/// Replies to one detection request.
pub trait DetectionSender: Send + Sync {
fn send_proposals(&self, proposals: DetectionProposals) -> Result<()>;
fn send_complete(&self, complete: DetectionComplete) -> Result<()>;
fn send_activity(&self, activity: ActivityEvent) -> Result<()>;
/// Reports what the worker saw while deciding. Admin caches the last one
/// per object and serves it back for display; nothing is scheduled from it.
fn send_observations(&self, observations: WorkerObservations) -> Result<()>;
}
/// Replies to one execution request.
pub trait ExecutionSender: Send + Sync {
fn send_progress(&self, progress: JobProgressUpdate) -> Result<()>;
fn send_completed(&self, completed: JobCompleted) -> Result<()>;
}
/// Both senders write to the single outbound stream, so they share one channel.
#[derive(Clone)]
pub struct StreamSender {
worker_id: String,
tx: mpsc::UnboundedSender<WorkerToAdminMessage>,
}
impl StreamSender {
pub fn new(worker_id: String, tx: mpsc::UnboundedSender<WorkerToAdminMessage>) -> Self {
Self { worker_id, tx }
}
pub fn send(&self, body: Body) -> Result<()> {
self.tx.send(WorkerToAdminMessage {
worker_id: self.worker_id.clone(),
sent_at: Some(std::time::SystemTime::now().into()),
body: Some(body),
})?;
Ok(())
}
}
impl DetectionSender for StreamSender {
fn send_proposals(&self, proposals: DetectionProposals) -> Result<()> {
self.send(Body::DetectionProposals(proposals))
}
fn send_complete(&self, complete: DetectionComplete) -> Result<()> {
self.send(Body::DetectionComplete(complete))
}
fn send_activity(&self, _activity: ActivityEvent) -> Result<()> {
// Activity rides inside progress and completion messages rather than
// being a body of its own, so there is nothing to send on its own here.
Ok(())
}
fn send_observations(&self, observations: WorkerObservations) -> Result<()> {
self.send(Body::Observations(observations))
}
}
impl ExecutionSender for StreamSender {
fn send_progress(&self, progress: JobProgressUpdate) -> Result<()> {
self.send(Body::JobProgressUpdate(progress))
}
fn send_completed(&self, completed: JobCompleted) -> Result<()> {
self.send(Body::JobCompleted(completed))
}
}
/// Wraps a sender so what passes through it is counted. The handlers report
/// their results to admin and nowhere else, so this is where a scrape can learn
/// what happened without every handler having to know about metrics.
pub struct MeteredSender<'a> {
inner: &'a StreamSender,
proposals: std::sync::atomic::AtomicUsize,
failed: std::sync::atomic::AtomicBool,
}
impl<'a> MeteredSender<'a> {
pub fn new(inner: &'a StreamSender) -> Self {
Self {
inner,
proposals: std::sync::atomic::AtomicUsize::new(0),
failed: std::sync::atomic::AtomicBool::new(false),
}
}
/// How many proposals went out, for the detection counter.
pub fn proposals(&self) -> usize {
self.proposals.load(std::sync::atomic::Ordering::Relaxed)
}
/// Whether the handler reported a failure of its own. A handler that fails
/// by returning an error is counted by the caller; this catches the one that
/// reports failure and returns Ok.
pub fn reported_failure(&self) -> bool {
self.failed.load(std::sync::atomic::Ordering::Relaxed)
}
}
impl DetectionSender for MeteredSender<'_> {
fn send_proposals(&self, proposals: DetectionProposals) -> Result<()> {
// Counted after the send, not before: a stream that closed mid-sweep
// would otherwise leave proposals_total claiming work admin never saw.
let count = proposals.proposals.len();
self.inner.send_proposals(proposals)?;
self.proposals
.fetch_add(count, std::sync::atomic::Ordering::Relaxed);
Ok(())
}
fn send_complete(&self, complete: DetectionComplete) -> Result<()> {
if !complete.success {
self.failed
.store(true, std::sync::atomic::Ordering::Relaxed);
}
self.inner.send_complete(complete)
}
fn send_activity(&self, activity: ActivityEvent) -> Result<()> {
self.inner.send_activity(activity)
}
fn send_observations(&self, observations: WorkerObservations) -> Result<()> {
self.inner.send_observations(observations)
}
}
impl ExecutionSender for MeteredSender<'_> {
fn send_progress(&self, progress: JobProgressUpdate) -> Result<()> {
self.inner.send_progress(progress)
}
fn send_completed(&self, completed: JobCompleted) -> Result<()> {
if !completed.success {
self.failed
.store(true, std::sync::atomic::Ordering::Relaxed);
}
self.inner.send_completed(completed)
}
}
#[cfg(test)]
mod tests {
use super::*;
use crate::pb::JobProposal;
fn proposals(n: usize) -> DetectionProposals {
DetectionProposals {
proposals: (0..n).map(|_| JobProposal::default()).collect(),
..Default::default()
}
}
#[test]
fn proposals_are_counted_once_they_are_sent() {
let (tx, mut rx) = mpsc::unbounded_channel();
let stream = StreamSender::new("worker-1".to_string(), tx);
let metered = MeteredSender::new(&stream);
metered.send_proposals(proposals(3)).expect("send");
assert_eq!(metered.proposals(), 3);
assert!(rx.try_recv().is_ok(), "the proposals reached the stream");
}
// Admin never saw these, so counting them would report work that was not
// handed over.
#[test]
fn proposals_are_not_counted_when_the_stream_is_gone() {
let (tx, rx) = mpsc::unbounded_channel();
let stream = StreamSender::new("worker-1".to_string(), tx);
let metered = MeteredSender::new(&stream);
drop(rx);
assert!(metered.send_proposals(proposals(3)).is_err());
assert_eq!(metered.proposals(), 0);
}
// A handler can report failure and still return Ok; the outcome has to come
// from what it said, not only from what it returned.
#[test]
fn a_reported_failure_is_remembered() {
let (tx, _rx) = mpsc::unbounded_channel();
let stream = StreamSender::new("worker-1".to_string(), tx);
let metered = MeteredSender::new(&stream);
assert!(!metered.reported_failure());
metered
.send_complete(DetectionComplete {
success: false,
..Default::default()
})
.expect("send");
assert!(metered.reported_failure());
}
}
+491
View File
@@ -0,0 +1,491 @@
use std::sync::Arc;
use std::time::{Duration, Instant};
use anyhow::{anyhow, Context, Result};
use tokio::sync::{mpsc, Semaphore};
use tokio_stream::wrappers::UnboundedReceiverStream;
use tonic::transport::{Certificate, Channel, ClientTlsConfig, Identity};
use tracing::{info, warn};
use crate::config::WorkerOptions;
use crate::metrics::Metrics;
use crate::pb::{
admin_to_worker_message::Body as AdminBody,
plugin_control_service_client::PluginControlServiceClient,
worker_to_admin_message::Body as WorkerBody, ConfigSchemaResponse, ExecuteJobRequest,
JobCompleted, ObjectPreviewResponse, PreviewRow, RequestObjectPreview, RunDetectionRequest,
RunningWork, WorkerHeartbeat, WorkerHello,
};
use crate::registry::Registry;
use crate::senders::{MeteredSender, StreamSender};
/// The protocol version this worker speaks, sent in WorkerHello.
const PROTOCOL_VERSION: &str = "1";
/// Connect to admin and serve the registry until the context is cancelled,
/// reconnecting on failure. The stream is the only channel: everything admin
/// asks for and everything the worker reports flows through it.
pub async fn run(options: WorkerOptions, registry: Registry) -> Result<()> {
let metrics = Metrics::new(&options.worker_id, &options.worker_version)?;
run_with_metrics(options, registry, metrics).await
}
/// Runs with metrics a caller has already made, so it can serve them and add
/// collectors of its own before the stream starts.
pub async fn run_with_metrics(
options: WorkerOptions,
registry: Registry,
metrics: Metrics,
) -> Result<()> {
if registry.is_empty() {
return Err(anyhow!("no job handlers registered"));
}
if options.max_detection_concurrency < 1 || options.max_execution_concurrency < 1 {
return Err(anyhow!(
"concurrency limits must be at least 1, got detection={} execution={}",
options.max_detection_concurrency,
options.max_execution_concurrency
));
}
let slots = Slots::new(&options);
metrics.set_slots("detection", 0, slots.detection_total as i64);
metrics.set_slots("execution", 0, slots.execution_total as i64);
loop {
match serve_once(&options, &registry, &slots, &metrics).await {
Err(err) => {
metrics.stream_ended("failed");
warn!("worker stream ended: {err:#}")
}
// Admin asked this worker to stop, so stop. Reconnecting here would
// make shutdown impossible: the worker would log back in.
Ok(Outcome::ShutdownRequested) => {
metrics.stream_ended("shutdown");
return Ok(());
}
// Admin closing a healthy stream is not an error, but reconnecting
// in silence hides the reason - two workers sharing an id evict
// each other and produce nothing but a login every few seconds.
Ok(Outcome::StreamClosed) => {
metrics.stream_ended("closed");
warn!(
"admin closed the stream; reconnecting in {:?}. If this repeats, check for \
another worker using the id {}",
options.reconnect_delay, options.worker_id
)
}
}
tokio::time::sleep(options.reconnect_delay).await;
}
}
/// Why a stream ended. Only one of these means "do not come back".
enum Outcome {
StreamClosed,
ShutdownRequested,
}
/// The capacity this worker advertises in WorkerHello. Admin schedules against
/// those numbers, so the worker has to actually hold to them - and the heartbeat
/// has to report what is in use, or admin is scheduling blind.
#[derive(Clone)]
struct Slots {
detection: Arc<Semaphore>,
execution: Arc<Semaphore>,
detection_total: i32,
execution_total: i32,
}
impl Slots {
fn new(options: &WorkerOptions) -> Self {
Self {
detection: Arc::new(Semaphore::new(options.max_detection_concurrency as usize)),
execution: Arc::new(Semaphore::new(options.max_execution_concurrency as usize)),
detection_total: options.max_detection_concurrency,
execution_total: options.max_execution_concurrency,
}
}
fn detection_used(&self) -> i32 {
self.detection_total - self.detection.available_permits() as i32
}
fn execution_used(&self) -> i32 {
self.execution_total - self.execution.available_permits() as i32
}
}
/// Dials admin, over mTLS when certificates are configured. Plaintext is the
/// default and is fine over loopback; anything else carries preview rows and
/// execution commands in the clear, and a cluster with grpc TLS on refuses the
/// connection anyway.
async fn connect(options: &WorkerOptions, grpc_address: &str) -> Result<Channel> {
let Some(tls) = options.tls.as_ref() else {
return Ok(Channel::from_shared(format!("http://{grpc_address}"))?
.connect_timeout(Duration::from_secs(10))
.connect()
.await?);
};
let ca = tokio::fs::read(&tls.ca_path)
.await
.with_context(|| format!("read CA certificate {}", tls.ca_path))?;
let cert = tokio::fs::read(&tls.client_cert_path)
.await
.with_context(|| format!("read client certificate {}", tls.client_cert_path))?;
let key = tokio::fs::read(&tls.client_key_path)
.await
.with_context(|| format!("read client key {}", tls.client_key_path))?;
let mut config = ClientTlsConfig::new()
.ca_certificate(Certificate::from_pem(ca))
.identity(Identity::from_pem(cert, key));
if let Some(server_name) = tls.server_name.as_ref() {
config = config.domain_name(server_name.clone());
}
Ok(Channel::from_shared(format!("https://{grpc_address}"))?
.tls_config(config)?
.connect_timeout(Duration::from_secs(10))
.connect()
.await?)
}
async fn serve_once(
options: &WorkerOptions,
registry: &Registry,
slots: &Slots,
metrics: &Metrics,
) -> Result<Outcome> {
// Operators give the admin's HTTP address; the gRPC port is derived, the
// same way the Go worker does it.
let grpc_address = crate::address::server_to_grpc_address(&options.admin_address)
.ok_or_else(|| anyhow!("cannot parse admin address {}", options.admin_address))?;
let channel = connect(options, &grpc_address).await?;
let mut client = PluginControlServiceClient::new(channel);
let (tx, rx) = mpsc::unbounded_channel();
let sender = StreamSender::new(options.worker_id.clone(), tx);
sender.send(WorkerBody::Hello(WorkerHello {
worker_id: options.worker_id.clone(),
worker_instance_id: options.worker_id.clone(),
address: options.worker_address.clone(),
worker_version: options.worker_version.clone(),
protocol_version: PROTOCOL_VERSION.to_string(),
capabilities: registry.capabilities(),
metadata: Default::default(),
}))?;
let mut inbound = client
.worker_stream(UnboundedReceiverStream::new(rx))
.await?
.into_inner();
let heartbeat = spawn_heartbeat(
sender.clone(),
options.clone(),
slots.clone(),
metrics.clone(),
);
while let Some(message) = inbound.message().await? {
let request_id = message.request_id.clone();
match message.body {
Some(AdminBody::Hello(hello)) => {
if !hello.accepted {
return Err(anyhow!("admin rejected this worker: {}", hello.message));
}
metrics.stream_connected();
info!(
"connected to admin at {} ({})",
options.admin_address, grpc_address
);
}
Some(AdminBody::RequestConfigSchema(request)) => {
let response = match registry.get(&request.job_type) {
Some(handler) => ConfigSchemaResponse {
request_id: request_id.clone(),
job_type: request.job_type.clone(),
success: true,
error_message: String::new(),
job_type_descriptor: Some(handler.descriptor()),
},
None => ConfigSchemaResponse {
request_id: request_id.clone(),
job_type: request.job_type.clone(),
success: false,
error_message: format!("unknown job type {}", request.job_type),
job_type_descriptor: None,
},
};
sender.send(WorkerBody::ConfigSchemaResponse(response))?;
}
Some(AdminBody::RequestObjectPreview(request)) => {
spawn_preview(
registry.clone(),
sender.clone(),
metrics.clone(),
request_id.clone(),
request,
);
}
Some(AdminBody::RunDetectionRequest(request)) => {
spawn_detection(
registry.clone(),
sender.clone(),
slots.clone(),
metrics.clone(),
request,
);
}
Some(AdminBody::ExecuteJobRequest(request)) => {
spawn_execution(
registry.clone(),
sender.clone(),
slots.clone(),
metrics.clone(),
request,
);
}
Some(AdminBody::CancelRequest(request)) => {
// Cancellation needs a per-request handle to be honoured; until
// then say so rather than silently continuing to run the job.
warn!(
"cancel requested for {} ({}) but is not implemented",
request.target_id, request.reason
);
}
Some(AdminBody::Shutdown(shutdown)) => {
info!("admin asked this worker to stop: {}", shutdown.reason);
heartbeat.abort();
return Ok(Outcome::ShutdownRequested);
}
None => {}
}
}
heartbeat.abort();
Ok(Outcome::StreamClosed)
}
/// Answers one preview request off the stream loop. Reading rows takes as long
/// as it takes, and the stream has heartbeats to keep up meanwhile.
fn spawn_preview(
registry: Registry,
sender: StreamSender,
metrics: Metrics,
request_id: String,
request: RequestObjectPreview,
) {
tokio::spawn(async move {
let limit = request.row_limit.max(1) as usize;
let response = match registry.preview_provider(&request.format) {
None => ObjectPreviewResponse {
request_id,
success: false,
error_message: format!("this worker does not read {} objects", request.format),
..Default::default()
},
Some(provider) => match provider.preview(&request.object_id, limit).await {
Ok(preview) => ObjectPreviewResponse {
request_id,
success: true,
error_message: String::new(),
columns: preview.columns,
rows: preview
.rows
.into_iter()
.map(|values| PreviewRow { values })
.collect(),
total_rows: preview.total_rows,
},
Err(err) => ObjectPreviewResponse {
request_id,
success: false,
error_message: format!("{err:#}"),
..Default::default()
},
},
};
metrics.preview_finished(if response.success { "ok" } else { "failed" });
let _ = sender.send(WorkerBody::ObjectPreviewResponse(response));
});
}
fn spawn_heartbeat(
sender: StreamSender,
options: WorkerOptions,
slots: Slots,
metrics: Metrics,
) -> tokio::task::JoinHandle<()> {
// The handle has to be the heartbeat's own, or aborting it aborts nothing
// and every reconnect leaves another ticker running.
tokio::spawn(async move {
let mut ticker = tokio::time::interval(options.heartbeat_interval);
loop {
ticker.tick().await;
// The heartbeat already computes this for admin; publish the same
// numbers so a scrape and the admin UI cannot disagree.
metrics.set_slots(
"detection",
slots.detection_used() as i64,
slots.detection_total as i64,
);
metrics.set_slots(
"execution",
slots.execution_used() as i64,
slots.execution_total as i64,
);
let beat = WorkerHeartbeat {
worker_id: options.worker_id.clone(),
running_work: Vec::<RunningWork>::new(),
detection_slots_used: slots.detection_used(),
detection_slots_total: slots.detection_total,
execution_slots_used: slots.execution_used(),
execution_slots_total: slots.execution_total,
queued_jobs_by_type: Default::default(),
metadata: Default::default(),
};
if sender.send(WorkerBody::Heartbeat(beat)).is_err() {
return;
}
}
})
}
fn spawn_detection(
registry: Registry,
sender: StreamSender,
slots: Slots,
metrics: Metrics,
request: RunDetectionRequest,
) {
tokio::spawn(async move {
let Some(handler) = registry.get(&request.job_type) else {
return;
};
// Held until the sweep finishes, so the worker keeps to the capacity it
// advertised and the heartbeat reports the truth while it works.
let _permit = slots.detection.acquire().await;
let metered = MeteredSender::new(&sender);
let started = Instant::now();
let outcome = handler.detect(&request, &metered).await;
let result = if let Err(err) = &outcome {
warn!("detection for {} failed: {err:#}", request.job_type);
let _ = sender.send(WorkerBody::DetectionComplete(
crate::pb::DetectionComplete {
request_id: request.request_id.clone(),
job_type: request.job_type.clone(),
success: false,
error_message: format!("{err:#}"),
total_proposals: 0,
},
));
"failed"
} else if metered.reported_failure() {
"failed"
} else {
"ok"
};
metrics.detection_finished(
&request.job_type,
result,
started.elapsed().as_secs_f64(),
metered.proposals(),
);
});
}
fn spawn_execution(
registry: Registry,
sender: StreamSender,
slots: Slots,
metrics: Metrics,
request: ExecuteJobRequest,
) {
tokio::spawn(async move {
let _permit = slots.execution.acquire().await;
let job_type = request
.job
.as_ref()
.map(|job| job.job_type.clone())
.unwrap_or_default();
let job_id = request
.job
.as_ref()
.map(|job| job.job_id.clone())
.unwrap_or_default();
let Some(handler) = registry.get(&job_type) else {
return;
};
let metered = MeteredSender::new(&sender);
let started = Instant::now();
let outcome = handler.execute(&request, &metered).await;
let result = if let Err(err) = &outcome {
warn!("job {job_id} failed: {err:#}");
let _ = sender.send(WorkerBody::JobCompleted(JobCompleted {
request_id: request.request_id.clone(),
job_id,
job_type: job_type.clone(),
success: false,
error_message: format!("{err:#}"),
..Default::default()
}));
"failed"
} else if metered.reported_failure() {
"failed"
} else {
"ok"
};
metrics.job_finished(&job_type, result, started.elapsed().as_secs_f64());
});
}
#[cfg(test)]
mod tests {
use super::*;
fn options(detection: i32, execution: i32) -> WorkerOptions {
WorkerOptions {
max_detection_concurrency: detection,
max_execution_concurrency: execution,
..Default::default()
}
}
// The heartbeat is admin's only view of how busy this worker is; before the
// permits existed it reported zero however much was running.
#[tokio::test]
async fn slots_report_what_is_held() {
let slots = Slots::new(&options(2, 3));
assert_eq!(slots.detection_used(), 0);
assert_eq!(slots.execution_used(), 0);
let held = slots.detection.acquire().await.unwrap();
assert_eq!(slots.detection_used(), 1);
assert_eq!(slots.execution_used(), 0, "the two do not share capacity");
drop(held);
assert_eq!(slots.detection_used(), 0);
}
// A limit of one means the second request waits, rather than running anyway
// as it did when every request simply spawned a task.
#[tokio::test]
async fn a_full_lane_makes_the_next_request_wait() {
let slots = Slots::new(&options(1, 1));
let held = slots.execution.clone().acquire_owned().await.unwrap();
assert_eq!(slots.execution_used(), 1);
let waiter = tokio::spawn({
let execution = slots.execution.clone();
async move { execution.acquire_owned().await.unwrap() }
});
tokio::task::yield_now().await;
assert!(!waiter.is_finished(), "the second request must not start");
drop(held);
let _second = waiter.await.expect("the waiter should be handed the slot");
assert_eq!(slots.execution_used(), 1);
}
}
+40
View File
@@ -0,0 +1,40 @@
[package]
name = "weed-lance-worker"
version.workspace = true
edition.workspace = true
description = "SeaweedFS maintenance worker for Lance tables"
[lib]
name = "weed_lance_worker"
[[bin]]
name = "weed-lance-worker"
path = "src/main.rs"
[dependencies]
seaweed-worker-core = { path = "../core" }
prometheus.workspace = true
# Only the S3 backend: the other object stores lance enables by default are
# build time this worker never spends.
lance = { version = "10", default-features = false, features = ["aws"] }
lance-index = "10"
arrow-schema = "58"
arrow-cast = "58"
chrono = "0.4"
futures = "0.3"
anyhow.workspace = true
async-trait.workspace = true
clap = { version = "4", features = ["derive", "env"] }
reqwest = { version = "0.12", default-features = false, features = ["json", "rustls-tls"] }
serde = { version = "1", features = ["derive"] }
serde_json = "1"
tokio.workspace = true
tracing.workspace = true
tracing-subscriber.workspace = true
[dev-dependencies]
tokio = { workspace = true }
arrow-array = "58"
arrow-schema = "58"
arrow-cast = "58"
lance-linalg = "10"
@@ -0,0 +1,10 @@
//! How the worker finds Lance tables and gets at their bytes.
//!
//! It goes through the Lance namespace rather than the filer: the namespace is
//! the catalog of record, it already knows which tables are Lance, and asking it
//! to describe a table with vend_credentials is how the worker gets storage
//! credentials without holding any of its own.
pub mod namespace;
pub use namespace::{parse_id, NamespaceClient, TableDescription};
@@ -0,0 +1,129 @@
use std::collections::HashMap;
use std::time::Duration;
use anyhow::{Context, Result};
use serde::Deserialize;
/// The delimiter the Lance namespace joins identifier parts with.
const DELIMITER: &str = "$";
#[derive(Debug, Deserialize)]
pub struct TableDescription {
pub location: String,
#[serde(default)]
pub storage_options: HashMap<String, String>,
#[serde(default)]
pub managed_versioning: bool,
}
#[derive(Debug, Deserialize)]
struct ListNamespacesResponse {
#[serde(default)]
namespaces: Vec<String>,
}
#[derive(Debug, Deserialize)]
struct ListTablesResponse {
#[serde(default)]
tables: Vec<String>,
}
/// A thin client for the operations this worker needs. It deliberately does not
/// wrap the whole spec: a maintenance worker lists, describes, and commits.
pub struct NamespaceClient {
base_url: String,
http: reqwest::Client,
}
/// A namespace call that has not answered by now is not going to. Without this
/// a gateway that accepts the connection and then goes quiet holds a detection
/// slot open forever, and the sweep never finishes.
const REQUEST_TIMEOUT: Duration = Duration::from_secs(30);
const CONNECT_TIMEOUT: Duration = Duration::from_secs(10);
impl NamespaceClient {
pub fn new(base_url: impl Into<String>) -> Self {
let http = reqwest::Client::builder()
.timeout(REQUEST_TIMEOUT)
.connect_timeout(CONNECT_TIMEOUT)
.build()
// The builder only fails on a bad TLS backend, which would break
// every call anyway; a client with no timeouts is worse than a panic
// at startup, so keep the default only as a last resort.
.unwrap_or_else(|err| {
tracing::warn!("falling back to an untimed HTTP client: {err}");
reqwest::Client::new()
});
Self {
base_url: base_url.into().trim_end_matches('/').to_string(),
http,
}
}
/// Every table the namespace holds, as delimiter-joined identifiers.
pub async fn list_all_tables(&self) -> Result<Vec<String>> {
let url = format!("{}/v1/table", self.base_url);
let response: ListTablesResponse = self
.http
.get(&url)
.send()
.await
.context("list tables")?
.error_for_status()?
.json()
.await?;
Ok(response.tables)
}
/// Child namespaces of `id`; the root lists table buckets.
pub async fn list_namespaces(&self, id: &[String]) -> Result<Vec<String>> {
let url = format!("{}/v1/namespace/{}/list", self.base_url, encode_id(id));
let response: ListNamespacesResponse = self
.http
.get(&url)
.send()
.await
.context("list namespaces")?
.error_for_status()?
.json()
.await?;
Ok(response.namespaces)
}
/// Resolve a table to a location and the credentials to reach it. The
/// credentials expire, so a long compaction re-describes rather than
/// carrying one set for the whole job.
pub async fn describe_table(&self, id: &[String]) -> Result<TableDescription> {
let url = format!("{}/v1/table/{}/describe", self.base_url, encode_id(id));
let body = serde_json::json!({ "id": id, "vend_credentials": true });
let description: TableDescription = self
.http
.post(&url)
.json(&body)
.send()
.await
.context("describe table")?
.error_for_status()?
.json()
.await?;
Ok(description)
}
}
fn encode_id(id: &[String]) -> String {
if id.is_empty() {
DELIMITER.to_string()
} else {
id.join(DELIMITER)
}
}
/// Splits a delimiter-joined identifier back into parts.
pub fn parse_id(encoded: &str) -> Vec<String> {
encoded
.split(DELIMITER)
.filter(|part| !part.is_empty())
.map(|part| part.to_string())
.collect()
}
@@ -0,0 +1,98 @@
//! Opening a Lance dataset with credentials the namespace vended.
//!
//! The worker holds no storage credentials of its own: it asks the namespace to
//! describe a table with `vend_credentials`, and the `storage_options` that come
//! back are handed to lance as-is. They expire, so a job that runs longer than
//! their lifetime re-describes rather than carrying one set throughout.
use std::collections::HashMap;
use anyhow::{Context, Result};
use lance::dataset::builder::DatasetBuilder;
use lance::dataset::Dataset;
use crate::catalog::{NamespaceClient, TableDescription};
/// A table the worker is about to work on.
pub struct OpenTable {
pub id: Vec<String>,
pub location: String,
pub dataset: Dataset,
}
/// What detection needs to decide whether a table is worth a job. Reading it
/// opens the dataset but touches no data files.
pub struct TableStats {
pub fragments: usize,
pub version: u64,
pub total_versions: usize,
pub rows: usize,
/// The Arrow schema as JSON, which is the only description of this table
/// anything outside the format can produce.
pub schema: Option<String>,
}
/// Storage options an operator supplies for deployments that vend none.
pub type FallbackOptions = HashMap<String, String>;
/// Resolve a table through the namespace and open it.
pub async fn open(
client: &NamespaceClient,
id: &[String],
fallback: &FallbackOptions,
) -> Result<OpenTable> {
let description = client.describe_table(id).await?;
let dataset = open_at(&description, fallback).await?;
Ok(OpenTable {
id: id.to_vec(),
location: description.location,
dataset,
})
}
/// The namespace vends object_store's own option names, so they pass straight
/// through. What it vends always wins: the fallback exists because a deployment
/// without STS vends no credentials at all, and then the worker has no other way
/// to reach the data.
async fn open_at(description: &TableDescription, fallback: &FallbackOptions) -> Result<Dataset> {
let mut options: HashMap<String, String> = fallback.clone();
options.extend(description.storage_options.clone());
DatasetBuilder::from_uri(&description.location)
.with_storage_options(options)
.load()
.await
.with_context(|| format!("open lance dataset at {}", description.location))
}
impl OpenTable {
pub async fn stats(&self) -> Result<TableStats> {
let versions = self.dataset.versions().await?;
Ok(TableStats {
fragments: self.dataset.get_fragments().len(),
version: self.dataset.version().version,
total_versions: versions.len(),
rows: self.dataset.count_rows(None).await.unwrap_or(0),
schema: schema_json(&self.dataset),
})
}
}
/// Renders the dataset's schema as JSON. Best effort: a schema that will not
/// serialise is not a reason to fail a maintenance sweep.
fn schema_json(dataset: &Dataset) -> Option<String> {
let arrow: arrow_schema::Schema = dataset.schema().into();
serde_json::to_string(
&arrow
.fields()
.iter()
.map(|f| {
serde_json::json!({
"name": f.name(),
"type": f.data_type().to_string(),
"nullable": f.is_nullable(),
})
})
.collect::<Vec<_>>(),
)
.ok()
}

Some files were not shown because too many files have changed in this diff Show More