Files
seaweedfs/design-lance-catalog.md
Chris LuandGitHub 8c7d714d5e Lance catalog, and a Rust plugin worker to maintain it (#10841)
* iceberg: skip tables the maintenance worker does not own

A Lance dataset registered through the Lance namespace's Iceberg REST adapter
arrives as an Iceberg table with a placeholder schema and table_type=lance, and
keeps its fragments under data/ - the same subdirectory the orphan cleaner
walks. Every fragment is unreferenced by the Iceberg metadata, so a maintenance
pass deletes the dataset. Views share the entry shape and were only skipped
because parsing their metadata happened to fail first.

Gate the scan and the execution path on the entry actually being an Iceberg
table. Maintenance is off by default, so this was latent rather than live.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: let a table declare a format the catalog does not interpret

CreateTable accepted ICEBERG and nothing else. A Lance table has no metadata
file for the catalog to maintain - the entry records a name and the dataset
root, and the client owns everything under it - so accept LANCE, and carry the
declared format on the entry instead of hardcoding it back on the way out.

ListTables now reports format and metadataLocation, so listing a catalog that
holds both kinds takes one pass rather than a GetTable per row. AWS omits both
fields; adding them is additive.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: move the in-memory filer into its own package

The Lance namespace tests need the same harness, and copying it would leave two
of them to keep in step. Extracted as it was, plus the two fidelity gaps that
only surface once a paginating caller uses it: ListEntries ignored
startFromFileName and limit, so a caller that paginates re-read the first page
until it hit its own cap and reported the same entry over and over, and
GetFilerConfiguration was missing, which CreateTableBucket needs to resolve the
buckets directory.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: serve the Lance Namespace REST spec

A second catalog surface beside the Iceberg one, over the same table buckets:
the namespace and table metadata operations, the $-delimited identifier codec,
the spec's numeric error model, the directory-catalog marker files, and
storage_options vending through the STS path the Iceberg catalog already uses.
Listens on -port.lance, 9101 by default, and inherits ARNs, policies and tags
from the storage layer, so a Lance table needs no second permission model.

Identifiers map bucket / namespace / table onto the three levels Lance clients
already use, which is why there is no warehouse selector to invent. The data
plane needs Lance format support that does not exist in Go and answers with the
spec's Unsupported code rather than a bare 404.

Two things it deliberately will not do: create a table bucket as a side effect
of creating a namespace inside one, since a bucket carries its own policy and
lifecycle, and resolve an Iceberg table's location for a Lance client, which
would hand it a table another engine owns.

The design note this follows is in design-lance-catalog.md, including the
.lance directory suffix it proposed and this does not implement.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* mini: give the Lance port the same treatment as the Iceberg one

The flag was registered but nothing else knew about it, so mini would start the
server without reserving its port, waiting for it, or saying where it is. Adds
it to the startup service list, the conflict resolver, the gRPC allocator's
reserved set, the readiness wait, the stop reporting and the banner.

The admin server still takes only the Iceberg port, because there is no Lance
page for it to link to.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: stop deregister and repoint from deleting the dataset

Deregistering preserves data by definition, and this did the opposite: the
catalog entry is the dataset directory, so DeleteTable took the files with it.
Registering over an existing name had the same shape, destroying the dataset
the name used to hold. Found by driving the running server rather than the
in-memory filer, where both looked like success because the table did stop
being listed.

Deregistering is now a state on the entry - the marker file hides it, and
declaring or registering the name again brings it back. Repointing a name at
another dataset is an UpdateTable against the version token, so neither dataset
loses files. Drop is left alone; it is the operation that does remove data.

The storage endpoint now falls back to the advertised -ip where the Iceberg
derivation gives up. An Iceberg client brings its own s3.endpoint and
advertising the wrong one hijacks it, but storage_options is the only place a
Lance client learns where the store is, and without it object_store quietly
talks to real AWS.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: refuse to create a table over one of another format

Creating a table that already exists is idempotent, and that path returned the
existing table without looking at its format. A Lance declare over an Iceberg
table answered 200 and handed back a directory Iceberg owns, so the client
would write its dataset on top. The view check immediately above it already
guards the same class of collision.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: let a table bucket hold a format other than Iceberg

The S3 door validated every object written into a table bucket against
Iceberg's file layout, so a Lance client could not write its dataset at all: it
got 403 on data/*.lance, on _versions/, and on the _transactions/ directory it
turned out to write as well. Table buckets were only neutral containers by
intention; in practice they were Iceberg-shaped and enforced as such.

The allowed set is now the union of what the supported formats write, because
the validator runs where the table's format is not in hand. Underscore-prefixed
directories are treated as belonging to the format, since enumerating them
means guessing at the next one - _transactions is exactly the one this missed -
and their contents are checked only for traversal. Iceberg writes none of them,
so it loses nothing. Marker files at the table root are admitted too, which the
namespace/table/dir/file shape had rejected as too shallow.

Describe also honours the request-body spellings of with_table_uri,
load_detailed_metadata and check_declared. The spec puts them in the query
string, but real clients send both.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record what the implementation found

The table bucket being an Iceberg-shaped container, enforced at the S3 door,
was the premise this design never questioned and the one that had to change
before anything worked end to end.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* iceberg: prove the data loss the foreign-format guard prevents

The guard landed with a unit test for the predicate and nothing showing what it
saves. These seed what the Lance namespace's Iceberg REST adapter actually
leaves behind - an Iceberg table with a placeholder schema and table_type=lance
whose directory holds a Lance dataset - and assert both halves: orphan
collection does flag the dataset's fragments, because the Iceberg metadata
beside them references nothing, and the scan never reaches the table. An
ordinary Iceberg table in the same shape is still scanned, so the guard is not
just skipping everything.

Confirmed against a running gateway first: our Iceberg catalog accepts the
adapter's registration, and a real Lance client then writes a dataset into that
table's location.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tablestest: make the in-memory filer safe to race against

Two gaps that only matter once a test drives concurrent writers, which is what
an exclusive create has to be tested with: the entry map had no lock, and
CreateEntry ignored O_EXCL entirely, so both writers of the same name would
have won and the test would have passed while proving nothing.

The BeforeUpdate hook runs before the lock is taken. Its whole purpose is to
land a competing write in a handler's read-to-write window, and that write
needs the lock the hook would otherwise be holding.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: make the namespace an external manifest store

Lance commits a version by writing _versions/{v}.manifest with
put-if-not-exists. The S3 layer in front of this same filer evaluates
If-None-Match by looking the entry up and then writing without a precondition,
so two writers can both pass the check and one commit is lost. The filer itself
has the primitive: CreateEntry with o_excl.

Adds the four version operations a Lance client actually calls - create, list,
describe and batch-delete - recording one entry per version under
_lance_versions/, and advertises managed_versioning so the client routes its
commits here. Reserving a version is the exclusive create, so exactly one of
several racing writers wins and the rest rebase.

Off by default, behind -lance.managedVersioning. Turning it on moves where a
table's version history lives, and a reader that does not come through this
namespace no longer sees all of it; that is the operator's call, not a default.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record what managed versioning does and does not reach

The first commit through a namespace-backed store works and is recorded the way
the protocol specifies. Later commits do not, because lance 4.0.0 refuses
put_if_exists on that path in its own code, so the feature is capped upstream
rather than here.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: integration tests for the Lance namespace

Everything this surface got wrong so far - a deregister that deleted the
dataset, an S3 door that refused every Lance file, a version reservation that
could not actually be exclusive - passed against an in-memory filer first. So
these run against a live gateway, and where the claim is about data they check
storage rather than visibility.

Five Go tests on the shared harness: namespace and table lifecycle including
that deregister keeps the bytes and drop removes them, that a Lance client
cannot resolve or declare over an Iceberg table, that a Lance dataset's files
get past the table-bucket layout guard while junk still does not, and that
eight writers racing for one version produce exactly one winner.

One Docker-gated test drives the real Lance client, which is the only way to
check that the location and storage_options the namespace vends are between
them enough to write and read a dataset. It overrides the endpoint with the
container's view of the same gateway, because the shared harness binds a
wildcard address and so vends none.

The harness gains a Lance port and turns managed versioning on; the flag
touches nothing outside that surface.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3tables: a directory with no namespace metadata is a missing namespace

Three callers resolved a namespace by reading its metadata attribute and each
tested only for a missing entry, so a directory that carried no metadata came
back as an internal error saying "attribute not found". Creating a table under
a namespace that does not exist answered 500.

Collapses the three copies into one helper that reports both conditions as
absent, which is what they are: a directory without namespace metadata is not a
namespace.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* iceberg: stop reporting storage-layer refusals as server faults

writeManagerError recognised a missing table bucket and sent everything else to
500, so a missing namespace, a duplicate name and a commit conflict all reached
the client as InternalServerError with nothing to act on. Creating a table in a
namespace that does not exist is the case that turned up: 500 where the spec
wants 404 NoSuchNamespaceException.

Maps the storage error types onto the exception names this package already
uses, and keeps the existing bucket message, which explains how to select a
table bucket.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* iceberg: skip a foreign-format table by name, not by failing to parse it

A table the namespace created as LANCE carries no Iceberg metadata, so the
worker skipped it only because the parse failed, and logged that as damaged
metadata. The catalog records the format on the entry and this never read it.

Reading it turns an accident into a decision, and separates a mixed catalog
from a corrupt one in the logs. The property check beside it still covers the
other shape: a real Iceberg table wearing table_type=lance, which is what the
Lance namespace's Iceberg REST adapter writes.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: answer whether a Lance table needs maintenance

It does, and index optimization has no Iceberg equivalent: rows written after
an index was built are not covered by it, so a vector search quietly misses
them. None of the three jobs can run in the Go worker, and there is no useful
subset, because deciding what an old version still references means parsing
Lance manifests. Version cleanup at least has an answer that needs nothing from
us - Lance can enable it on the dataset itself.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: the Lance maintenance worker is a plugin worker, in Rust

Framing it as a sidecar was wrong. plugin.proto already defines a
language-agnostic gRPC contract for external maintenance workers, and
"weed worker -admin=..." is the Go reference implementation of it from outside
the admin process. seaweed-volume already compiles protos out of weed/pb with
tonic_build, so a Lance worker is that build plus plugin.proto and the lance
crate.

Scheduling, retries, dedupe, progress and the admin settings page all come from
the protocol: a worker that answers RequestConfigSchema with a descriptor gets
its configuration form rendered without a line of Go.

The data plane is the part that genuinely does need a process answering HTTP,
and this had the two conflated.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* seaweed-worker: Rust plugin worker workspace, with Lance as the first one

plugin.proto is language-agnostic and the Rust toolchain was already in the
tree, so a Lance maintenance worker needs no new integration surface: core is
the contract and nothing else, and a worker crate beside it supplies handlers
and a binary. A second worker is a new member here rather than a fork of the
protocol, which is why this is seaweed-worker and not seaweed-lance-worker.

Verified against a running admin: it connects, is accepted, and admin prefetches
descriptors for lance_compact, lance_optimize_indices and lance_cleanup_versions,
so their settings pages render from the Rust side without a line of Go. The
stream stays up across heartbeats.

The job bodies are stubs that report failure. Doing the work means adding the
lance crate and opening the dataset, and claiming success before that would be
worse than saying so.

Two things running it caught that reading the proto did not: the admin address
has to be converted to the gRPC port the way pb.ServerToGrpcAddress does, or the
dial fails as an h2 frame error; and the generated field names differ from the
Go ones in several places, so JobCompleted carries success rather than a state
enum.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: implement compaction

Detection lists tables from the namespace, opens each one, and proposes a job
for any with more fragments than the policy allows; opening a dataset reads its
manifest and not its data, so a sweep stays cheap. Execution re-resolves the
table rather than trusting what detection saw - it may have been repointed, and
the vended credentials expire - then compacts and reports the fragment counts
either side.

Verified against a live gateway: a twelve-fragment dataset became one fragment
with all twelve rows intact. The test drives the handler directly and skips
unless WEED_LANCE_NAMESPACE names a namespace, the way the Go integration tests
skip without Docker.

Running it turned up a gap the design had not: a gateway without STS vends no
credentials at all, so the worker could not open anything and detection quietly
proposed nothing. --access-key/--secret-key are the fallback, and whatever the
namespace vends still wins over them.

Two API assumptions did not survive contact either. Datasets open through
DatasetBuilder::with_storage_options, not ReadParams, and lance 10's
ObjectStoreParams has no storage_options field at all.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: implement index optimization and version cleanup

Index optimization is the job with no Iceberg equivalent: rows appended after
an index was built are invisible to a search of it until this runs. Detection
reads num_unindexed_rows from each index's statistics and proposes a table once
more rows sit outside its indices than the budget allows; a table with no
indices is skipped, which is different from one whose indices have fallen
behind.

Cleanup applies a retention window, refusing rather than silently dropping a
tagged version, and leaving unverified files alone because they may belong to a
commit still in flight.

Both verified against a live gateway: 512 uncovered rows became 0, and a
fourteen-version table lost its old ones. Each test now seeds what it needs,
including building an IVF_PQ index and appending rows outside it. The first
version of these depended on state a script had left, so the second run found
the work already done and asserted nothing - a test that passes by doing
nothing is worse than no test.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: answer an empty catalog with an empty list, not null

ListAllTables built its result from a nil slice, so a namespace holding no
tables answered {"tables":null} on a field the spec marks required. A generated
client may decode that differently from an empty list. Found running the
namespace on a dev box, where the catalog was empty.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: give Lance maintenance its own scheduler lane

Lane assignment is a hardcoded map, so the three lance_* job types fell through
to the default lane. That lane serialises its work under the cluster admin lock
because volume management shares global state, which would queue a table's
compaction behind volume balancing for no reason - Iceberg has its own
lock-free lane for exactly this.

Adds the lane, maps the three job types to it, and puts it in the sidebar
beside Iceberg and Lifecycle. The lane routes were already generic, so only the
nav was hand-written.

The lane-coverage test spelled out the three known lanes, so a fourth failed
it. It now checks against AllLanes(), which is the property it was reaching for
and does not need editing next time.

Found by connecting the Rust worker to a real admin: it registered fine and its
job types were known, but they were filed under "default" and had no page.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: log what detection saw

"Detection proposed nothing" and "the worker could not read the table" look
identical from the admin side, and the second is what a missing credential
produces. One line per table separates them.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: fix a leaked heartbeat and a silent reconnect loop

spawn_heartbeat returned a handle to an empty task rather than the ticker it
had just spawned, so aborting it aborted nothing and every reconnect left
another heartbeat running against a dead channel.

A stream that admin closes cleanly is not an error, but reconnecting in silence
hides why. Two workers sharing an id evict each other forever and the log shows
nothing but a login every five seconds - which is exactly how this presented on
a dev box, and it took a look at the admin's own log to see it. The message now
names the id to check.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: a namespace cannot be created without its parent

Storage keeps a namespace's parts flattened, so creating "a.b" with no "a"
was accepted and left an intermediate that only existed inside a name. Listing
derives child names by slicing those parts, so it reported "a", while describe
and exists on "a" both answered 404 - a client walking the tree got a 404 on
something the listing had just handed it.

The spec asks for NamespaceNotFound when the parent is missing, which is also
what keeps listing and describe telling the same story.

Namespaces created through the S3 Tables API still bypass this, so listing
keeps deriving intermediates rather than hiding whatever is already there.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: say why a non-Iceberg table shows no schema

The table pages read Iceberg metadata for schema and snapshots, and a Lance
table has none, so both panels rendered "No schema available" - which reads as
an empty table rather than a table this page cannot describe. The dataset
behind the one that prompted this holds 1024 rows.

The format is already on the entry and shown two rows above, so the empty
states now use it: the catalog records where a LANCE table lives, not what is
in it.

Reading the schema for real needs Lance format code, which is the same wall as
the data plane.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* seaweed-worker: run rustfmt over the workspace

Committed the crates unformatted, so `cargo fmt --all --check` failed on
files nothing had touched since.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* plugin: let a worker report what it saw about an object

Admin cannot read a Lance table: it knows where the dataset lives and
nothing else, so the details page had a location and two empty panels.
The worker already opens every dataset during detection to decide whether
it needs compacting, so it knows the schema, the row count and the
fragment count at that moment. It just had no way to say so.

Add a WorkerObservations body to the worker stream. Admin caches the last
observation per object and serves it back, timestamped, for display;
nothing schedules from it. The Lance compaction sweep reports what it
opened, and the S3 Tables details page fills its schema panel from the
cache when it has no metadata of its own, badged with when the worker
looked and which worker it was.

Nothing about this is Lance-specific past the reporting side, which is
the point: any format admin cannot parse can describe itself the same way.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record the observation channel

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* plugin: ask a worker for sample rows of a table admin cannot read

Browse Data reads an Iceberg table's Parquet files directly, so it shows
real rows. For a Lance table it showed "Table has no Iceberg metadata"
and an empty grid, because there is no Go Lance reader and never will be
one worth maintaining.

The worker has the reader. Add RequestObjectPreview / ObjectPreviewResponse
to the stream, mirroring the config-schema round trip that already exists,
and give the Rust worker a PreviewProvider that scans the dataset and
formats the rows with Arrow's own formatter, so a vector column reads as a
vector. Admin picks the worker from the observation store: whichever one
last described this table is the one that can read it.

Unlike an observation the rows are not cached. They are the table's data
rather than a description of it, and a copy sitting in admin would be both
stale and nobody's business. The page fetches on load, bounded at 200 rows
and a 15 second round trip, and drops the snapshot and data-file panels
that only mean something for Iceberg.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record the preview channel

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: disable the lance listener when two gateways share a host

* test: keep AllocatePorts away from the lance default port

* s3tables: let a table bucket declare the format it holds

A bucket is a catalog, and a catalog serves one protocol. Format was
recorded per table, so nothing could answer "where do I point a client at
this bucket" without opening a table first, and an empty bucket had no
answer at all.

CreateTableBucket takes an optional format, stored with the rest of the
bucket metadata and returned by Get and List. Empty means ICEBERG, which
is what AWS S3 Tables serves and therefore what an SDK that has never
heard of the field means. CreateTable refuses a table of another format,
and CreateView refuses outright in a bucket that is not Iceberg, since a
view is Iceberg metadata.

Buckets that already exist carry no declaration and keep accepting
anything, so nothing is migrated and nothing that worked stops working.
The Lance namespace declares LANCE for the buckets it creates, which is
what stops one of them being described to a client as an Iceberg catalog.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: take the Lance port the way it takes the Iceberg one

The UI cannot name the endpoint that serves a Lance bucket without it,
and every format-aware page below needs to.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: show which format a table bucket holds

The bucket list printed an Iceberg endpoint for every bucket, including
ones holding Lance datasets, where that endpoint serves nothing. It was
the most visible place the UI assumed one format.

The list gains a Format column and its endpoint column follows the
bucket's declaration. The banner names both endpoints rather than
asserting everything is Iceberg, and says so only for the servers that
are actually running. Create Bucket picks a format with two cards rather
than a dropdown, since what matters is not the name but which clients can
read the result, and the endpoint under them updates as you choose so the
operator leaves the modal knowing where to point one.

A bucket from before the declaration existed shows "unset" in an outline
badge, explained on hover. It is a fact about the bucket's age, not a
fault, so nothing nags about it.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: carry the bucket's format into the pages inside it

Namespaces and tables are reached through a bucket, so both now say which
catalog they belong to rather than making you go back up to find out. The
tables list gains a Format column and a Rows column filled from what a
worker last observed, since for a format admin cannot read that is the
only row count there is; a table nothing has looked at shows a dash, not
a zero.

Create Table stops offering a choice the bucket has already made: in a
declared bucket the format is fixed and says why, and only an undeclared
one still offers both. Before this the select had exactly one option,
hardcoded, which made a Lance table impossible to create from the UI at
all.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: let the table page speak the table's own format

Partitions and Snapshot History are Iceberg's shape. Rendering them empty
for a Lance table reads as a fault; a Lance table has neither, and says
so by not showing them. In their place is a Versions panel, which is what
that format calls its history, carrying the worker's timestamp so it is
clear the numbers are a cached look rather than something read live.

The breadcrumb carries the format badge, so the page names what it is
looking at before you read a panel and wonder why it is empty.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: show how to connect to either catalog, and group the two format workers

The client examples on the buckets page were Iceberg's alone, so the one
thing an operator wants after creating a Lance bucket - what to type to
reach it - was not written down anywhere in the UI. Both formats now get
a pair of snippets, and only for a server that is running.

In the Workers menu, Iceberg moves below Lifecycle so it sits next to
Lance: the two table-format workers together, the two cluster-wide ones
above them.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* shell: create a table bucket of either format

s3tables.bucket -create takes -format, so a Lance bucket can be made
without going through the UI. The integration harness passes it too: its
Lance tests were creating Iceberg buckets and getting away with it only
because nothing checked.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* design: record that a bucket declares its format

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: drop managed versioning; the store already orders commits

The namespace offered itself as an external manifest store, so that a
commit could reserve a version through a real put-if-not-exists. That was
designed around a gateway that no longer exists: If-None-Match: * is
reduced to a filer WriteCondition and evaluated at the object's owner
under its per-path lock, or under the object write lock on the fallback
path. Sixteen writers racing one fresh key get a single 200 and fifteen
412s, every time.

Lance needs nothing else. commit_handler_from_url hands every s3:// dataset
a ConditionalPutCommitHandler, which puts with PutMode::Create, which
object_store sends as If-None-Match: *. So the feature solved a problem
this store does not have, while moving a table's version history out of
the dataset and into the catalog - and lance could not use it past the
first commit anyway, since its own namespace-backed store answers
"put_if_not_exists is not supported" to the second.

The version operations answer Unsupported with the rest, managed_versioning
is false, and the flag is gone. In place of the reserve-once test there is
one that races eight writers at the manifest key through S3, which is the
path a commit actually takes.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: honour the version floor, the slot limits, and a shutdown

Five findings from review, all of them things the worker claimed to do and
did not.

The version floor was checked when a cleanup job was proposed and ignored
when it ran, so a table whose versions had aged past the retention window
in between could be taken below the count the operator asked to keep.
Execution now computes the floor itself and passes it as before_version;
CleanupPolicy ANDs its clauses, so a version has to be both too old and
below the floor to go. Both settings are clamped to the range the form
offers, since Duration::hours panics on a large enough value and a
negative min-versions wraps to a huge usize.

Admin's shutdown was answered by returning from the stream, which the
reconnect loop read as a healthy close and logged straight back in: the
worker could not be stopped. serve_once now says which of the two
happened.

The advertised concurrency limits bounded nothing - every request spawned
a task - and the heartbeat reported zero slots in use whatever was
running. Both now go through semaphores sized from the limits, with the
permits held for the life of the request and reported in the heartbeat.

A namespace call had no timeout, so a gateway that accepted the connection
and went quiet held a detection slot forever. And one table whose stats
could not be read failed the whole sweep, losing the proposals for every
table already scanned; it is now skipped and warned about, like a table
that cannot be opened.

The tests drove one shared catalog concurrently, which is why one of them
asserted "no proposals at all" and passed by luck. They now take a lock
and judge only their own tables.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* admin: fix the review findings on the format-aware pages

The endpoint hint in Create Bucket built its HTML by concatenating the
bucket name the operator is typing, so a name like <img onerror=...> ran
in the admin origin as they typed it. It is built from DOM nodes now.

A preview reply looked its channel up under the lock and then sent outside
it, which Shutdown can close in between: a Gosched in that gap panics with
"send on closed channel" every time. The send now happens under the lock.

Observations were looked up by path alone, so a table dropped and remade
in another format at the same path was described by the observation left
behind. Lookups now have to agree on the format.

Also: the Lance namespace caps a request body rather than reading whatever
arrives; the details action no longer says "Iceberg" over a Lance table;
mini stops advertising a catalog port when it is not running S3; a format
whose server this cluster does not run cannot be picked in the modal or
accepted by the API, since a bucket nothing can reach is not worth
creating; and the unused catalogPortFor helper is gone.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: let the control stream use mTLS

The channel was hardcoded to http://, so off loopback the stream carried
preview rows and execution commands in the clear - and a cluster with grpc
TLS turned on would refuse the worker outright.

--tls-ca, --tls-cert and --tls-key take the same certificates the Go
worker reads from the [grpc.worker] section of security.toml, and must be
given together: a CA on its own would quietly mean one-way TLS, which a
mutual setup rejects anyway. Without them the stream stays plaintext,
which is what the Go worker also does when nothing is configured.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: answer null properties rather than an empty map

The catalog does not keep a table's properties. Declare echoed the
request's back and describe answered {}, both of which claim they were
stored and are empty. Null says the catalog does not keep them, which is
what the spec distinguishes and what is true here.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance worker: test the slot accounting

The heartbeat reporting and the waiting are the two things the semaphores
are for, and neither is observable from outside without catching a sweep
mid-flight.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* test: fix the mixed-format catalog test, and name the binary it drives

The integration suite passed locally and failed in CI on
TestLanceRefusesIcebergTables. Both were right: CI builds the binary
first, my tree had one from the day before, so locally the test drove a
gateway with no format enforcement at all.

The test itself no longer holds as written. It made a bucket, put an
Iceberg table in it, and checked the Lance surface hid it - but a bucket
that declares LANCE now refuses the Iceberg table outright. The invariant
still matters from the other side, so it starts from an Iceberg bucket
instead: Lance must not describe or list a table whose format it does not
serve, and must refuse to declare one beside it.

The harness now prints which weed binary it is about to run and when that
was built. `make test` rebuilds first; a plain `go test` will happily
drive a weeks-old binary and report a pass for code it never ran, which is
exactly what happened here.

Also make the row-limit conversion in the preview request explicitly
bounded: CodeQL flagged the int-to-int32 conversion, and clamping by
reassignment beforehand is not a form it recognises.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* lance: prove concurrent commits are kept, and preselect the only format on offer

Two more from review.

The commit test asserted that exactly one writer wins the conditional PUT,
which is the mechanism, not the claim. The claim is that nothing is lost:
the losers see the conflict, rebase and commit again. So there is now a
test that has eight writers append to one dataset at once and counts the
rows afterwards - all eight batches survive. That is also the sequence
managed versioning could not finish, since its store refuses the second
commit outright.

And when Iceberg's endpoint is not running, the format picker offered two
options with neither selected, so Create Bucket submitted no format at
all, fell back to ICEBERG, and was refused by the guard added last round.
Lance is preselected when it is the only format this cluster serves.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* Clamp the remaining worker settings, and bootstrap buckets in a served format

Compaction and index optimization read their thresholds and cast straight
to usize and u64, so a negative arrives as an enormous number and turns
the threshold into "never": compaction and reindexing both go quiet with
nothing to say. The cleanup job was fixed last round; these are the same
bug. Clamped to the values that stay meaningful rather than to what the
form offers - zero uncovered rows is a real setting, meaning reindex as
soon as anything is not covered, so the floor there is zero and not the
form's thousand.

mini pre-creates the buckets named by -tableBucket, and did so without a
format, which now means Iceberg. Started with the Iceberg endpoint off
and the Lance one on, that left buckets nothing could reach and which
refused every Lance table. It takes the format from the endpoint that is
actually running, and creates nothing when neither is.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm

* s3: allow-unordered is a listing parameter, not an unimplemented subresource

The guard that stops a bucket GET with an unknown subresource from being
answered with a listing does not know about allow-unordered, so it answers
501 NotImplemented - to a parameter the listing handlers already read and
already validate against delimiter.

This is why test_bucket_list_unordered and test_bucket_listv2_unordered
fail in the Ceph s3-tests suite. They fail on master too; this is not a
Lance change and can be taken on its own.

Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
2026-08-19 22:59:56 -07:00

39 KiB

Lance Catalog for SeaweedFS

A second catalog surface next to the Iceberg REST catalog, speaking the Lance Namespace REST spec, over the same table buckets and the same filer.

Why

Gravitino 1.1 added a Lance REST service and 1.3 ships it as a standalone server; Lakekeeper added Lance in the same window by a completely different route. That is the useful signal: two unrelated catalogs decided independently that Lance had to be first-class, not a niche. The client side is already there — lance-spark (LanceNamespaceSparkCatalog with impl=rest), lance-ray, and the generated Python/Java/Rust clients all talk the same OpenAPI. Implementing the spec means those engines work against SeaweedFS with no SeaweedFS-specific code on the client.

The second reason is that Gravitino's own documentation names the gap it cannot close: DuckDB, pandas and DataFusion "do not support Lance REST natively yet" and have to fetch a location from the catalog and then open the dataset directly. Gravitino cannot help there, because it does not own the storage. SeaweedFS does. That is the whole design opportunity below.

Prior art: three families

Upstream lists twelve catalog implementations, and they fall into three shapes. Knowing which one we are building matters more than any individual API decision.

1. Storage-native, no service. The Lance Directory Catalog. V1 is a directory listing where every <name>.lance/ child of a prefix is a table; V2 adds a __manifest table — itself a Lance table — holding object_id/object_type/location rows, with nested namespaces, hash-prefixed table directories, and optional managed versioning. No server, no credentials, no governance. This is the floor every other implementation has to beat.

2. Protocol-native server. Someone implements the Lance Namespace REST OpenAPI and clients connect with impl=rest. Gravitino is the only one of the twelve that does this, and it is what this design proposes.

3. Client-side adapters onto an existing catalog. Nine of the twelve. The Lance client translates namespace operations into whatever the backing catalog already speaks: Apache Polaris, Unity Catalog, AWS Glue, Hive Metastore v2 and v3, Google BigLake, Dataproc, Microsoft OneLake — and Apache Iceberg REST. Two flavors:

  • Catalogs with a real non-Iceberg table concept mark the format directly. Polaris uses its Generic Table API with format = lance; Unity uses an EXTERNAL table with table_type=lance in properties and the path in storage_location; Glue uses EXTERNAL_TABLE plus table_type=lance in Parameters, path in StorageDescriptor.Location.
  • Catalogs with no such concept fake one. The Iceberg REST adapter registers a regular Iceberg table with a dummy schema — a single nullable string column named dummy — carrying the property table_type=lance, and treats the Iceberg table location as the Lance dataset root.

Every adapter in family 3 lands in the same place: DeclareTable/ListTables/ DescribeTable/DeregisterTable only, DropNamespace in RESTRICT mode only, load_detailed_metadata=false only, and managed_versioning=false. They are a name-to- location map and nothing more.

Lakekeeper is the instructive outlier. It has the same generic-table concept Polaris has, but no upstream adapter exists for it — there is no lance-namespace reference anywhere in its repository and no page for it in the supported-catalogs list. So Polaris's generic tables are reachable from a stock Lance client and Lakekeeper's are not, despite being the same idea. Shipping the concept is not the same as shipping the integration.

Gravitino and Lakekeeper: the two opposite bets

Both shipped Lance support in the same window and did not build the same thing.

Gravitino implements the protocol. Its lance/ module serves the Lance Namespace REST spec on its own port (:9101/lance), so stock lance-spark and lance-ray connect with impl=rest and no vendor-specific client. The cost is governance: storage credentials are static properties on the catalog (lance.storage.access_key_id, secret_access_key, endpoint, region, allow_http), optionally overridden per table, handed to the engine as-is. No STS, no expiry, no per-table scoping.

Lakekeeper refuses the protocol and governs the object instead. There is no lance-namespace anywhere in the repository; Lance arrived in 0.13.0 (2026-06-30, issue #1673 Generic Table API with Lance) as one format string on a Lakekeeper-native Generic Table API:

POST/GET/DELETE /lakekeeper/v1/{prefix}/namespaces/{ns}/generic-tables[/{table}]
GET             /lakekeeper/v1/{prefix}/namespaces/{ns}/generic-tables/{table}/credentials
POST            /lakekeeper/v1/{prefix}/generic-tables/rename

format is opaque, schema and statistics are stored but never validated, and the catalog writes no format-specific metadata — engines go straight to the location. In exchange Lance tables get everything Iceberg tables get: STS-vended prefix-scoped credentials, OpenFGA per-action permissions (16 actions), soft-delete with undrop, a protection flag, rename, pagination, and name uniqueness across Iceberg tables, views and generic tables in one namespace. The price is that no stock Lance client can talk to it — you need pylakekeeper, which exists mainly to translate vended credentials into lance_storage_options.

So: protocol fidelity and weak governance, or strong governance and client lock-in. Both documented their limit honestly, and it is the same limit. Lakekeeper's capability table says it outright — "Commit coordination: the catalog does not arbitrate writes — engines write directly." Gravitino does not claim it either. Neither of them coordinates a Lance commit, which is exactly the thing a store can do and a control plane cannot.

We do not have to choose. Serve the Lance protocol natively the way Gravitino does, over the s3tables entries that already carry ARNs, policies, tags and maintenance config, and the governance comes from the layer underneath rather than from a proprietary API on top. That is only available to us because we are the store, which is also what makes the third option — arbitrating the commit — available.

We are probably already a Lance catalog, and that is a problem

The Iceberg REST adapter does not care whose Iceberg catalog it is talking to. It needs /v1/config?warehouse=, /v1/{prefix}/namespaces, /v1/{prefix}/namespaces/{ns}/tables and unit-separator (\x1F) multi-level namespaces. We serve all of those, and parseNamespace in weed/s3api/iceberg/utils.go:22 already splits on \x1F. So a stock Lance client pointed at our Iceberg catalog on :8181 with the Iceberg impl should already create, list, describe and deregister Lance tables today, with no SeaweedFS change at all.

That is worth testing before writing a line of the design above, for two reasons. It is a free baseline — and possibly a free announcement. And it is a data-loss hazard.

A Lance table registered this way is an Iceberg table whose metadata references no data files, sitting on top of a Lance dataset that uses data/ for its fragments — the same subdirectory name Iceberg uses. The maintenance worker's orphan cleaner walks exactly <table>/metadata and <table>/data, and deletes every file not referenced by a snapshot and older than orphan_older_than_hours (weed/worker/tasks/iceberg/operations.go:331, default 72). Against an adapter-registered Lance table, every fragment is unreferenced by construction. Run maintenance and the dataset is deleted.

Maintenance is disabled by default (handler.go:334), so this is a latent hazard rather than a live one: it needs an operator to enable Iceberg maintenance on a bucket that also holds adapter-registered Lance tables. But it costs nothing to close — detection should skip any table carrying a non-Iceberg format marker (table_type property, or Format != "ICEBERG" once the format field is honest), and that guard is worth landing on its own regardless of whether the rest of this design ever gets built. It is the same "catalog-only, no maintenance" marker the generic-format question needs.

Where we differ from Gravitino

Gravitino is a metadata service in front of somebody else's object store:

  Spark / Ray                  Spark / Ray / pandas / duckdb
      |                              |
  Lance REST                    Lance REST          (direct S3)
      |                              |                   |
  Gravitino                     SeaweedFS S3 gateway ----+
      |                              |
   S3 keys handed out           SeaweedFS filer + volumes
      |
  somebody else's S3

It resolves a name to a location plus lance.storage.* credentials, and steps out of the way. Everything a Lance table actually is — _versions/, data/, _indices/ — is opaque to it.

We are the store. Three things follow that Gravitino cannot do:

  1. The catalog and a plain directory listing can be made to agree, so a client with no catalog at all still sees the right tables.
  2. _versions/ is a filer directory listing, not an object-store LIST. Version history is cheap and can back the admin UI.
  3. We can offer a genuinely atomic commit reservation. Lance's commit protocol needs put-if-not-exists; our S3 layer does not currently provide one (see Commit safety). The filer does.

Placement

The Iceberg catalog is a thin HTTP shell over s3tables.Manager; the storage work lives in weed/s3api/s3tables. Table buckets live under TablesPath = s3_constants.DefaultBucketsPath, i.e. the same filer tree the S3 gateway serves, so s3://bucket/ns/table/ is simultaneously a catalog entry and an S3 prefix. Catalog entries are filer directories carrying s3tables.* extended attributes. Table.Format already exists and is hard-checked against "ICEBERG" in weed/s3api/s3tables/handler_table.go:48.

So:

weed/s3api/lance/          new: HTTP surface, id codec, error model
weed/s3api/s3tables/       extended: Format "LANCE", lance state xattr, version entries
weed/command/s3.go         new: -port.lance (default 9101), startLanceServer

Format: "LANCE" on the table entry is the whole storage-model change for phase 1. Everything else — namespaces, ARNs, policies, tags, ownership — is shared verbatim.

                        s3tables.Manager (filer)
                                 |
        +------------------------+------------------------+
        |                                                 |
  weed/s3api/iceberg                              weed/s3api/lance
   Iceberg REST :8181                              Lance REST :9101
        |                                                 |
  Iceberg tables                                   Lance datasets
        \                                                 /
         +-------------------- s3 :8333 -----------------+
                                 |
                         SeaweedFS volumes

Identifier mapping

Lance identifiers are ["ns", ..., "table"], encoded in the URL as a single string joined by a delimiter that defaults to $. The delimiter alone means the root namespace, so /v1/namespace/$/list lists the root's children.

Iceberg had to invent a warehouse selector because its identifier is flat and every table bucket is a separate catalog. Lance does not need that — its identifier is already hierarchical, and Gravitino uses exactly three levels (["lance_catalog", "sales", "orders"]). That maps onto us without inventing anything:

  $                            root      -> list of table buckets
  $analytics                   level 1   -> a table bucket
  $analytics$sales             level 2   -> a namespace in that bucket
  $analytics$sales$orders      table

spark.sql.catalog.lance.parent = analytics then makes sales.orders resolve, which is the same shape Gravitino's Spark example uses.

Levels 2..N join into one s3tables namespace with ., matching what the Iceberg catalog already does with flattenNamespacePath. The flattened form is only the directory name — namespaceMetadata.Namespace []string in the xattr keeps the authoritative parts, so the mapping stays invertible even though . is a legal character inside a namespace part. Reject $ in any name part with InvalidInput; our charsets already exclude it, so no escaping scheme is needed.

Root-level ListNamespaces returning table buckets means an unauthenticated or broadly-scoped caller can enumerate buckets. Filter it through the same s3tables/permissions.go check ListTableBuckets uses, not a separate path.

CreateNamespace on a one-part identifier creates a table bucket, and it does so only if the caller is permitted to — the namespace never creates a bucket as a side effect of creating something inside it. A table bucket is a tenant resource with its own policy, ARN and lifecycle, and conjuring one because a client said CREATE SCHEMA is a privilege escalation dressed as a convenience. Lakekeeper draws the same line explicitly: its client creates tables, not warehouses.

Storage layout

Lay tables out as:

s3://<table-bucket>/<flattened-namespace>/<table>/
    data/
    _versions/
    _indices/

Built without the .lance suffix this design originally proposed. The suffix would have made every namespace prefix a valid Lance Directory Catalog V1 root, since V1 recognises a table by exactly that naming. It does not survive contact with the storage layer: the catalog entry is the dataset directory, validateTableName excludes . from the charset, and a suffixed entry name would leak into ARNs, policy documents and the S3 Tables API, where the same table would answer to two different names. Making GetTablePath format-aware instead spreads an "unless it is Lance" branch through code that has no business knowing — the exact cross-cutting cost this design rejects family 3 for.

So one name, one directory. What survives is direct access by URI, which is the larger half of the story and needs no naming convention at all:

# with the catalog
spark.sql("SELECT * FROM lance.sales.orders")

# without it, same bytes
lance.dataset("s3://analytics/sales/orders")

DuckDB, pandas and DataFusion still reach the data with no catalog running, which is the gap Gravitino's documentation admits to. What they no longer get for free is enumeration — a directory-catalog client pointed at the namespace prefix will not list these as tables. If that turns out to matter, the cheapest fix is a repair-style tool that materialises .lance aliases, not a rename of the catalog entry.

Note also that the directory catalog's own V2 mode puts child-namespace tables in <hash>_<ns$table> directories at the root and creates no physical subdirectories for namespaces, so full directory-catalog fidelity was never on offer anyway. We are a server-backed catalog; the human-readable prefix layout is worth more than partial V1 lookalike behaviour.

The table bucket was not a neutral container

This design assumed a table bucket is a place to put a table's files. It is not: validateTableBucketObjectPath runs on every S3 write into one and validated the path against Iceberg's layout, so a Lance client got 403 on data/*.lance, on _versions/, and on _transactions/ — a directory Lance writes that neither the spec documentation nor this design anticipated. Nothing about the catalog worked end to end until that changed.

The layout guard now admits the union of what the supported formats write, and treats any underscore-prefixed top-level directory as belonging to the format, checking only that the path stays inside the table. Enumerating Lance's internal directories by name is exactly the mistake that missed _transactions. Iceberg writes none of them, so it loses nothing.

Found by pointing the real Python client at a running gateway, not by reading the spec. Worth remembering for the next format: the premise to check first is whether the storage layer will accept its files at all.

Table lifecycle

Lance has three table states, and the spec pins them to marker files:

State Marker Created by Visible in ListTables
declared .lance-reserved DeclareTable yes, when include_declared=true
created _versions/ present client writes, or CreateTable yes
deregistered .lance-deregistered DeregisterTable no; data preserved

Record the state in an xattr (s3tables.lanceState) on the catalog entry and write the marker file into the table directory. The xattr is what the catalog reads; the marker is what keeps a directory-catalog client honest. Dual-write is the price of the interop claim above, and it is one extra filer write on three rarely-called operations.

DeclareTable is the operation lance-spark actually calls on CREATE TABLE (it replaced the legacy create-empty), so it is not optional in practice even though the spec marks only a subset as required.

DeregisterTable preserving data is the same shape as our Iceberg rename, where the catalog entry moves and the data stays put — reuse TableDataDirFromMetadataLocation's idea rather than re-deriving the data path from the catalog name.

Commit safety

This is the part I got wrong, and the correction removed a feature rather than adding one.

Lance commits a version by writing _versions/{v}.manifest with put-if-not-exists: exactly one writer is supposed to win, and the loser rebases. In lance 10 that path is not optional and needs nothing bolted on — commit_handler_from_url hands every s3:// dataset a ConditionalPutCommitHandler, which calls put_opts with PutMode::Create, which object_store's S3 backend sends as If-None-Match: *.

I originally read our gateway as evaluating that header check-then-act, and designed around it. That was already out of date. buildWriteCondition (weed/s3api/s3api_object_routed_write.go) reduces If-None-Match: * to a filer WriteCondition{IF_NOT_EXISTS}, and putToFiler routes the create to the object's owner filer, which evaluates the precondition under its per-path lock; when routing is not available it falls back to the object write lock, which evaluates it under the lock too. Either way it is atomic. Sixteen concurrent writers of the same fresh key get one 200 and fifteen 412s, repeatedly.

So the store already has the primitive Lance needs, cluster-wide, for every conditional-PUT client and not just this one.

What that removed

An earlier draft of this design offered the catalog as an external manifest store: managed_versioning: true plus CreateTableVersion and friends, with the reserve step as a filer CreateEntry with o_excl. It was implemented, tested, and shipped behind a default-off flag — and it should not exist.

  • It solves a problem this store does not have. The spec offers that path for stores that cannot order commits themselves.
  • It moves a table's version history out of the dataset and into the catalog, so a reader that does not go through this namespace no longer sees the whole picture. That is a real cost paid for nothing.
  • lance 10 cannot even use it past the first commit: NamespaceManifestStore::put_if_not_exists answers "put_if_not_exists is not supported for namespace-backed stores", which is exactly what a second append needs.

The version operations now answer Unsupported alongside the other operations the catalog does not serve, and managed_versioning is answered false. The property they were protecting is covered instead by a test that races eight writers at the manifest key through S3 and asserts one wins — testing the path Lance actually takes.

Credential vending

Iceberg needed a header (X-Iceberg-Access-Delegation: vended-credentials) and a bespoke response shape. Lance has it in the spec: vend_credentials: true on the request, storage_options on the response, with expires_at_millis as the well-known expiry key.

Reuse the existing vendor interface unchanged — iceberg.CredentialVendor / STSService.AssumeRoleForPrincipal scoped to the table prefix (#10777) — and map its output to the storage options Lance passes through to object_store:

aws_access_key_id, aws_secret_access_key, aws_session_token,
aws_region, aws_endpoint, allow_http, expires_at_millis

Those are the names pylakekeeper emits as lance_storage_options, which is the shape Lakekeeper's tested S3 path actually feeds to Lance. object_store also accepts the un-prefixed aliases (endpoint, region) that the directory catalog's storage. prefix strips down to and that Gravitino's lance.storage.endpoint resolves to, but the aws_ forms are the ones with a tested integration behind them, so emit those. aws_endpoint should come from deriveS3AdvertisedEndpoint(), the same source the Iceberg FileIO config uses, and allow_http must be set when that endpoint is plain HTTP or every read fails with a TLS error that looks like a credential problem — Lakekeeper vends both automatically for exactly this reason, and calls out that there is then no per-vendor branch in client code.

We emit this server-side, in the storage_options field the Lance spec already defines, which is strictly better than Lakekeeper's arrangement: no client library has to translate anything, so vending works from any stock Lance client rather than only from theirs.

Guard the same way #10777 had to after review: bucket-scoped list grants need an s3:prefix condition, and a location containing * or ? must be refused rather than widened into a resource pattern.

Auth and authorization

Authentication reuses S3Authenticator and CredentialValidator as-is. The Lance spec maps identity to headers — api_key to x-api-key, auth_token to Authorization: Bearer — and SigV4 keeps working because it is the same authenticator the Iceberg catalog already fronts.

Authorization needs nothing new. A Lance table gets the same ARN shape, arn:aws:s3tables:...:bucket/B/table/NS/T, so every existing table-bucket policy covers Lance tables with no new policy language and no second permission model. Route it through s3tables/permissions.go and inherit the DefaultAllow semantics the Iceberg server already mirrors from the S3 port.

One spec quirk worth honoring: request context entries prefixed header. become request headers, and every response header comes back as a header.-prefixed context entry. Echoing x-request-id through it costs nothing and makes tracing work.

What to take from Lakekeeper

Rejecting Lakekeeper's API shape does not mean rejecting what it learned building it.

Deregister is soft-delete, so implement it as one. Lakekeeper gives generic tables soft-deletion with undrop and a protected flag that makes a drop require force=true. Lance already has the concept — DeregisterTable preserves the data and hides the table — so the .lance-deregistered marker is a soft-delete by another name, and a re-register is an undrop. A protection flag on table-bucket entries is worth having regardless of Lance: it is a few lines against the existing xattrs and it applies to Iceberg tables too.

Enforce one identifier space across entry kinds. Lakekeeper rejects a generic table whose name collides with an Iceberg table or view in the same namespace. Our catalog entries already share one filer directory and already carry s3tables.entryType, so this is structurally true — but it has to be enforced deliberately on every path, or a Lance handler happily loads an Iceberg table's directory and vice versa. That is the same crossover bug class as the view/table rename authorization fixed in #10776; the catalogEntryKind pattern from that change is the thing to reuse rather than re-derive.

A re-vend path matters more than it looks. Lakekeeper exposes /credentials separately from load, because STS credentials expire in the middle of long jobs and re-loading the whole table to refresh them is wasteful. In Lance the spec's answer is another DescribeTable with vend_credentials: true, which is fine — but it means DescribeTable must stay cheap when load_detailed_metadata is false, which is another reason not to open the dataset on that path.

Generic tables are a cheap orthogonal win. Lakekeeper's real insight is that Delta, Parquet, CSV, Vortex and Paimon all get governance for free once the catalog stops caring what the format is. Our Table.Format field already exists and the only thing stopping it is the hard "ICEBERG" check in handler_table.go:48. Loosening that and letting the S3 Tables API register a table with an arbitrary format and a location — no metadata, no commits — is a small change that makes every format cataloguable. It is independent of this design and probably worth doing first, since Format: "LANCE" is then just a value rather than a special case.

Skip remote signing. It is Lakekeeper's fallback for S3-compatible stores with no STS, and their own documentation notes that Lance will not use it — format libraries with their own S3 client expect static credentials and do not implement the Iceberg signer protocol. We have STS, so vended credentials are the path, and the signer is not worth building for a client that cannot consume it.

Errors

Lance uses {code, error, detail, instance} with numeric codes, not Iceberg's exception-type strings. The mapping is mechanical:

HTTP code when
400 13 InvalidInput charset violations, malformed id, route/body mismatch
401 16 Unauthenticated
403 15 PermissionDenied
404 1 NamespaceNotFound, 4 TableNotFound, 11 TableVersionNotFound
409 2/5 AlreadyExists, 3 NamespaceNotEmpty, 14 ConcurrentModification
501 0 Unsupported every phase-3 data operation

Route/body mismatch is a spec requirement, not a nicety: when the identifier appears in both the path and the body and they disagree, the server must return 400. Cheap to get right at the decode step, annoying to retrofit.

Route surface

Phase 0 is not in this table: point a stock Lance client at the existing Iceberg catalog with the Iceberg impl, see how far it gets, and land the maintenance guard either way. That tells us what the native server actually has to beat.

Phase 1, the whole lance-spark and lance-ray contract:

POST /v1/namespace/{id}/create         CreateNamespace       mode: Create|ExistOk|Overwrite
GET  /v1/namespace/{id}/list           ListNamespaces
POST /v1/namespace/{id}/describe       DescribeNamespace
POST /v1/namespace/{id}/drop           DropNamespace         mode: Fail|Skip, behavior: Restrict|Cascade
POST /v1/namespace/{id}/exists         NamespaceExists
GET  /v1/namespace/{id}/table/list     ListTables            ?include_declared, ?page_token, ?limit
GET  /v1/table                         ListAllTables
POST /v1/table/{id}/declare            DeclareTable
POST /v1/table/{id}/describe           DescribeTable         ?with_table_uri, ?load_detailed_metadata, ?check_declared
POST /v1/table/{id}/exists             TableExists
POST /v1/table/{id}/register           RegisterTable         mode: Create|Overwrite
POST /v1/table/{id}/deregister         DeregisterTable
POST /v1/table/{id}/drop               DropTable
POST /v1/table/{id}/rename             RenameTable

DescribeTable with load_detailed_metadata=false needs only location, which is the common case and which we can answer from xattrs alone. With load_detailed_metadata=true the spec wants version, schema and stats, which means reading the Lance manifest. For phase 1, return the fields we can derive from the filer — version from the highest entry in _versions/, given V2 naming is {u64::MAX - version:020}.manifest and V1 is {version}.manifest — and omit schema/stats rather than fabricating them. The spec tolerates a partial response here; it does not tolerate a wrong one.

Phase 2 was the five version operations plus managed_versioning; it was built and then removed, for the reasons under Commit safety.

Phase 3 is the data plane: CreateTable, InsertIntoTable, MergeInsertIntoTable, UpdateTable, DeleteFromTable, QueryTable, CountTableRows, and the index and tag operations. These exchange Arrow IPC, and more to the point they require reading and writing the Lance file format, for which no Go implementation exists. Return Unsupported (code 0) and say so in the docs. arrow-go/v18 is already an indirect dependency, so Arrow framing is not the blocker — Lance is.

Does a Lance table need maintenance?

Yes, and one part of it has no Iceberg equivalent. The client exposes three jobs:

  • optimize.compact_files() — Lance writes a fragment per write batch, so a table fed by small appends accumulates small files exactly the way an Iceberg table does.
  • optimize.optimize_indices()rows written after an index was built are not covered by it. A vector search against a stale index silently misses recent data. That is a correctness-shaped failure, not a slow query, and it is specific to what people use Lance for.
  • cleanup_old_versions() — every version is retained until something removes it. Lance can do this itself: optimize.enable_auto_cleanup() sets it on the dataset, so this one need not be an external job at all.

None of it can run in the Go worker. All three read and rewrite Lance files, which needs Lance format code that does not exist in Go, and there is no useful subset either: deciding which fragments an old version still references means parsing Lance manifests.

So the maintenance worker must not touch a Lance table, and it declines by reading the format the catalog recorded rather than by failing to parse Iceberg metadata.

The worker can be Rust, and it is not a sidecar

The Go worker is not the only worker. weed/pb/plugin.proto defines PluginControlService, a language-agnostic gRPC stream that external maintenance workers connect on: the worker opens WorkerStream, sends WorkerHello with the job types it can detect and execute, answers RequestConfigSchema with a JobTypeDescriptor, replies to RunDetectionRequest with JobProposals and to ExecuteJobRequest with JobProgressUpdates and JobCompleted. weed worker -admin=host:23646 is the Go reference implementation of exactly that contract, from outside the admin process.

Nothing in it is Go-specific, and the Rust toolchain is already in the tree. seaweed-volume/build.rs compiles protos straight out of ../weed/pb/ with tonic_build, including filer.proto, on tonic 0.12 and prost 0.13. A Lance worker is that same build with plugin.proto added and the lance crate as a dependency — the real one, no FFI and no Python.

Three job types, one per real maintenance operation:

Job type Calls Detected from
lance_compact optimize.compact_files fragment count and sizes
lance_optimize_indices optimize.optimize_indices rows an index does not cover
lance_cleanup_versions cleanup_old_versions version count and age

What the existing machinery then supplies for free is the part worth noticing. Scheduling, retries, dedupe by dedupe_key, progress reporting, per-job concurrency limits and the admin settings page all come from the protocol: a worker that answers RequestConfigSchema with a descriptor gets its configuration form rendered in the admin UI without a line of Go or templ. A Rust worker is a first-class maintenance worker, not an appendage.

The remaining wiring is small and mostly decided already. RunDetectionRequest carries a ClusterContext with filer and S3 addresses plus a free-form metadata map, which is where the Lance namespace URL goes; the worker lists Lance tables from the namespace, which is the catalog of record and already filters by format. It gets at the data by asking DescribeTable for storage_options with vend_credentials, so the worker is just another client of the STS path rather than a component with its own credentials. And when it commits a compaction it goes through CreateTableVersion like any other writer, which is what managed versioning was for.

The worker is also the only thing that can describe the table

Admin can render an Iceberg table because it can read Iceberg metadata. It cannot read Lance: it knows the dataset's location and its format string, and that is the whole of it. The details page showed a location and two empty panels, which is an honest answer and a useless one.

The worker already knows. Detection opens every dataset to decide whether it needs compacting, so at that moment it holds the schema, the row count, the fragment count and the version count. It just had no way to say so — every message on the stream was about work.

So WorkerObservations is a body on WorkerToAdminMessage: a repeated ObjectObservation of object_id, object_kind, format, and a ConfigValue map the worker fills with whatever it can cheaply say. Admin keeps the last observation per object and serves it back with the time it was taken and the worker that took it. Nothing schedules from it, and it is not authoritative — it is a cache with its staleness on the label, which is why the page badges it rather than presenting it as metadata it read itself.

The keys are the worker's to choose, which keeps the protocol out of the business of knowing what a Lance table is. A worker for any other format admin cannot parse describes itself the same way.

A bucket declares its format

Format was recorded per table, which is enough for the storage layer and not enough for anything that has to answer a question about a bucket. The admin UI printed one Iceberg endpoint for every bucket, including the ones holding Lance datasets, where that endpoint serves nothing; an empty bucket had no format at all.

So CreateTableBucket takes an optional format, stored with the rest of the bucket metadata. Empty means ICEBERG - what AWS S3 Tables serves, and therefore what an SDK that has never heard of the field means. CreateTable refuses another format, and CreateView refuses outright outside an Iceberg bucket, a view being Iceberg metadata. The Lance namespace declares LANCE for the buckets it creates.

Enforced rather than defaulted, because the point of showing a format at all is the endpoint that follows from it, and that endpoint is only truthful if the bucket holds one format. Buckets that already exist stay undeclared and keep taking anything: nothing is migrated, and the UI shows "unset" as a fact about the bucket's age rather than a fault. That state is also the only way to hold both formats at once, which is what the Iceberg-REST adapter path produces.

Sample rows are fetched, not cached

The same asymmetry has a second half. Admin renders an Iceberg table's rows by reading its Parquet files directly; for Lance it has nothing to read with, so the data page offered a Browse Data button that led to an empty grid.

RequestObjectPreview / ObjectPreviewResponse mirror the config-schema round trip already on the stream: admin asks, the worker scans the dataset and hands back rows it has already rendered as text, because it is the only side that knows the types. Admin picks the worker from the observation store, so the one that last described a table is the one asked to read it.

The rows are deliberately not cached, and that is the line between the two channels. An observation describes an object, so a copy with a timestamp on it is useful. Rows are the object's contents: a copy held in admin would be stale, larger, and nobody's business. The page fetches on load, bounded, or says why it cannot.

The sidecar question

The data plane is a different problem, and this design previously conflated the two. Maintenance rides the worker protocol; QueryTable and InsertIntoTable do not, because they are synchronous REST operations on the namespace's own surface. Serving those means a Rust process that answers HTTP, either behind the Go namespace as a proxy target or in front of it. It would make SeaweedFS a store you can run vector search in rather than one you read vectors out of, which is the larger prize and the reason to keep the option open.

Neither should gate phase 1. Phases 1 and 2 are pure Go over the filer and are worth shipping on their own — they are what makes Spark and Ray work.

Testing

Mirror the Iceberg package: httptest plus a fake filer client for the handler tests, in weed/s3api/lance. Then an integration suite under test/s3tables/catalog/ next to the existing pyiceberg_test.go, driving the generated Python lance-namespace client against a live gateway. Three things that suite must cover and unit tests cannot:

  • the storage-options key names actually work, i.e. a client that gets storage_options from DescribeTable can open the dataset;
  • a table created through the catalog is visible to lance.dataset() by URI and to a V1 directory-catalog client rooted at the namespace prefix;
  • concurrent writers do not lose a commit, which is the phase-2 acceptance test and the thing that justifies the external manifest store.

Phase 1 is validated: lance_namespace 0.11.1 with impl=rest drives the namespace, lance.write_dataset writes to the vended location with the vended storage_options, and the rows read back. Note that this client version drops check_declared and include_declared on the wire, so is_only_declared reads null through it however the server behaves.

The commit path is validated at both levels. The mechanism: eight writers race the same manifest key through S3 with If-None-Match: *, and exactly one wins. The property that actually matters, which single-winner exclusivity does not by itself establish: eight writers append to one dataset concurrently through lance, and afterwards every batch is still there — the losers saw the conflict, rebased, and committed again. That second test is also the sequence managed versioning could not complete at all, since its store answers "put_if_not_exists is not supported" to the second commit.

One more that belongs in the Iceberg suite, not this one: a Lance dataset registered through the Iceberg adapter must survive a full maintenance pass. Reading the code, that test should fail today; it has not been run.

Open questions

  • Root-level ListNamespaces enumerating table buckets is convenient and is a listing surface we do not have on the Iceberg side. Decide whether it is gated behind a flag.
  • Whether the .lance directory suffix is worth the divergence from the Iceberg layout. I think yes — it is what makes the catalog optional — but it means the two catalogs' tables do not look alike on disk, and the admin UI has to know that.
  • Names: our charsets are lowercase-only and Lance identifiers are arbitrary strings. Reject and document, as Iceberg does, or case-fold. Rejecting is right, but see #10734 for how case handling bites when only one side normalizes.
  • Whether to land generic-format registration first. Dropping the "ICEBERG" check and letting a table carry an arbitrary format plus a location is smaller than this whole design, gets Delta and Parquet catalogued as a side effect, and turns Format: "LANCE" into an ordinary value. The argument against is that it invites tables the maintenance worker cannot service, so it needs a "catalog-only, no maintenance" marker to be honest.