Files
seaweedfs/seaweed-worker/crates/lance/tests/compaction.rs
T
166af06a2b rust: cargo fmt both crates, with a commented-out fmt --check CI step (#11329)
* rust: migrate seaweed-volume and seaweed-worker to tonic 0.14 / prost 0.14

tonic 0.14 boxes the contents of tonic::Status, which is what made every
RPC path trip clippy's result_large_err; the allow for that lint goes in
the next commit. The prost codec moved out of tonic into tonic-prost and
tonic-prost-build, so both build scripts now call
tonic_prost_build::configure() and both crates depend on tonic-prost for
the generated code. The `tls` feature was split into a per-backend
feature; `tls-aws-lc` is the same backend both crates already install
through rustls::crypto::aws_lc_rs.

tonic 0.14 depends on axum 0.8 and tower 0.5, which would have left a
second axum and a second tower in each tree next to the 0.7 / 0.4 the
crates named themselves. Bumping them keeps one copy of each: axum 0.8
only changes the path-parameter syntax for the routes here (`/:vid` ->
`/{vid}`, `/*path` -> `/{*path}`), tower 0.5 needs the `util` feature
named explicitly for ServiceExt::oneshot (it used to arrive through
tonic's feature unification), and tower-http 0.6 is the matching
release.

Lock files move only through cargo's own resolution for the new
versions; no other dependency was refreshed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust: drop the result_large_err allow now that tonic::Status is boxed

tonic 0.14 stores Status behind a Box, so Result<_, Status> is no longer
a large-Err type and clippy has nothing to say about it. Both crates
pass `cargo clippy --all-targets -- -D warnings` without the allow
(seaweed-volume in both feature sets), so the policy entry and its
comment go.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: drop the unused headers argument of try_expand_chunk_manifest

The parameter was already named `_headers`; nothing in the body reads it.
With it gone the function is under clippy's argument threshold and the
expect goes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: pass EC peer reads an EcInterval instead of ten arguments

fetch_one_interval, read_remote_ec_shard_interval,
do_read_remote_ec_shard_interval and recover_one_remote_ec_shard_interval
all took the same (vid, needle_id, shard_id, shard_offset, size,
expected_encode_ts_ns) tuple, and the two that reconstruct also took the
location map with the data/parity counts. Those are now EcInterval (Copy)
and EcShardMap (a borrow of the map plus the counts). The fan-out inside
recovery builds its per-shard request with `EcInterval { shard_id: sid,
..iv }`, which is the one place the old argument list was easy to get
wrong. Bodies destructure at the top, so the code below the signatures
is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: give the EC encoder an EcEncodeLayout and an EncodeRun

encode_dat_file took the Reed-Solomon shape and three block sizes as five
loose integers; they are now one Copy struct, EcEncodeLayout, which is
what Go calls ECContext. The per-row and per-batch helpers took the same
six sinks and the offsets; they become methods on EncodeRun, which owns
the borrows for one run, so each call names only the offset and block
size that vary. The byte-level work is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: describe a .dat rebuild with DatRebuild instead of nine arguments

write_dat_file_from_shards, its _with_dirs twin and the private
write_dat_file were three layers over one nine-argument signature. One
public function now takes a DatRebuild, whose shard_dirs is None when
every shard sits beside the .dat and Some(dirs) for the cross-disk
reconciled layout. The field docs carry what the function doc used to
say about the encode-time size and the block layout.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: split copy_file_from_source's fifteen arguments into two structs

CopyFileSpec is the per-file request (what to ask the source for, where
it lands, whether its bytes count as progress); CopyProgress is the
sender, throttler and report state that all three files of one
VolumeCopy share, held by &mut across the calls. The three production
call sites now read as the .dat/.idx/.vif literals they are, instead of
positional trues and falses.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: create volumes from a VolumeSpec

Volume::new, DiskLocation::create_volume and Store::add_volume each
took the same five-value tail of Go's NewVolume argument list:
collection, replica placement, TTL, preallocation and needle version.
That tail is now VolumeSpec, a Copy struct whose Default is what almost
every test wanted anyway (empty collection, no replication, no TTL, no
preallocation, current version), so most of the 104 call sites shrink
to `&VolumeSpec::default()` or name the one field they set. The id,
directories, index kind and disk type stay positional because they
differ at every site.

Two imports that only test modules use moved into those modules, and
DiskLocation no longer imports ReplicaPlacement.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-worker: run cargo fmt

Layout only; no token in the workspace changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* rust-volume: run cargo fmt

Layout only; no token in the crate changes. Every earlier Rust PR here
formatted only the blocks it touched so as not to drown its diff in
this one, and this commit is that debt paid in a single place. rustfmt
needed two passes to settle one block in handlers.rs; the committed
form is the fixed point, so `cargo fmt --check` is clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

* ci: add a commented-out cargo fmt --check step to both Rust workflows

Same shape as the commented clippy step from #11312: the check is
written out so that making formatting a gate is a one-line uncomment,
and whether to do that stays a maintainer call.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CjZY429aVU74SLDmo1wiuU

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Chris Lu <chrislusf@users.noreply.github.com>
2026-09-15 09:29:22 -07:00

541 lines
19 KiB
Rust

//! Drives the compaction handler against a live namespace.
//!
//! Skipped unless WEED_LANCE_NAMESPACE names one, the way the Go integration
//! tests skip without Docker: compaction rewrites real files, and there is
//! nothing to learn from it against a fake.
use std::collections::HashMap;
use anyhow::Result;
use seaweed_worker_core::pb::{
ExecuteJobRequest, JobSpec, RunDetectionRequest, config_value::Kind,
};
use seaweed_worker_core::{JobHandler, PreviewProvider};
use weed_lance_worker::catalog::NamespaceClient;
use weed_lance_worker::jobs::cleanup::CleanupVersionsHandler;
use weed_lance_worker::jobs::compact::{CompactHandler, JOB_TYPE};
use weed_lance_worker::jobs::indices::OptimizeIndicesHandler;
use weed_lance_worker::preview::LancePreview;
mod common;
use common::{Recorder, fallback, int_config, namespace_url};
/// These tests drive one live gateway and one shared catalog: `list_all_tables`
/// sweeps everything, so a table another test is writing shows up in this test's
/// detection. Rust runs a binary's tests concurrently, so take a lock.
static GATEWAY: tokio::sync::Mutex<()> = tokio::sync::Mutex::const_new(());
/// Declares a table through the namespace and writes `fragments` one-row
/// appends into it, so a test brings its own state instead of depending on
/// whatever a previous run left behind.
async fn seed_fragmented_table(url: &str, name: &str, fragments: usize) -> Result<String> {
seed_table(url, name, fragments, 1, false).await
}
/// Writes `batches` appends of `rows_each` into a freshly declared table, and
/// optionally builds a vector index after the first batch so the later ones are
/// rows no index covers.
async fn seed_table(
url: &str,
name: &str,
batches: usize,
rows_each: usize,
with_index: bool,
) -> Result<String> {
use arrow_array::{
FixedSizeListArray, Float32Array, Int64Array, RecordBatch, RecordBatchIterator,
};
use arrow_schema::{DataType, Field, Schema};
use lance::dataset::{Dataset, WriteMode, WriteParams};
use lance::io::{ObjectStoreParams, StorageOptionsAccessor};
use std::sync::Arc;
// Declaring is the namespace's job, not the worker's, so the test asks for
// it directly rather than widening the client the worker uses. The bucket and
// namespace come first: a table cannot be declared under a parent that does
// not exist, and a test that assumes one is a test that only passes twice.
let http = reqwest::Client::new();
for parent in ["vec", "vec$ml"] {
http.post(format!("{url}/v1/namespace/{parent}/create"))
.json(&serde_json::json!({"mode": "EXIST_OK"}))
.send()
.await?
.error_for_status()?;
}
let encoded = format!("vec$ml${name}");
http.post(format!("{url}/v1/table/{encoded}/declare"))
.json(&serde_json::json!({}))
.send()
.await?
.error_for_status()?;
let client = NamespaceClient::new(url.to_string());
let id = vec!["vec".to_string(), "ml".to_string(), name.to_string()];
let description = client.describe_table(&id).await?;
let mut options = description.storage_options.clone();
options.extend(fallback());
const DIM: i32 = 16;
let schema = Arc::new(Schema::new(vec![
Field::new("id", DataType::Int64, false),
Field::new(
"vec",
DataType::FixedSizeList(Arc::new(Field::new("item", DataType::Float32, true)), DIM),
false,
),
]));
for i in 0..batches {
let ids: Vec<i64> = (0..rows_each).map(|r| (i * rows_each + r) as i64).collect();
let values: Vec<f32> = ids
.iter()
.flat_map(|id| (0..DIM).map(move |d| (*id as f32) + d as f32))
.collect();
let vectors = FixedSizeListArray::new(
Arc::new(Field::new("item", DataType::Float32, true)),
DIM,
Arc::new(Float32Array::from(values)),
None,
);
let batch = RecordBatch::try_new(
schema.clone(),
vec![Arc::new(Int64Array::from(ids)), Arc::new(vectors)],
)?;
let reader = RecordBatchIterator::new(vec![Ok(batch)], schema.clone());
let params = WriteParams {
mode: if i == 0 {
WriteMode::Overwrite
} else {
WriteMode::Append
},
store_params: Some(ObjectStoreParams {
storage_options_accessor: Some(std::sync::Arc::new(
StorageOptionsAccessor::with_static_options(options.clone()),
)),
..Default::default()
}),
..Default::default()
};
let dataset = Dataset::write(reader, description.location.as_str(), Some(params)).await?;
// The index is built after the first batch, so everything appended
// afterwards is a row it does not cover.
if with_index && i == 0 {
use lance::index::DatasetIndexExt;
use lance::index::vector::VectorIndexParams;
use lance_index::IndexType;
use lance_index::vector::{ivf::IvfBuildParams, pq::PQBuildParams};
let mut dataset = dataset;
let params = VectorIndexParams::with_ivf_pq_params(
lance_linalg::distance::MetricType::L2,
IvfBuildParams::new(1),
PQBuildParams::new(4, 8),
);
dataset
.create_index(&["vec"], IndexType::Vector, None, &params, true)
.await?;
}
}
Ok(encoded)
}
/// A table with more fragments than the policy allows is proposed, and running
/// the proposal leaves it with fewer than it started with.
#[tokio::test]
async fn compacts_a_fragmented_table() {
let _gateway = GATEWAY.lock().await;
let Some(url) = namespace_url() else {
eprintln!("WEED_LANCE_NAMESPACE is unset, skipping");
return;
};
// Seeded here rather than by a script, so the test is repeatable: a previous
// run compacts the table it depended on.
let encoded = seed_fragmented_table(&url, "compactme", 12)
.await
.expect("seed a fragmented table");
let handler = CompactHandler::new(url).with_fallback(fallback());
let recorder = Recorder::default();
let request = RunDetectionRequest {
request_id: "detect-1".to_string(),
job_type: JOB_TYPE.to_string(),
worker_config_values: int_config("min_fragments", 4),
..Default::default()
};
handler
.detect(&request, &recorder)
.await
.expect("detection failed");
let proposals = recorder.proposals.lock().unwrap().clone();
assert!(
!proposals.is_empty(),
"expected a proposal for the fragmented table"
);
let proposal = proposals
.iter()
.find(|p| p.summary.contains(encoded.as_str()))
.expect("no proposal for the seeded table");
// Detection opened the dataset to decide, so it reports what it saw. This is
// the only description of a Lance table anything outside the format can give.
let observations = recorder.observations.lock().unwrap().clone();
let observed = observations
.iter()
.find(|o| o.object_id.last().map(String::as_str) == Some("compactme"))
.expect("detection reported no observation for the seeded table");
assert_eq!(observed.format, "LANCE");
for attribute in ["fragments", "rows", "versions", "schema"] {
assert!(
observed.attributes.contains_key(attribute),
"observation is missing {attribute}: {:?}",
observed.attributes.keys().collect::<Vec<_>>()
);
}
let execute = ExecuteJobRequest {
request_id: "execute-1".to_string(),
job: Some(JobSpec {
job_id: "job-1".to_string(),
job_type: JOB_TYPE.to_string(),
parameters: proposal.parameters.clone(),
..Default::default()
}),
worker_config_values: int_config("target_rows_per_fragment", 1_048_576),
..Default::default()
};
handler
.execute(&execute, &recorder)
.await
.expect("execution failed");
let completed = recorder.completed.lock().unwrap().clone();
let result = completed.first().expect("no completion reported");
assert!(
result.success,
"compaction reported failure: {}",
result.error_message
);
let summary = result
.result
.as_ref()
.map(|r| r.summary.clone())
.unwrap_or_default();
assert!(
summary.contains("became"),
"completion carried no fragment counts: {summary}"
);
eprintln!("compaction result: {summary}");
}
/// A table with more versions than the floor is proposed, and running the job
/// reports what it removed. The compaction test above leaves one behind.
#[tokio::test]
async fn cleans_up_old_versions() {
let _gateway = GATEWAY.lock().await;
let Some(url) = namespace_url() else {
eprintln!("WEED_LANCE_NAMESPACE is unset, skipping");
return;
};
let encoded = seed_table(&url, "cleanme", 6, 4, false)
.await
.expect("seed a table with versions to clean");
let handler = CleanupVersionsHandler::new(url).with_fallback(fallback());
let recorder = Recorder::default();
let request = RunDetectionRequest {
request_id: "detect-cleanup".to_string(),
job_type: "lance_cleanup_versions".to_string(),
worker_config_values: int_config("min_versions_to_keep", 2),
..Default::default()
};
handler
.detect(&request, &recorder)
.await
.expect("detection failed");
let proposals = recorder.proposals.lock().unwrap().clone();
let proposal = proposals
.iter()
.find(|p| p.summary.contains(encoded.as_str()))
.cloned()
.expect("no cleanup proposal for the seeded table");
// Retain nothing, so every version outside the current one is fair game and
// the job has something to report rather than a no-op.
let execute = ExecuteJobRequest {
request_id: "execute-cleanup".to_string(),
job: Some(JobSpec {
job_id: "job-cleanup".to_string(),
job_type: "lance_cleanup_versions".to_string(),
parameters: proposal.parameters.clone(),
..Default::default()
}),
worker_config_values: int_config("retain_hours", 0),
..Default::default()
};
handler
.execute(&execute, &recorder)
.await
.expect("cleanup failed");
let completed = recorder.completed.lock().unwrap().clone();
let result = completed.last().expect("no completion reported");
assert!(
result.success,
"cleanup reported failure: {}",
result.error_message
);
eprintln!(
"cleanup result: {}",
result.result.as_ref().unwrap().summary
);
}
/// A table with no indices has nothing to optimize, so detection proposes
/// nothing rather than queueing work that would do nothing.
#[tokio::test]
async fn skips_tables_without_indices() {
let _gateway = GATEWAY.lock().await;
let Some(url) = namespace_url() else {
eprintln!("WEED_LANCE_NAMESPACE is unset, skipping");
return;
};
let encoded = seed_table(&url, "noindex", 2, 8, false)
.await
.expect("seed a table without an index");
let handler = OptimizeIndicesHandler::new(url).with_fallback(fallback());
let recorder = Recorder::default();
let request = RunDetectionRequest {
request_id: "detect-indices".to_string(),
job_type: "lance_optimize_indices".to_string(),
worker_config_values: int_config("max_unindexed_rows", 1),
..Default::default()
};
handler
.detect(&request, &recorder)
.await
.expect("detection failed");
// Judge this table only: the catalog holds every other test's tables too,
// and an indexed one with uncovered rows is supposed to be proposed.
assert!(
!recorder
.proposals
.lock()
.unwrap()
.iter()
.any(|p| p.summary.contains(encoded.as_str())),
"a table with no indices must not be proposed for reindexing"
);
}
/// The job with no Iceberg equivalent: rows appended after an index was built
/// are invisible to a search of it until this runs. Needs a table with an index
/// and rows outside it, which `indexed.py` in the scratchpad seeds.
#[tokio::test]
async fn reindexes_rows_an_index_does_not_cover() {
let _gateway = GATEWAY.lock().await;
let Some(url) = namespace_url() else {
eprintln!("WEED_LANCE_NAMESPACE is unset, skipping");
return;
};
let encoded = seed_table(&url, "reindexme", 2, 512, true)
.await
.expect("seed an indexed table with uncovered rows");
let handler = OptimizeIndicesHandler::new(url).with_fallback(fallback());
let recorder = Recorder::default();
let request = RunDetectionRequest {
request_id: "detect-reindex".to_string(),
job_type: "lance_optimize_indices".to_string(),
worker_config_values: int_config("max_unindexed_rows", 100),
..Default::default()
};
handler
.detect(&request, &recorder)
.await
.expect("detection failed");
let proposals = recorder.proposals.lock().unwrap().clone();
let proposal = proposals
.iter()
.find(|p| p.summary.contains(encoded.as_str()))
.cloned()
.expect("the seeded indexed table was not proposed");
let execute = ExecuteJobRequest {
request_id: "execute-reindex".to_string(),
job: Some(JobSpec {
job_id: "job-reindex".to_string(),
job_type: "lance_optimize_indices".to_string(),
parameters: proposal.parameters.clone(),
..Default::default()
}),
..Default::default()
};
handler
.execute(&execute, &recorder)
.await
.expect("reindex failed");
let completed = recorder.completed.lock().unwrap().clone();
let result = completed.last().expect("no completion reported");
assert!(
result.success,
"reindex reported failure: {}",
result.error_message
);
let output = &result.result.as_ref().unwrap().output_values;
let after = match output
.get("unindexed_rows_after")
.and_then(|v| v.kind.as_ref())
{
Some(Kind::Int64Value(value)) => *value,
other => panic!("no unindexed_rows_after in {other:?}"),
};
assert_eq!(
after, 0,
"rows are still outside the index after optimizing"
);
eprintln!(
"reindex result: {}",
result.result.as_ref().unwrap().summary
);
}
/// The UI's whole reason for asking a worker: admin cannot read a Lance table,
/// so the rows have to come back already rendered.
#[tokio::test]
async fn previews_rows_of_a_table() {
let _gateway = GATEWAY.lock().await;
let Some(url) = namespace_url() else {
eprintln!("set WEED_LANCE_NAMESPACE to run this test");
return;
};
seed_table(&url, "previewme", 2, 3, false)
.await
.expect("seed a table to preview");
let provider = LancePreview::new(url, fallback());
let id = vec!["vec".to_string(), "ml".to_string(), "previewme".to_string()];
let preview = provider.preview(&id, 4).await.expect("preview the table");
assert_eq!(preview.columns, vec!["id".to_string(), "vec".to_string()]);
assert_eq!(preview.total_rows, 6, "total is the table, not the sample");
assert_eq!(preview.rows.len(), 4, "row_limit bounds the sample");
assert!(
preview.rows[0][1].starts_with('['),
"a vector column should render as a list, got {:?}",
preview.rows[0][1]
);
}
/// The claim that removed managed versioning is not "one writer wins the
/// conditional PUT" - that is only the mechanism. It is that concurrent writers
/// lose nothing: the loser sees the conflict, rebases, and commits again. Eight
/// writers appending at once must leave all eight batches in the table.
#[tokio::test]
async fn concurrent_writers_keep_every_commit() {
let _gateway = GATEWAY.lock().await;
let Some(url) = namespace_url() else {
eprintln!("set WEED_LANCE_NAMESPACE to run this test");
return;
};
const WRITERS: i64 = 8;
const ROWS_EACH: i64 = 4;
seed_table(&url, "racers", 1, ROWS_EACH as usize, false)
.await
.expect("seed the table the writers will append to");
let client = NamespaceClient::new(url.clone());
let id = vec!["vec".to_string(), "ml".to_string(), "racers".to_string()];
let description = client.describe_table(&id).await.expect("describe");
let mut options = description.storage_options.clone();
options.extend(fallback());
let writes = (0..WRITERS).map(|writer| {
let location = description.location.clone();
let options = options.clone();
tokio::spawn(
async move { append_rows(&location, &options, writer * 1000, ROWS_EACH).await },
)
});
for (writer, handle) in writes.enumerate() {
handle
.await
.expect("writer panicked")
.unwrap_or_else(|err| panic!("writer {writer} failed to commit: {err:#}"));
}
let table = weed_lance_worker::dataset::open(&client, &id, &fallback())
.await
.expect("reopen the table");
let rows = table.dataset.count_rows(None).await.expect("count rows");
let expected = (ROWS_EACH + WRITERS * ROWS_EACH) as usize;
assert_eq!(
rows, expected,
"concurrent commits lost data: {rows} rows, want {expected}"
);
}
/// Appends one batch to an existing dataset, the way an independent writer would.
async fn append_rows(
location: &str,
options: &HashMap<String, String>,
first_id: i64,
rows: i64,
) -> Result<()> {
use arrow_array::{
FixedSizeListArray, Float32Array, Int64Array, RecordBatch, RecordBatchIterator,
};
use arrow_schema::{DataType, Field, Schema};
use lance::dataset::{Dataset, WriteMode, WriteParams};
use lance::io::{ObjectStoreParams, StorageOptionsAccessor};
use std::sync::Arc;
const DIM: i32 = 16;
let schema = Arc::new(Schema::new(vec![
Field::new("id", DataType::Int64, false),
Field::new(
"vec",
DataType::FixedSizeList(Arc::new(Field::new("item", DataType::Float32, true)), DIM),
false,
),
]));
let ids: Vec<i64> = (0..rows).map(|r| first_id + r).collect();
let values: Vec<f32> = ids
.iter()
.flat_map(|id| (0..DIM).map(move |d| (*id as f32) + d as f32))
.collect();
let vectors = FixedSizeListArray::new(
Arc::new(Field::new("item", DataType::Float32, true)),
DIM,
Arc::new(Float32Array::from(values)),
None,
);
let batch = RecordBatch::try_new(
schema.clone(),
vec![Arc::new(Int64Array::from(ids)), Arc::new(vectors)],
)?;
let params = WriteParams {
mode: WriteMode::Append,
store_params: Some(ObjectStoreParams {
storage_options_accessor: Some(Arc::new(StorageOptionsAccessor::with_static_options(
options.clone(),
))),
..Default::default()
}),
..Default::default()
};
Dataset::write(
RecordBatchIterator::new(vec![Ok(batch)], schema.clone()),
location,
Some(params),
)
.await?;
Ok(())
}