scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-04-21 00:50:35 +00:00

Author	SHA1	Message	Date
Yaniv Michael Kaul	d1b4fd5683	cql3, service, test: add explicit direct includes Add direct includes for result_message.hh and unimplemented.hh at files that use those declarations directly instead of relying on transitive includes.	2026-04-20 12:45:04 +03:00
Piotr Szymaniak	378bcd69e3	tree: add AGENTS.md router and improve AI instruction files Add AGENTS.md as a minimal router that directs AI agents to the relevant instruction files based on what they are editing. Improve the instruction files: - cpp.instructions.md: clarify seastarx.hh scope (headers, not "many files"), explain std::atomic restriction (single-shard model, not "blocking"), scope macros prohibition to new ad-hoc only, add coroutine exception propagation pattern, add invariant checking section preferring throwing_assert() over SCYLLA_ASSERT (issue #7871) - python.instructions.md: demote PEP 8 to fallback after local style, clarify that only wildcard imports are prohibited - copilot-instructions.md: show configure.py defaults to dev mode, add frozen toolchain section, clarify --no-gather-metrics applies to test.py, fix Python test paths to use .py extension, add license header guidance for new files Closes scylladb/scylladb#29023	2026-04-19 21:59:52 +03:00
Dario Mirovic	f77ff28081	test: manager_client: use safe_driver_shutdown for exclusive_clusters Using cluster.shutdown() is an incorrect way to shut down a Cassandra Cluster. The correct way is using safe_driver_shutdown. Fixes SCYLLADB-1434 Closes scylladb/scylladb#29390	2026-04-19 21:31:18 +03:00
Avi Kivity	a15294d601	Revert "Update seastar submodule" This reverts commit `2943d30b0c`. It introduces a regression where --unsafe-bypass-fsync is not honored. Fixes https://scylladb.atlassian.net/browse/SCYLLADB-1496	2026-04-19 15:14:48 +03:00
Avi Kivity	9fb67e3e96	Revert "alternator: optional stripping of http response headers" This reverts commit `73f0deef6d`. It prevents `2943d30b0c`, which causes high flakiness, from being reverted.	2026-04-19 15:14:48 +03:00
Szymon Malewski	73f0deef6d	alternator: optional stripping of http response headers In Alternator's HTTP API, response headers can dominate bandwidth for small payloads. The Server, Date, and Content-Type headers were sent on every response but many clients never use them. This patch introduces three Alternator config options: - alternator_http_response_server_header, - alternator_http_response_disable_date_header, - alternator_http_response_disable_content_type_header, which allow customizing or suppressing the respective HTTP response headers. All three options support live update (no restart needed). The Server header is no longer sent by default; the Date and Content-Type defaults preserve the existing behavior. The Server and Date header suppression uses Seastar's set_server_header() and set_generate_date_header() APIs added in https://github.com/scylladb/seastar/pull/3217. This patch also fixes deprecation warnings from older Seastar HTTP APIs. Tests are in test/alternator/test_http_headers.py. Fixes https://scylladb.atlassian.net/browse/SCYLLADB-70 Closes scylladb/scylladb#28288	2026-04-19 09:22:04 +03:00
Nadav Har'El	f83270df12	Merge 'alternator/streams: Block tablet merges for Alternator Streams on tablet tables' from Piotr Szymaniak DynamoDB Streams API can only convey a single parent per stream shard. Tablet merges produce two parents, making them incompatible with Alternator Streams. This series blocks tablet merges when streams are active on a tablet table. For CreateTable, a freshly created table has no pending merges, so streams are enabled immediately with tablet merges blocked. For UpdateTable on an existing table, stream enablement is deferred: the user's intent is stored via `enable_requested`, tablet merges are blocked (new merge decisions are suppressed and any active merge decision is revoked), and the topology coordinator finalizes enablement once no in-flight merges remain. The topology coordinator is woken promptly on error injection release and tablet split completion, reducing finalization latency from ~60s to seconds. `test_parent_children_merge` is marked xfail (merges are now blocked), and downward (merge) steps are removed from `test_parent_filtering` and `test_get_records_with_alternating_tablets_count`. Not addressed here: using a topology request to preempt long-running operations like repair (tracked in SCYLLADB-1304). Refs SCYLLADB-461 Closes scylladb/scylladb#29224 * github.com:scylladb/scylladb: topology: Wake coordinator promptly for stream enablement lifecycle test/cluster: Test deferred stream enablement on tablet tables alternator/streams: Block tablet merges when Alternator Streams are enabled	2026-04-19 09:15:13 +03:00
Piotr Szymaniak	a2a0868c7d	topology: Wake coordinator promptly for stream enablement lifecycle The topology coordinator sleeps on a condition variable between iterations. Several events relevant to Alternator stream enablement did not wake it, causing delays of up to 60s (the periodic load stats refresh interval) at each step: 1. Error injection release: when a test disables the delay_cdc_stream_finalization injection, the coordinator was not notified. Add an on_disable callback mechanism to the error injection framework (register_on_disable / unregister_on_disable) so subsystems can react when an injection is released. The topology coordinator uses this to broadcast its event. 2. Tablet split completion: after all local storage groups for a table finish splitting, split_ready_seq_number is set but the coordinator only discovered this via the periodic stats refresh. Add an on_tablet_split_ready callback to topology_state_machine that the coordinator sets to trigger_load_stats_refresh(). The split monitor in storage_service calls it when all compaction groups are split-ready, giving the coordinator fresh stats immediately so it can finalize the resize. These changes reduce test_deferred_stream_enablement_on_tablets from ~120s to ~13s and fix a production issue where Alternator stream enablement could be delayed by up to 60s at each step of the lifecycle (error injection release, split completion).	2026-04-19 03:54:33 +02:00
Piotr Szymaniak	a5d35d2b4c	test/cluster: Test deferred stream enablement on tablet tables Async cluster test exercising the deferred enablement lifecycle: ENABLING -> ENABLED -> disabled, verifying tablet merge blocking and unblocking at each stage. Uses delay_cdc_stream_finalization error injection and CQL ALTER TABLE with tablet count constraints. Also adds tablet scheduler config to test_config.yaml (fast refresh interval, scale factor 1) for reliable tablet count changes.	2026-04-19 03:54:33 +02:00
Piotr Szymaniak	4b6937b570	alternator/streams: Block tablet merges when Alternator Streams are enabled DynamoDB Streams API can only convey a single parent per stream shard. Tablet merges produce 2 parents, which is incompatible. When streams are requested on a tablet table, block tablet merges via tablet_merge_blocked (the allocator suppresses new merge decisions and revokes any active merge decision). add_stream_options() sets tablet_merge_blocked=true alongside enabled=true, so CreateTable needs no special handling — the flag is inert on vnode tables and immediately effective on tablet tables. For UpdateTable, CDC enablement is deferred: store the user's intent via enable_requested, and let the topology coordinator finalize enablement once no in-progress merges remain. A new helper, defer_enabling_streams_block_tablet_merges(), amends the CDC options to this deferred state. Disabling streams clears all flags, immediately re-allowing merges. The tablet allocator accesses the merge-blocked flag through a schema::tablet_merges_forbidden() accessor rather than reaching into CDC options directly. Mark test_parent_children_merge as xfail and remove downward (merge) steps from tablet_multipliers in test_parent_filtering and test_get_records_with_alternating_tablets_count.	2026-04-19 03:54:33 +02:00
Avi Kivity	f5886b4fdd	Merge 'Add virtual task for vnodes-to-tablets migrations' from Nikos Dragazis This PR exposes vnodes-to-tablets migrations through the task manager API via a virtual task. This allows users to list, query status, and wait on ongoing migrations through a standard interface, consistent with other global operations such as tablet operations and topology requests are already exposed. The virtual task exposes all migrations that are currently in progress. Each migrating keyspace appears as a separate task, identified by a deterministic name-based (v3) UUID derived from the keyspace name. Progress is reported as the number of nodes that have switched to tablets vs. the total. The number increases on the forward path and decreases on rollback. The task is not abortable - rolling back a migration requires a manual procedure. The `wait` API blocks until the migration either completes (returning `done`) or is rolled back (returning `suspended`). Example output: ``` $ scylla nodetool tasks list vnodes_to_tablets_migration task_id type kind scope state sequence_number keyspace table entity shard start_time end_time 1747b573-6cd6-312d-abb1-9b66c1c2d81f vnodes_to_tablets_migration cluster keyspace running 0 ks 0 $ scylla nodetool tasks status 1747b573-6cd6-312d-abb1-9b66c1c2d81f id: 1747b573-6cd6-312d-abb1-9b66c1c2d81f type: vnodes_to_tablets_migration kind: cluster scope: keyspace state: running is_abortable: false start_time: end_time: error: parent_id: none sequence_number: 0 shard: 0 keyspace: ks table: entity: progress_units: nodes progress_total: 3 progress_completed: 0 ``` Fixes SCYLLADB-1150. New feature, no backport needed. Closes scylladb/scylladb#29256 * github.com:scylladb/scylladb: test: cluster: Verify vnodes-to-tablets migration virtual task distributed_loader: Link resharding tasks to migration virtual task distributed_loader: Make table_populator aware of migration rollbacks service: Add virtual task for vnodes-to-tablets migrations storage_service: Guard migration status against uninitialized group0 compaction: Add parent_id to table_resharding_compaction_task_impl storage_service: Add keyspace-level migration status function storage_service: Replace migration status string with enum utils: Add UUID::is_name_based()	2026-04-19 00:56:33 +03:00
Nadav Har'El	2943d30b0c	Update seastar submodule * seastar 4d268e0e...22a5aa13 (36): > apps/httpd: replace deprecated reply::done() with write_body() > missing header(s) > net: Fix missing throw for runtime_error in create_native_net_device > tests/io_queue: account for token bucket refill granularity in bandwidth checks > Merge 'iovec: fix iovec_trim_front infinite loop on zero-length iovecs' from Travis Downs tests: add regression tests for zero-length iovec handling iovec: fix iovec_trim_front infinite loop on zero-length iovecs > util/process: graduate process management API from experimental > cooking: don't register ready.txt as a build output > sstring: make make_sstring not static > Add SparkyLinux to debian list in install-dependencies.sh > http: allow control over default response headers > Merge 'chunked_fifo: make cached chunk retention configurable' from Brandon Allard tests/perf: add chunked_fifo microbenchmarks chunked_fifo: set the default free chunk retention to 0 chunked_fifo: make free chunk retention configurable > Merge 'reactor_backend: fix pollable_fd_state_completion reuse in io_uring' from Kefu Chai tests: add regression test for pollable_fd_state_completion reuse reactor_backend: use reset() in AIO and epoll poll paths reactor_backend: fix pollable_fd_state_completion reuse after co_await in io_uring > Merge 'coroutine: Generator cleanups' from Kefu Chai coroutine/generator: extract schedule_or_resume helper coroutine/generator: remove unused next_awaiter classes coroutine/generator: remove write-only _started field coroutine/generator: assert on unreachable path in buffered await_resume coroutine/generator: add elements_of tag and #include <ranges> coroutine/generator: add empty() to bounded_container concept > cmake: bump minimum Boost version to 1.79.0 > seastar_test: remove unnecessary headers > cmake: bump minimum GnuTLS version to 3.7.4 > Merge 'reactor: add get_all_io_queues() method' from Travis Downs tests: add unit test for reactor::get_all_io_queues() reactor: add get_all_io_queues() method reactor: move get_io_queue and try_get_io_queue to .cc file > http: deprecate reply::done(), remove _response_line dead field > core: Deprecate scattered_message > ci: add workflow dispatch to tests workflow > perf_tests: exit non-zero when -t pattern matches no tests > Replace duplicate SEGV_MAPERR check in sigsegv_action() with SEGV_ACCERR. > perf_tests: add total runtime to json output > Merge 'Relax large allocation error originating from json_list_template' from Robert Bindar implement move assignment operator for json_list_template json_list_template copy assignment operator reserves capacity upfront > perf_tests: add --no-perf-counters option > Merge 'Fix to_human_readable_value() ability to work with large values' from Pavel Emelyanov memory: Add compile-time test for value-to-human-readable conversion memory: Extend list of suffixes to have peta-s memory: Fix off-by-one in suffix calculation memory: Mark to_human_readable_value() and others constexpr > http: Improve writing of response_line() into the output > Merge 'websocket: add template parameter for text/binary frame mode and implement client-side WebSocket' from wangyuwei websocket: add template parameter for text/binary frame mode websocket: impl client side websocket function > file: Fix checks for file being read-only > reactor: Make do_dump_task_queue a task_queue method > Merge 'Implement fully mixed mode for output_stream-s' from Pavel Emelyanov tests/output_stream: sample type patterns in sanitizer builds tests/output_stream: extend invariant test to cover mixed write modes iostream: allow unrestricted mixing of buffered and zero-copy writes tests/output_stream: remove obsolete ad-hoc splitting tests tests/output_stream: add invariant-based splitting tests iostream: rename output_stream::_size to ::_buffer_size > reactor_backend: replace virtual bool methods with const bool_class members > resource: Avoid copying CPU vector to break it into groups > perf_tests: increase overhead column precision to 3 decimal places > Merge 'Move reactor::fdatasync() into posix_file_impl' from Pavel Emelyanov reactor: Deprecate fdatasync() method file: Do fdatasync() right in the posix_file_impl::flush() file: Propagate aio_fdatasync to posix_file_impl reactor: Move reactor::fdatasync() code to file.cc reactor,file: Make full use of file_open_options::durable bit file: Add file_open_options::durable boolean file: Account io_stats::fsyncs in posix_file_impl::flush() reactor: Move _fsyncs counter onto io_stats > http: Remove connection::write_body()	2026-04-18 11:52:33 +03:00
Nadav Har'El	31e0315710	Merge 'alternator: fix unnecesary cdc log entries' from Radosław Cybulski Fix cdc writing unnecesary entries to it's log, like for example when Alternator deletes an item which in reality doesn't exist. Originally @wps0 tackled this issue. This patch is an extension of his work. His work involved adding `should_skip` function to cdc, which would process a `mutation` object and decide, wherever changes in the object should be added to cdc log or not. The issue with his approach is that `mutation` object might contain changes for more than one row. If - for example - the `mutation` object contains two changes, delete of non-existing row and create of non-existing row, `should_skip` function will detect changes in second item and allow whole `mutation` (BOTH items) to be added. For example (using python's boto3) running this on empty table: ``` with table.batch_writer() as batch: batch.put_item({'p': 'p', 'c': 'c0'}) batch.delete_item(Key={'p': 'p', 'c': 'c1'}) ``` will emit two events ("put" event and "delete" event), even though the item with `c` set to `c1` does not exist (thus can't be deleted). Note, that both entries in batch write must use the same partition key, otherwise upper layer with split them into separate `mutation` objects and the issue will not happen. The solution is to do similar processing, but consider each change separated from others. This is tricky to implement due to a way cdc works. When cdc processes `mutation` object (containing X changes), it emits cdc entries in phases. Phase 1 - emit `preimage` (old state) for each change (if requested). Phase 2 - for each change emit actual "diff" (update / delete and so on). Phase 3 - emit `postimage` (new state). We will know if change needs to be skipped during phase 2. By that time phase 1 is completed and preimage for the change is emited. At that moment we set a flag that the change (identified by clustering key value) needs to be skipped - we add a clustering key to a `ignore-rows` set (`_alternator_clustering_keys_to_ignore` variable) and continue normally. Once all phases finish we add a `postprocess` phase (`clean_up_noop_rows` function). It will go through generated cdc mutations and skip all modifications, for which clustering key is in `ignore-rows` set. After skipping we need to do a "cleanup" operation - each generated cdc mutation contain index (incremented by one), if we skipped some parts, the index is not consecutive anymore, so we reindex final changes. There's a special case worth mentioning - Alternator tables without clustering keys. At that point `mutation` object passed to cdc can contain exactly one change (since different partition keys are splitted by upper layers and Alternator will never emit `mutation` object containing two (or more) changes with the same primary key. Here, when we decide the change is to be skipped we add empty `bytes` object to `ignore-rows` set. When checking `ignore-rows` set, we check if it's empty or not (we don't check for presence of empty `bytes` object). Note: there might be some confusion between this patch and #28452 patch. Both started from the same error observation and use similar tests for validation, as both are easily triggered by BatchWrite commands (both needs `mutation` object passed to cdc to contain more than one single change). This issue tho is about wrong data written in cdc log and is fixed at cdc, where #28452 is about wrong way of parsing correct cdc data and is fixed at Alternator side of things. Note, that we need #28452 to truly verify (otherwise we will emit correct cdc entries, but Alternator will incorrectly parse them). Note: to benefit / notice this patch you need `alternator_streams_increased_compatibility` flag turned on. Note: rework is quite "broad" and covers a lot of ground - every operation, that might result in a no-change to the database state should be tested. An additional test was added - trying to remove a column from non-existing item, as well as trying to remove non-existing column from existing item. Fixes: #28368 Fixes: SCYLLADB-1528 Fixes: SCYLLADB-538 Closes scylladb/scylladb#28544 * github.com:scylladb/scylladb: alternator: remove unnecesary code alternator: fix Alternator writing unnecesary cdc entries alternator: add failing tests for Streams	2026-04-18 00:07:51 +03:00
Nadav Har'El	32060d73df	Merge 'alternator: Add stream support for tablets' from Radosław Cybulski Implements neccesary changes for Streams to work with tablet based tables. - add utility functions to `system_keyspace` that helps reading cdc content from cdc log tables for tablet based base tables (similar api to ones for vnodes) - remove antitablet `if` checks, update tests that fail / skip if tablets are selected - add two tests to extensively test tablet based version, especially while manipulating stream count Fixes #23838 Fixes SCYLLADB-463 Closes scylladb/scylladb#28500 * github.com:scylladb/scylladb: alternator: add streams with tablets tests alternator: remove antitablet guards when using Streams alternator: implement streams for tablets treewide: add cdc helper functions to system_keyspace alternator: add system_keyspace reference	2026-04-17 23:48:31 +03:00
Radosław Cybulski	586bb1d345	alternator: fix issues with stream_arn copy / move `stream_arn` object holds a full ARN as `std::string` and two `std::string_view` fields (`table_name_` and `keyspace_name_`) pointing into ARN itself. This prevents object from being safely copied (as in that case both `table_name_` and `keyspace_name_` will point into original object's ARN). Similar issue might happen with move, when ARN contains string short enough for small string optimization to kick in (although in practice this is not possible, as ARN has requirements which make it's minimal length above 15 characteres - current limit for small string optimizations in most popular string libraries). The patch drops `std::string_view` objects in favor of integer offsets and sizes. The offset equal to 0 means beginning of ARN string. The api is preserved - both `table_name` and `keyspace_name` function will return `std::string_view` reconstructed on the fly. Closes scylladb/scylladb#29507	2026-04-17 23:13:17 +03:00
Piotr Szymaniak	caaef45b7a	audit: restore static_cast for batch inspect Closes scylladb/scylladb#29545	2026-04-17 23:11:18 +03:00
Nikos Dragazis	d361a0dd83	test: cluster: Verify vnodes-to-tablets migration virtual task Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 21:13:52 +03:00
Nikos Dragazis	295e434781	distributed_loader: Link resharding tasks to migration virtual task When a table is loaded on startup during a vnodes-to-tablets migration (forward or rollback), the `table_populator` runs a resharding compaction. Set the migration virtual task as parent of the resharding task. This enables users to easily find all node-local resharding tasks related to a particular migration. Make `migration_virtual_task::make_task_id()` public so that the `distributed_loader` can compute the migration's task ID. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	a3aa4f6cb4	distributed_loader: Make table_populator aware of migration rollbacks The `table_populator` uses a `migrate_to_tablets` flag to distinguish normal tables from tables under vnodes-to-tablets migration (forward path), since the two require different resharding. The next patch will set the parent info of migration-related resharding compaction tasks so they appear as children of the migration virtual task. For that, the table populator needs to recognize not only migrations in the forward path, but rollbacks as well. Replace the flag with a tri-state `migration_direction` enum (none, forward, rollback). Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	696f9f8954	service: Add virtual task for vnodes-to-tablets migrations Add a virtual task that exposes in-progress vnodes-to-tablets migrations through the task manager API. The task is synthesized from the current migration state, so completed migrations are not shown. Progress is reported as the number of nodes that currently use tablets: it increases on the forward path and decreases on rollback. For simplicity, per-node storage modes are not exposed in the task status; callers that need them should use the migration status REST endpoint. Unlike regular tasks that use time-based UUIDs, this task uses deterministic named UUIDs derived from the keyspace names. This keeps the implementation simple (no need to persist them) and gives each keyspace a stable task ID. The downside is that the start time of each task is unknown and repeated migrations of the same keyspace (migration -> rollback -> new migration) cannot be distinguished. Introduce a new task manager module to keep them separate from other tasks. Add support for `wait()`. While its practical value is debatable (migration is a manual procedure, rolling restart will interrupt it), it keeps the task consistent with the task manager interface. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	d1ca01b25d	storage_service: Guard migration status against uninitialized group0 `storage_service::get_tablets_migration_status()` reads a group0 virtual table, so it requires group0 to be initialized. When invoked via the migration REST API, this condition is satisfied since the API is only available after joining group0. However, once this function is integrated into the task API later in this series, the assumption will no longer hold, as the task API is exposed earlier in the startup process. Add a guard to detect this condition and return a clear error message. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	ca830c7bce	compaction: Add parent_id to table_resharding_compaction_task_impl Required to link it with the migration task in the next patches. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	46e3902daa	storage_service: Add keyspace-level migration status function `storage_service::get_tablets_migration_status()` returns the keyspace-level migration status, indicating whether migration has not started, is in progress, or has completed, and for migrating keyspaces also returns per-node migration statuses. Rename it to `get_tablets_migration_status_with_node_details()` and introduce a new `get_tablets_migration_status()` that returns only the keyspace-level status. This prepares the function for reuse in the next patches, which will add a virtual task for vnodes-to-tablets migrations. Several task-manager paths will only need the keyspace-level migration state, not per-node information. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	3096ba0577	storage_service: Replace migration status string with enum Using a string was sufficient while this status was only exposed through the REST API, but the next patches will also consume it internally. Use an enum for the internal representation and convert it back to the existing string values in the REST API. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:59:05 +03:00
Nikos Dragazis	a00056381f	utils: Add UUID::is_name_based() The UUID class already provides `is_timestamp()` for identifying time-based (version 1) UUIDs. Add the analogous `is_name_based()` predicate for version 3 (name-based) UUIDs, along with a test. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com>	2026-04-17 20:58:39 +03:00
Radosław Cybulski	9a6aed721b	alternator: add streams with tablets tests Add tests for Streams, when table uses tablets underneath. One test verifies filtering using CHILD_SHARDS feature. Other one makes sure we get read all data while the table undergoes tablet count change. Add `--tablet-load-stats-refresh-interval-in-seconds=1` to `alternator/run` script, as otherwise newly added tests will fail. The setting changes how often scylla refreshes tablet metadata. This can't be done using `scylla_config_temporary`, as 1) default is 60 seconds 2) scylla will wait full timeout (60s) to read configuration variable again.	2026-04-17 18:58:27 +02:00
Radosław Cybulski	6be16cf224	alternator: remove antitablet guards when using Streams Remove `if` condition, that prevented tables with tablets working with Streams. Remove a test, that verifies, that Alternator will reject tables with tablets underneath working with Streams feature enabled on them. Update few tests, that were expected to fail on tablets to enable their normal execution.	2026-04-17 18:58:26 +02:00
Radosław Cybulski	d5df3ec07c	alternator: implement streams for tablets Add a code, that will handle Streams reading, when table is using tablets underneath. Fixes #23838	2026-04-17 18:57:44 +02:00
Radosław Cybulski	eb35a7b6ce	treewide: add cdc helper functions to system_keyspace Add helper functions to `system_keyspace` object, that deal with reading cdc content for tablet based table's. `read_cdc_for_tablets_current_generation_timestamp` will read current generation's timestamp. `read_cdc_for_tablets_versioned_streams` will build timestamp -> `cdc::streams_version` map similar to how `system_distributed_keyspace::cdc_get_versioned_streams` works. We're adding those helper functions, because their siblings in `system_distributed_keyspace` work only, when base table is backed up by vnodes. New additions work only, when base table is backed up by tablets.	2026-04-17 18:57:44 +02:00
Radosław Cybulski	d93299b605	alternator: add system_keyspace reference Add a reference to `system_keyspace` object to `executor` object in alternator. The reference is needed, because in future commit we will add there (and use) helper functions that read `cdc_log` tables for tablet based tables similarly to already existing siblings for vnodes living in `system_distributed_keyspace`.	2026-04-17 18:57:43 +02:00
Radosław Cybulski	04b9d3875f	alternator: remove unnecesary code After our fix, that prevents no-op changes being written into cdc log we will remove Piotr Wieczorek's previous attempt, which is now unnecesary.	2026-04-17 18:02:00 +02:00
Radosław Cybulski	6e5aaa85b6	alternator: fix Alternator writing unnecesary cdc entries Work in this patch is a result of two bugs - spurious MODIFY event, when remove column is used in `update_item` on non-existing item and spurious events, when batch write item mixed noop operations with operations involving actual changes (the former would still emit cdc log entries). The latter issue required rework of Piotr Wieczorek's algorithm, which fixed former issue as well. Piotr Wieczorek previously wrote checks, that should prevent unnecesary cdc events from being written. His implementation missed the fact, that a single `mutation` object passed to cdc code to be analysed for cdc log entries can contain modifications for multiple rows (with the same timestamp - for example as a result to BatchWriteItem call). His code tries to skip whole `mutation`, which in such case is not possible, because BatchWriteItem might have one item that does nothing and second item that does modification (this is the reason for the second bug). His algorithm was extended and moved. Originally it was working as follows - user would sent a `mutation` object with some changes to be "augmented". The cdc would process those changes and built a set of cdc log changes based on them, that would be added to cdc log table. Piotr added a `should_skip` function, which processes user changes and tried to determine if they all should be dropped or not. New version, instead of trying to skip adding rows to cdc log `mutation` object, builds a rows-to-ignore set. After whole cdc log `mutation` object is completed, it processes it and go through it row by row. Any row that was previously added to a `rows_to_ignore` set will now be removed. Remaining rows are written to new cdc log `mutation` with new clustering key (`cdc$batch_seq_no` index value should probably be consecutive - we just want to be safe here) and returns new `mutation` object to be sent to cdc log table. The first bug is fixed as a side effect of new algorithm, which contains more precise checks detecting, if given mutation actually made a difference. Fixes: #28368 Fixes: SCYLLADB-538 Fixes: SCYLLADB-1528 Refs: #28452	2026-04-17 18:00:25 +02:00
Botond Dénes	6ce0968960	compaction: release GC'ed sstables incrementally during compaction Garbage collected sstables created during incremental compaction are deleted only at the end of the compaction, which increases the memory footprint. This is inefficient, especially considering that the related input sstables are released regularly during compaction. This commit implements incremental release of GC sstables after each output sstable is sealed. Unlike regular input sstables, GC sstables use a different exhaustion predicate: a GC sstable is only released when its token range no longer overlaps with any remaining input sstable. This is because GC sstables hold tombstones that may shadow data in still-alive overlapping input sstables; releasing them prematurely would cause data resurrection. Fixes #5563 Closes scylladb/scylladb#28984	2026-04-17 18:20:47 +03:00
Radosław Cybulski	2894542e57	alternator: add failing tests for Streams Add failing tests for Streams functionality. Trying to remove column from non-existing item is producing a MODIFY event (while it should none). Doing batch write with operations working on the same partition, where one operation is without side effects and second with will produce events for both operations, even though first changes nothing. First test has two versions - with and without clustering key. Second has only with clustering key, as we can't produce batch write with two items for the same partition - batch write can't use primary key more than once in single call. We also add a test for batch write, where one of three operations has no observable side effects and should not show up in Streams output, but in current scylla's version it does show.	2026-04-17 16:28:14 +02:00
Botond Dénes	6eb2d15f39	Merge 'Replace CAS estimated histogram with estimated_histogram_with_max' from Amnon Heiman ScyllaDB uses estimated_histogram in many places. We already have a more efficient alternative: estimated_histogram_with_max. It is both CPU- and memory-efficient, and it can be exported as Prometheus native histograms. Its main limitation (which also has benefits) is that the bucket layout is fixed at compile time, so histograms with different configurations cannot be mixed. The end goal is to replace all uses of estimated_histogram in the codebase. That migration requires a few small API adjustments, so it is done in steps. This PR replaces estimated_histogram for CAS contention. The PR includes a patch that adds functionality to the base approx_exponential_histogram, which will be used by the API. The specific histograms are defined in a single place and cover the range 1-100; this makes future changes easy. New feature, no need to backport Closes scylladb/scylladb#29017 * github.com:scylladb/scylladb: storage_proxy: migrate CAS contention histograms to estimated_histogram_with_max estimated_histogram.hh: Add bucket offset and count to approx_exponential_histogram	2026-04-17 13:12:59 +03:00
Andrzej Jackowski	e256d9f69d	test: retry get_coordinator_host() after topology coordinator stop After stopping the topology coordinator, a new topology coordinator may not yet be started when get_coordinator_host() is called. Make the function always retry via wait_for so that every caller is protected against this race. Fixes SCYLLADB-1553 Closes scylladb/scylladb#29489	2026-04-17 12:08:26 +02:00
Botond Dénes	fbcfe3f88f	test: use uuid4 for DockerizedServer container names to avoid collisions Container names were generated as {name}-{pid}-{counter}, where the counter is a per-process itertools.count. This scheme breaks across CI runs on the same host: if a prior job was killed abruptly (SIGKILL, cancellation) its containers are left running since --rm only removes containers on exit. A subsequent run whose worker inherits the same PID (common in containerized CI with small PID namespaces) and reaches the same counter value will collide with the orphaned container. Replace pid+counter with uuid.uuid4(), which generates a random UUID, making names unique across processes, hosts, and time without any shared state or leaking host identifiers. Fixes: SCYLLADB-1540 Closes scylladb/scylladb#29509	2026-04-17 11:56:51 +02:00
Botond Dénes	57f8be49e9	Merge 'Move ignore_component_digest_mismatch flag on sstables_manager' from Pavel Emelyanov The PR serves two purposes. First, it makes the flag usage be consistent across multiple ways to load sstables components. For example, the sstable::load_metadata() doesn't set it (like .load() does) thus potentially refusing to load "corrupted" components, as the flag assumes. Second, it removes the fanout of db.get_config().ignore_component_digest_mismatch() over the code. This thing is called pretty much everywhere to initialize the sstable_open_config, while the option in question is "scylla state" parameter, not "sstable opening" one. Code cleanup, not backporting Closes scylladb/scylladb#29513 * github.com:scylladb/scylladb: sstables: Remove ignore_component_digest_mismatch from sstable_open_config sstables: Move ignore_component_digest_mismatch initialization to constructor sstables: Add ignore_component_digest_mismatch to sstables_manager config	2026-04-17 12:54:17 +03:00
Avi Kivity	cad3c0de94	test: write minio log to testlog dir for Jenkins artifact collection Write the MinIO server log directly to tempdir_base (testlog/<arch>/) instead of the per-server temp directory that gets destroyed on shutdown. This preserves the log for Jenkins artifact collection, helping debug S3-related flaky test failures like the stcs_reshape_overlapping_s3_test hang (SCYLLADB-1481). Closes scylladb/scylladb#29458	2026-04-17 12:51:55 +03:00
Botond Dénes	facb50cbf9	Merge 'test.py: refactor test.py' from Andrei Chekun With the latest changes, there are a lot of code that is redundant in the test.py. This PR just cleans this code. Also, it narrows using dynamic scope for fixtures to test/alternator and test/cqlpy. All the rest by default will have module scope. test.py will be a wrapper for pytest mostly for CI use. As for now test.py have important part of calculating the number of threads to start pytest with. This is not possible to do in pytest itself. No backport needed, framework enhancement only. Fixes: https://scylladb.atlassian.net/browse/SCYLLADB-666 Closes scylladb/scylladb#28852 * github.com:scylladb/scylladb: test.py: remove testpy_test_fixture_scope test.py: add logger for 3rd party service test.py: delete dead code in test.py	2026-04-17 12:51:14 +03:00
Pawel Pery	7883f161bb	vector-store: fix creating local vector search indexes with a part of the partition key Users ought to have possibility to create the local index for Vector Search based only on a part of the partition key. This commits provides this by removing requirements of 'full partition key only' for custom local index. The commit updates docs to explain that local vector index can use only a part of the partition key. The commit implements cqlpy test to check fixed functionality. Fixes: SCYLLADB-953 Needs to be backported to 2026.1 as it is a fix for local vector indexes. Closes scylladb/scylladb#28931	2026-04-17 11:44:15 +02:00
Karol Nowacki	c643f321af	vector_search: decrease default connection timeout to 3s Decrease the default connection timeout to 3s to better align with the default CQL query timeout of 10s. The previous timeout allowed only one failover request in high availability scenario before hitting the CQL query timeout. By decreasing the timeout to 3s, we can perform up to three failover requests within the CQL query timeout, which significantly improves the chances of successfully completing the query in high availability scenarios. Fixes: SCYLLADB-95	2026-04-17 12:26:39 +03:00
Karol Nowacki	9269ca9cf7	vector_search: add unreachable node detection time config Add option `vector_store_unreachable_node_detection_time_in_ms` to control parameters related to detecting unreachable vector store nodes. This parameter is used to set the TCP connect timeout, keepalive parameters, and TCP_USER_TIMEOUT. By configuring these parameters, we can detect unreachable vector store nodes faster and trigger failover mechanisms in a timely manner.	2026-04-17 12:26:38 +03:00
Piotr Smaron	686029f52c	audit: disable caching for the audit log table The audit table had caching enabled by default, which provides no value since audit data is write-heavy and rarely read back through the cache. This wastes cache space that could be used for more important user data. Disable caching by setting keys and rows_per_partition to NONE and enabled to false, consistent with get_disabled_caching_options() and other system tables such as system.batchlog, system.large_partitions, and CDC log tables. Closes scylladb/scylladb#29506	2026-04-17 11:17:10 +02:00
Piotr Dulikowski	37fc1507f0	Merge 'Alternator: Add vector search support' from Nadav Har'El This series adds support for vector search in Alternator based on the existing implementation in CQL. The series adds APIs for `CreateTable` and `UpdateTable` to add or remove vector indexes to Alternator tables, `DescribeTable` to list them and check the indexing status, and `Query` to perform a vector search - which contacts the vector store for the actual ANN (approximate nearest neighbor) search. Correct functionality of these features depend on some features of the the vector store, that were already done (see https://github.com/scylladb/vector-store/pull/394). This initial implementation is fully functional, and can already be useful, but we do not yet support all the features we hope to eventually support. Here are things that we have not done yet, and plan to do later in follow-up pull requests: 1. Support a new optimized vector type ("V") - in addition to the "list of numbers" type supported in this version. 2. Allow choosing a different similarity function when creating an index, by SimilarityFunction in VectorIndex definition. 3. Allow choosing quantization (f32/f16/bf16/i8/b1) to ask the vector index to compress stored vectors. 4. Support oversampling and rescoring, defined per-index and per-query. 5. Support HNSW tuning parameters — maximum_node_connections, construction_beam_width, search_beam_width. 6. Support pre-filtering over key columns, which are available at the vector store, by sending the filter to the vector store (translated from DynamoDB filter syntax to the vector's store's filter syntax). A decision still need to be made if this will use KeyConditionExpression or FilterExpression. This version supports only post-filtering (with `FilterExpression`). 7. Support projecting non-key attributes into the index (Projection=INCLUDE and Projection=ALL), and then 1. pre-filtering using these attributes, and 2. efficiently return these attributes (using Select=ALL_PROJECTED_ATTRIBUTES, which today returns just the key columns). 8. Optimize the performance of `Query`, which today is inefficient for Select=ALL_ATTRIBUTES because it serially retrieves the matching items one at a time. 9. Returning the similarity scores with the items (the design proposes ReturnVectorSearchSimilarity). 10. Add more vector-search-specific metrics, beyond the metric we already have counting Query requests. For example separate latency and request-count metrics for vector-search Queries (distinct from GSI/LSI queries), and a metric accumulating the total Limit (K) across all vector search queries. 11. Consider how (and if at all) we want to run the tests in test/alternator/test_vector.py that need the vector store in the CI. Currently they are skipped in CI and only run manually (with `test/alternator/run --vs test_vector`). 12. UpdateTable 'Update' operation to modify index parameters. Only some can be modified, e.g., Oversampling. 13. Support for "local index" (separate index for each partition). 14. Make sure that vector search and Streams can be enabled concurrently on the same table - both need CDC but we need to verify that one doesn't confuse the other or disables options that the other needs. We can only do this after we have Alternator Streams running on tablets (since vector store requires tablets). Testing the new Alternator vector search end-to-end requires running both Scylla and the vector store together. We will have such end-to-end tests in the vector store repository (see https://github.com/scylladb/vector-store/pull/392), but we also add in this pull request many end-to-end tests written in Python, that can be run with the command "test/alternator/run --vs test_vector.py". The "--vs" option tells the run script to run both Scylla and the vector store (currently assumed to be in `.../vector-store/target/release/vector-store`). About 65% of the tests in this pull request check supported syntax and error paths so can run without the vector store, while about 35% of the tests do perform actual Query operations and require the vector store to be running. Currently, the tests that do require the vector store will not get run by CI, but can be easily re-run manually with `test/alternator/run --vs test_vector.py`. In total, this series includes 78 functional tests in 2200 lines of Python code. This series also includes documentation for the new Alternator feature and the new APIs introduced. You can see a more detailed design document here: https://docs.google.com/document/d/1cxLI7n-AgV5hhH1DTyU_Es8_f-t8Acql-1f58eQjZLY/edit Two patches in this series split the huge alternator/executor.cc, after this series continued to grow it and it reached a whoppng 7,000 lines. These patches are just reorganization of code, no functional changes. But it's time that we finally do this (Refs #5783), we can't just continue to grow executor.cc with no end... Closes scylladb/scylladb#29046 * github.com:scylladb/scylladb: test/alternator: add option to "run" script to run with vector search alternator: document vector search test/alternator: fix retries in new_dynamodb_session test/alternator: test for allowed characters in attribute names test/alternator: tests for vector index support alternator, vector: add validation of non-finite numbers in Query alternator: Query: improve error message when VectorSearch is missing alternator: add per-table metrics for vector query alternator: clean up duplicated code alternator: fix default Select of Query alternator: split executor.cc even more alternator: split alternator/executor.cc alternator: validate vector index attribute values on write alternator: DescribeTable for vector index: add IndexStatus and Backfilling alternator: implement Query with a vector index alternator: fix bug in describe_multi_item() alternator: prevent adding GSI conflicting with a vector index alternator: implement UpdateTable with a vector index alternator: implement DescribeTable with a vector index alternator: implement CreateTable with a vector index alternator: reject empty attribute names cdc: fix on_pre_create_column_families to create CDC log for vector search	2026-04-17 10:25:45 +02:00
Avi Kivity	04b54f363b	Merge 'Enable vnodes-to-tablets migrations with arbitrary tokens' from Nikos Dragazis This PR removes the power-of-two token constraint from vnodes-to-tablets migrations, allowing clusters with randomly generated tokens to migrate without manual token reassignment. Previously, migrations required vnode tokens to be a power of two and aligned. In practice, these conditions are not met with Scylla's default random token assignment, so the constraint is a blocker for real-world use. With the introduction of arbitrary tablet boundaries in PR #28459, the tablet layer can now support arbitrary tablet boundaries. This PR builds on that capability to allow arbitrary vnode tokens during migration. When the highest vnode token does not coincide with the end of the token ring, the vnode wraps around, but tablets do not support that. This is handled by splitting it into two tablets: one covering the tail end of the ring and one covering the beginning. Testing has been updated accordingly: existing cluster tests now use randomly generated tokens instead of precomputed power-of-two values, and a new Boost test validates the wrap-around tablet boundary logic. Fixes SCYLLADB-724. New feature, no backport is needed. Closes scylladb/scylladb#29319 * github.com:scylladb/scylladb: test: Use arbitrary tokens in vnodes->tablets migration tests test: boost: Add test for wrap-around vnodes storage_service: Support vnodes->tablets migrations w/ arbitrary tokens storage_service: Hoist migration precondition	2026-04-17 00:46:35 +03:00
Andrei Chekun	745debe9ec	test.py: remove testpy_test_fixture_scope With migration to pyest this fixture is useless. Removing and setting the session to the module for the most of the tests. Add dynamic_scope function to support running alternator fixtures in session scope, while Test and TestSuite are not deleted. This is for migration period, later on this function should be deleted.	2026-04-16 22:08:33 +02:00
Andrei Chekun	21addb2173	test.py: add logger for 3rd party service With migration of preparation environment and starting 3rd party services to the pytest, they're output the logs to the terminal. So this PR binds them their own log file to avoid polluting the terminal.	2026-04-16 22:08:33 +02:00
Andrei Chekun	13770ab394	test.py: delete dead code in test.py With the latest changes, there are a lot of code that is redundant in the test.py. This PR just cleans this code. Changes in other files are related to cleaning code from the test.py, especially with redundant parameter --test-py-init and moving prepare_environment to pytest itself.	2026-04-16 22:08:31 +02:00
Avi Kivity	999e108139	Merge 'test: lib: fix broken retry in start_docker_service' from Dario Mirovic The retry loop in `start_docker_service` passes the parse callbacks via `std::move` into `create_handler` on each iteration. After the first iteration, the moved-from `std::function` objects are empty. All subsequent retries skip output parsing entirely and immediately treat the service as successfully started. This defeats the entire purpose of the retry mechanism. Fix by passing the callbacks by copy instead of move, so the original callbacks remain valid across retries. Fixes SCYLLADB-1542 This is a CI stability issue and should be backported. Closes scylladb/scylladb#29504 * github.com:scylladb/scylladb: test/lib: fix typos in proc_utils, gcs_fixture, and dockerized_service test: gcs_fixture: rename container from "local-kms" to "fake-gcs-server" test: fix proc_utils.cc formatting from previous commit test: lib: use unique container name per retry attempt test: lib: fix broken retry in start_docker_service	2026-04-16 21:48:25 +03:00

1 2 3 4 5 ...

53361 Commits