scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-04-26 11:30:36 +00:00

Author	SHA1	Message	Date
Patryk Jędrzejczak	2b984f07cf	test: test_maintenance_mode: check that group0 is disabled by creating a keyspace In the following commit, we make the rest run with multiple keyspaces, and the old check becomes inconvenient. We also move it below to the part of the code that won't be executed for each keyspace. Additionally, we check if the error message is as expected. (cherry picked from commit `53f58b85b7`)	2026-02-03 11:33:52 +01:00
Patryk Jędrzejczak	59aee20992	test: test_maintenance_mode: get rid of the conditional skip This skip has already caused trouble. After `0668c642a2`, the skip was always hit, and the test was silently doing nothing. This made us miss #26816 for a long time. The test was fixed in `222eab45f8`, but we should get rid of the skip anyway. We increase the number of writes from 256 to 1000 to make the chance of not finding the key on server A even lower. If that still happens, it must be due to a bug, so we fail the test. We also make the test insert rows until server A is a replica of one row. The expected number of inserted rows is a small constant, so it should, in theory, make the test faster and cleaner (we need one row on server A, so we insert exactly one such row). It's possible to make the test fully deterministic, by e.g., hardcoding the key and tokens of all nodes via `initial_token`, but I'm afraid it would make the test "too deterministic" and could hide a bug. (cherry picked from commit `408c6ea3ee`)	2026-02-03 11:33:52 +01:00
Patryk Jędrzejczak	7f133e72af	test: test_maintenance_mode: remove the redundant value from the query result (cherry picked from commit `c92962ca45`)	2026-02-03 11:33:52 +01:00
Patryk Jędrzejczak	0df478e02e	storage_proxy: skip validate_read_replica in maintenance mode In maintenance mode, the local node adds only itself to the topology. However, the effective replication map of a keyspace with tablets enabled contains all tablet replicas. It gets them from the tablets map, not the topology. Hence, `network_topology_strategy::sanity_check_read_replicas` hits ``` throw std::runtime_error(format("Requested location for node {} not in topology. backtrace {}", id, lazy_backtrace())); ``` for tablet replicas other than the local node. As a result, all requests to a keyspace with tablets enabled and RF > 1 fail in debug mode (`validate_read_replica` does nothing in other modes). We don't want to skip maintenance mode tests in debug mode, so we skip the check in maintenance mode. We move the `is_debug_build()` check because: - `validate_read_replicas` is a static function with no access to the config, - we want the `!_db.local().get_config().maintenance_mode()` check to be dropped by the compiler in non-debug builds. We also suppress `-Wunneeded-internal-declaration` with `[[maybe_unused]]`. (cherry picked from commit `9d4a5ade08`)	2026-02-03 11:33:52 +01:00
Patryk Jędrzejczak	3ccfed37f6	storage_service: set up topology properly in maintenance mode We currently make the local node the only token owner (that owns the whole ring) in maintenance mode, but we don't update the topology properly. The node is present in the topology, but in the `none` state. That's how it's inserted by `tm.get_topology().set_host_id_cfg(host_id);` in `scylla_main`. As a result, the node started in maintenance mode crashes in the following way in the presence of a vnodes-based keyspace with the NetworkTopologyStrategy: ``` scylla: locator/network_topology_strategy.cc:207: locator::natural_endpoints_tracker::natural_endpoints_tracker( const token_metadata &, const network_topology_strategy::dc_rep_factor_map &): Assertion `!_token_owners.empty() && !_racks.empty()' failed. ``` Both `_token_owners` and `_racks` are empty. The reason is that `_tm.get_datacenter_token_owners()` and `_tm.get_datacenter_racks_token_owners()` called above filter out nodes in the `none` state. This bug basically made maintenance mode unusable in customer clusters. We fix it by changing the node state to `normal`. We also update its rack, datacenter, and shards count. Rack and datacenter are present in the topology somehow, but there is nothing wrong with updating them again. The shard count is also missing, so we better update it to avoid other issues. Fixes #27988 (cherry picked from commit `a08c53ae4b`)	2026-02-03 11:33:51 +01:00
Jenkins Promoter	6698ee205f	Update pgo profiles - aarch64	2026-02-01 04:32:56 +02:00
Jenkins Promoter	5942e636ea	Update pgo profiles - x86_64	2026-02-01 04:03:33 +02:00
Łukasz Paszkowski	960c952eb5	storage/test_out_of_space_prevention.py: Fix async/await bugs - Add missing await keywords for async operations on s2_log.wait_for() and coord_log.wait_for() - Fix incorrect regex: "compaction .* Split {cf}" → "compaction.*Split {cf}" - The commit https://github.com/scylladb/scylladb/commit/f7324a4 demoted compaction start/end log messages to debug level. Hence add compaction=debug log messages to the following tests: test_split_compaction_not_triggered test_node_restart_while_tablet_split test_repair_failure_on_split_rejection Fixes https://github.com/scylladb/scylladb/issues/27931 Closes scylladb/scylladb#27932 (cherry picked from commit `76b84b71d1`) Closes scylladb/scylladb#27949	2026-01-31 20:59:03 +02:00
Aleksandra Martyniuk	196143b830	service: node_ops: remove coroutine::lambda wrappers In storage_service::raft_topology_cmd_handler we pass a lambda wrapped in coroutine::lambda to a function that creates streaming_task_impl. The lambda is kept in streaming_task_impl that invokes it in its run method. The lambda captures may be destroyed before the lambda is called, leading to use after free. Do not wrap a lambda passed to streaming_task_impl into coroutine::lambda. Use this auto dissociate the lambda lifetime from the calling statement. Fixes: https://github.com/scylladb/scylladb/issues/28200. Closes scylladb/scylladb#28201 (cherry picked from commit `65cba0c3e7`) Closes scylladb/scylladb#28244	2026-01-30 16:12:11 +02:00
Botond Dénes	3573535167	Merge '[Backport 2025.4] schema: Apply `sstable_compression_user_table_options` to CQL aux and Alternator tables' from Scylladb[bot] In PR `5b6570be52` we introduced the config option `sstable_compression_user_table_options` to allow adjusting the default compression settings for user tables. However, the new option was hooked into the CQL layer and applied only to CQL base tables, not to the whole spectrum of user tables: CQL auxiliary tables (materialized views, secondary indexes, CDC log tables), Alternator base tables, Alternator auxiliary tables (GSIs, LSIs, Streams). This gap also led to inconsistent default compression algorithms after we changed the option’s default algorithm from LZ4 to LZ4WithDicts (`adf9c426c2`). This series introduces a general “schema initializer” mechanism in `schema_builder` and uses it to apply the default compression settings uniformly across all user tables. This ensures that all base and aux tables take their default compression settings from config. Fixes #26914. Backport justification: LZ4WithDicts is the new default since 2025.4, but the config option exists since 2025.2. Based on severity, I suggest we backport only to 2025.4 to maintain consistency of the defaults. - (cherry picked from commit `4ec7a064a9`) - (cherry picked from commit `76b2d0f961`) - (cherry picked from commit `5b4aa4b6a6`) - (cherry picked from commit `d5ec66bc0c`) - (cherry picked from commit `1e37781d86`) - (cherry picked from commit `7fa1f87355`) Parent PR: #27204 Closes scylladb/scylladb#28305 * github.com:scylladb/scylladb: db/config: Update sstable_compression_user_table_options description schema: Add initializer for compression defaults schema: Generalize static configurators into schema initializers schema: Initialize static properties eagerly db: config: Add accessor for sstable_compression_user_table_options test: Check that CQL and Alternator tables respect compression config test/cqlpy: test compression setting for auxiliary table test/alternator: tests for schema of Alternator table	2026-01-30 16:10:35 +02:00
Ernest Zaslavsky	c9fe14b79d	aws_error: handle all restartable nested exception types Previously we only inspected std::system_error inside std::nested_exception to support a specific TLS-related failure mode. However, nested exceptions may contain any type, including other restartable (retryable) errors. This change unwraps one nested exception per iteration and re-applies all known handlers until a match is found or the chain is exhausted. Closes scylladb/scylladb#28240 (cherry picked from commit `cb2aa85cf5`) Closes scylladb/scylladb#28344	2026-01-30 16:08:59 +02:00
Avi Kivity	498f74f53d	test/cqlpy: restore LWT tests marked XFAIL for tablets Commit `0156e97560` ("storage_proxy: cas: reject for tablets-enabled tables") marked a bunch of LWT tests as XFAIL with tablets enabled, pending resolution of #18066. But since that event is now in the past, we undo the XFAIL markings (or in some cases, use an any-keyspace fixture instead of a vnodes-only fixture). Ref #18066. Closes scylladb/scylladb#28336 (cherry picked from commit `ec70cea2a1`) [avi: skip counters-with-tablets test as not supported in 2025.4] Closes scylladb/scylladb#28364	2026-01-30 16:07:27 +02:00
Patryk Jędrzejczak	538295e97b	test: test_gossiper_orphan_remover: get host ID of the bootstrapping node before it crashes The test is currently flaky. It tries to get the host ID of the bootstrapping node via the REST API after the node crashes. This can obviously fail. The test usually doesn't fail, though, as it relies on the host ID being saved in `ScyllaServer._host_id` at this point by `ScyllaServer.try_get_host_id()` repeatedly called in `ScyllaServer.start()`. However, with a very fast crash and unlucky timings, no such call may succeed. We deflake the test by getting the host ID before the crash. Note that at this point, the bootstrapping node must be serving the REST API requests because `await log.wait_for("finished do_send_ack2_msg")` above guarantees that the node has started the gossip shadow round, which happens after starting the REST API. Fixes #28385 Closes scylladb/scylladb#28388 (cherry picked from commit `a2c1569e04`) Closes scylladb/scylladb#28416	2026-01-29 11:27:37 +01:00
Nikos Dragazis	371cff0e06	db/config: Update sstable_compression_user_table_options description Clarify what "user table" means. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> (cherry picked from commit `7fa1f87355`)	2026-01-28 12:42:10 +02:00
Nikos Dragazis	914d3f845a	schema: Add initializer for compression defaults In PR `5b6570be52` we introduced the config option `sstable_compression_user_table_options` to allow adjusting the default compression settings for user tables. However, the new option was hooked into the CQL layer and applied only to CQL base tables, not to the whole spectrum of user tables: CQL auxiliary tables (materialized views, secondary indexes, CDC log tables), Alternator base tables, Alternator auxiliary tables (GSIs, LSIs, Streams). Fix this by moving the logic into the `schema_builder` via a schema initializer. This ensures that the default compression settings are applied uniformly regardless of how the table is created, while also keeping the logic in a central place. Register the initializer at startup in all executables where schemas are being used (`scylla_main()`, `scylla_sstable_main()`, `cql_test_env`). Finally, remove the ad-hoc logic from `create_table_statement` (redundant as of this patch), remove the xfail markers from the relevant tests and adjust `test_describe_cdc_log_table_create_statement` to expect LZ4WithDicts as the default compressor. Fixes #26914. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> (cherry picked from commit `1e37781d86`)	2026-01-28 12:42:10 +02:00
Nikos Dragazis	b6deff2547	schema: Generalize static configurators into schema initializers Extend the `static_configurator` mechanism to support initialization of arbitrary schema properties, not only static ones, by passing a `schema_builder` reference to the configurator interface. As part of this change, rename `static_configurator` to `schema_initializer` to better reflect its broader responsibility. Add a checkpoint/restore mechanism to allow de-registering an initializer (useful for testing; will be used in the next patch). Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> (cherry picked from commit `d5ec66bc0c`)	2026-01-28 12:42:10 +02:00
Nikos Dragazis	a817aee457	schema: Initialize static properties eagerly Schemas maintain a set of so-called "static properties". These are not user-visible schema properties; they are internal values carried by in-memory `schema` objects for convenience (`349bc1a9b6`, https://github.com/scylladb/scylladb/pull/13170#issuecomment-1469848086). Currently, the initialization of these properties happens when a `schema_builder` builds a schema (`schema_builder::build()`), by invoking all registered "static configurators". This patch moves the initialization of static properties into the `schema_builder` constructor. With this change, the builder initializes the properties once, stores them in a data member, and reuses them for all schema objects that it builds. This doesn't affect correctness as the values produced by static configurators are "static" by nature; they do not depend on runtime state. In the next patch, we will replace the "static configurator" pattern with a more general pattern that also supports initialization of regular schema properties, not just static ones. Regular properties cannot be initialized in `build()` because users may have already explicitly set values via setters, and there is no way to distinguish between default values and explicitly assigned ones. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> (cherry picked from commit `5b4aa4b6a6`)	2026-01-28 12:42:10 +02:00
Nikos Dragazis	001581c69c	db: config: Add accessor for sstable_compression_user_table_options The `sstable_compression_user_table_options` config option determines the default compression settings for user tables. In patch `2fc812a1b9`, the default value of this option was changed from LZ4 to LZ4WithDicts and a fallback logic was introduced during startup to temporarily revert the option to LZ4 until the dictionary compression feature is enabled. Replace this fallback logic with an accessor that returns the correct settings depending on the feature flag. This is cleaner and more consistent with the way we handle the `sstable_format` option, where the same problem appears (see `get_preferred_sstable_version()`). As a consequence, the configuration option must always be accessed through this accessor. Add a comment to point this out. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> (cherry picked from commit `76b2d0f961`)	2026-01-28 12:42:07 +02:00
Nikos Dragazis	80e064d61c	test: Check that CQL and Alternator tables respect compression config In patches `11f6a25d44` and `7b9428d8d7` we added tests to verify that auxiliary tables for both CQL and Alternator have the same default compression settings as their base tables. These tests do not check where these defaults originate from; they just verify that they are consistent. Add some more tests to verify the actual source of the defaults, which is expected to be the `sstable_compression_user_table_options` from the configuration. Unlike the previous tests, these tests require dedicated Scylla instances with custom configuration, so they must be placed under `test/cluster/`. Mark them as xfail-ing. The marker will be removed later in this series. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> (cherry picked from commit `4ec7a064a9`)	2026-01-27 19:33:35 +02:00
Nadav Har'El	cac8096e41	test/cqlpy: test compression setting for auxiliary table In the previous patch we noticed that although recently (commit `adf9c42`, Refs #26610) we changed the default sstable compressor from LZ4Compressor to LZ4WithDictsCompressor, this change was only applied to CQL, not to Alternator. In this patch we add tests that demonstrate that it's even worse - the new compression only applies to CQL's base table - all the "auxiliary" tables - * Materialized views * Secondary index's materialized views * CDC log tables all still have the old LZ4Compressor, different from the base table's default compressor. The new test fails on Scylla, reproducing #26914, and passes on Cassandra (on Cassandra, we only compare the materialized view table, because SI and CDC is implemented differently). Signed-off-by: Nadav Har'El <nyh@scylladb.com> (cherry picked from commit `7b9428d8d7`)	2026-01-27 19:33:21 +02:00
Nadav Har'El	b395458af2	test/alternator: tests for schema of Alternator table This patch introduces a new test that exposed a previously unknown bug, Refs #26914: Recently we saw a lot of patches that change how we create new schemas (keyspaces and tables), sometimes changing various long-time defaults. We started to worry that perhaps some of these defaults were applied only to CQL and not to Alternator. For example, in Refs #26307 we wondered if perhaps the default "speculative_retry" option is different in Alternator than in CQL. This patch includes a new test file test/alternator/test_cql_schema.py, with tests for verifying how Alternator configures the underlying tables it creates. This test shows that the "speculative_retry" doesn't have this suspected bug - it defaults to "99.0PERCENTILE" in both CQL and Alternator. But unfortunately, we do have this bug with the "compression" option: It turns out that recently (commit `adf9c42`, Refs #26610) we changed the default sstable compressor from LZ4Compressor to LZ4WithDictsCompressor, but the change was only applied to CQL, not Alternator. So the test that "compression" is the same in both fails - and marked "xfails" and I created a new issue to track it - #26914. Another test verifies that Alternators "auxiliary" tables - holding GSIs, LSIs and Streams - have the same default properties as the base table. This currently seems to hold (there is no bug). Signed-off-by: Nadav Har'El <nyh@scylladb.com> (cherry picked from commit `11f6a25d44`)	2026-01-27 19:32:27 +02:00
Yaron Kaikov	2c4fef8638	.github/workflows/backport-pr-fixes-validation: support Atlassian URL format in backport PR fixes validation Add support for matching full Atlassian JIRA URLs in the format https://scylladb.atlassian.net/browse/SCYLLADB-400 in addition to the bare JIRA key format (SCYLLADB-400). This makes the validation more flexible by accepting both formats that developers commonly use when referencing JIRA issues. Fixes: https://github.com/scylladb/scylladb/issues/28373 Closes scylladb/scylladb#28374 (cherry picked from commit `3f10f44232`) Closes scylladb/scylladb#28393	2026-01-27 16:05:34 +02:00
Avi Kivity	2a457484d2	Merge '[Backport 2025.4] test_lwt_shutdown: fix flakiness by removing storage_proxy::stop injection' from Scylladb[bot] The storage_proxy::stop() is not called by main (it is commented out due to #293), so the corresponding message injection is never hit. When the test releases paxos_state_learn_after_mutate, shutdown may already be in progress or even completed by the time we try to trigger the storage_proxy::stop injection, which makes the test flaky. Fix this by completely removing the storage_proxy::stop injection. The injection is not required for test correctness. Shutdown must wait for the background LWT learn to finish, which is released via the paxos_state_learn_after_mutate injection. The shutdown process blocks on in-flight HTTP requests through seastar::httpd::http_server::stop and its _task_gate, so the HTTP request that releases paxos_state_learn_after_mutate is guaranteed to complete before the node is shut down. Fixes scylladb/scylladb#28260 backport: 2025.4, the `test_lwt_shutdown` test was introduced in this version - (cherry picked from commit `f5ed3e9fea`) - (cherry picked from commit `c45244b235`) Parent PR: #28315 Closes scylladb/scylladb#28331 * github.com:scylladb/scylladb: storage_proxy: drop stop() method test_lwt_shutdown: fix flakiness by removing storage_proxy::stop injection	2026-01-26 12:34:35 +02:00
Petr Gusev	a164887f15	storage_proxy: drop stop() method It's not called by main.cc and can be confusing. (cherry picked from commit `c45244b235`)	2026-01-23 19:24:06 +00:00
Petr Gusev	69a24bb933	test_lwt_shutdown: fix flakiness by removing storage_proxy::stop injection storage_proxy::stop() is not called by main (it is commented out due to #293), so the corresponding message injection is never hit. When the test releases paxos_state_learn_after_mutate, shutdown may already be in progress or even completed by the time we try to trigger the storage_proxy::stop injection, which makes the test flaky. Fix this by completely removing the storage_proxy::stop injection. The injection is not required for test correctness. Shutdown must wait for the background LWT learn to finish, which is released via the paxos_state_learn_after_mutate injection. The shutdown process blocks on in-flight api HTTP requests through seastar::httpd::http_server::stop and its _task_gate, so the shutdown will not prevent the HTTP request that released the paxos_state_learn_after_mutate from completing successfully. Fixes scylladb/scylladb#28260 (cherry picked from commit `f5ed3e9fea`)	2026-01-23 19:24:06 +00:00
Szymon Wasik	9c01dd8d49	Add vector search documentation links to CQL docs This patch adds links to the Vector Search documentation that is hosted together with Scylla Cloud docs to the CQL documentation. It also make the note about supported capabilities consistent and removes the experimental label as the feature is GAed. Fixes: SCYLLADB-371 Closes scylladb/scylladb#28312 (cherry picked from commit `927aebef37`) Closes scylladb/scylladb#28319	2026-01-23 12:00:22 +01:00
Patryk Jędrzejczak	d4277c95e8	test: test_zero_token_nodes_multidc: properly handle reads with CL=LOCAL_ONE The test is currently flaky. It incorrectly assumes that a read with CL=LOCAL_ONE will see the data inserted by a preceding write with CL=LOCAL_ONE in the same datacenter with RF=2. The same issue has already been fixed for CL=ONE in `21edec1ace`. The difference is that for CL=LOCAL_ONE, only dc1 is problematic, as dc2 has RF=1. We fix the issue for CL=LOCAL_ONE by skipping the check for dc1. Fixes #28253 The fix addresses CI flakiness and only changes the test, so it should be backported. Closes scylladb/scylladb#28274 (cherry picked from commit `1f0f694c9e`) Closes scylladb/scylladb#28304	2026-01-22 18:22:05 +01:00
Patryk Jędrzejczak	a247e19f56	test: test_raft_recovery_during_join: get host ID of the bootstrapping node before it crashes The test is currently flaky. It tries to get the host ID of the bootstrapping node via the REST API after the node crashes. This can obviously fail. The test usually doesn't fail, though, as it relies on the host ID being saved in `ScyllaServer._host_id` at this point by `ScyllaServer.try_get_host_id()` repeatedly called in `ScyllaServer.start()`. However, with a very fast crash and unlucky timings, no such call may succeed. We deflake the test by getting the host ID before the crash. Note that at this point, the bootstrapping node must be serving the REST API requests because `await coordinator_log.wait_for("delay_node_bootstrap: waiting for message")` above guarantees that the node has submitted the join topology request, which happens after starting the REST API. Fixes #28227 Closes scylladb/scylladb#28233 (cherry picked from commit `e503340efc`) Closes scylladb/scylladb#28310	2026-01-22 18:19:36 +01:00
Botond Dénes	62f399d8db	db/row_cache: make_nonpopulating_reader(): pass cache tracker to snapshot The API contract in partition_version.hh states that when dealing with evictable entries, a real cache tracker pointer has to be passed to all methods that ask for it. The nonpopulating reader violates this, passing a nullptr to the snapshot. This was observed to cause a crash when a concurrent cache read accessed the snapshot with the null tracker. A reproducer is included which fails before and passes after the fix. Fixes: #26847 Closes scylladb/scylladb#28163 (cherry picked from commit `a53f989d2f`) Closes scylladb/scylladb#28279	2026-01-22 12:38:00 +02:00
Jenkins Promoter	ade92555c3	Update ScyllaDB version to: 2025.4.3	2026-01-21 23:16:44 +02:00
Szymon Wasik	9d9ee8ad68	Improve documentation of vector search configuration parameters. This patch adds separate group for vector search parameters in the documentation and fixes small typos and formatting. Fixes: SCYLLADB-77. Closes scylladb/scylladb#27385 (cherry picked from commit `4f803aad22`) Closes scylladb/scylladb#27424	2026-01-21 13:21:32 +02:00
Anna Stuchlik	ea222186bc	doc: fix the default compaction strategy for Materialized Views Fixes https://github.com/scylladb/scylladb/issues/24483 Closes scylladb/scylladb#27725 (cherry picked from commit `84e9b94503`) Closes scylladb/scylladb#28285	2026-01-21 06:38:01 +02:00
Botond Dénes	e8f5ac5fb6	reader_concurrency_semaphore: improve handling of base resources reader_permit::release_base_resources() is a soft evict for the permit: it releases the resources aquired during admission. This is used in cases where a single process owns multiple permits, creating a risk for deadlock, like it is the case for repair. In this case, release_base_resources() acts as a manual eviction mechanism to prevent permits blockings each other from admission. Recently we found a bad interaction between release_base_resources() and permit eviction. Repair uses both mechanism: it marks its permits as inactive and later it also uses release_base_resources(). This partice might be worth reconsidering, but the fact remains that there is a bug in the reader permit which causes the base resources to be released twice when release_base_resources() is called on an already evicted permit. This is incorrect and is fixed in this patch. Improve release_base_resources(): * make _base_resources const * move signal call into the if (_base_resources_consumed()) { } * use reader_permit::impl::signal() instead of reader_concurrency_semaphore::signal() * all places where base resources are released now call release_base_resources() A reproducer unit test is added, which fails before and passes after the fix. Fixes: #28083 Closes scylladb/scylladb#28155 (cherry picked from commit `b7bc48e7b7`) Closes scylladb/scylladb#28245	2026-01-21 06:37:38 +02:00
Botond Dénes	ee59610b9b	Merge '[Backport 2025.4] The system_replicated_keys should be mark as a system keyspace' from Scylladb[bot] This PR marks system_replicated_keys as a system keyspace. It was missing when the keyspace was added. A side effect of that is that metrics that are not supposed to be reported are. Fixes #27903 - (cherry picked from commit `83c1103917`) - (cherry picked from commit `c6d1c63ddb`) Parent PR: #27954 Closes scylladb/scylladb#28237 * github.com:scylladb/scylladb: distributed_loader: system_replicated_keys as system keyspace replicated_key_provider: make KSNAME public	2026-01-21 06:36:43 +02:00
Michał Hudobski	4af195c866	cql: fail with a better error when null vector is passed to ann query Currently when a null vector is passed to an ANN query we fail with a quite confusing error ("NoHostAvailable: ('Unable to complete the operation against any hosts', {<Host: 127.0.0.1:9042 datacenter1>: <Error from server: code=0000 [Server error] message="to_bytes() called on raw value that is null">})"). This patch fixes that by throwing an InvalidRequestException with an appropriate message instead. We also add a test case that validates this behavior. Fixes: VECTOR-257 Closes scylladb/scylladb#26510 (cherry picked from commit `541b52cdbf`) Closes scylladb/scylladb#28052	2026-01-21 06:35:58 +02:00
Michał Hudobski	9a0849ef36	vector search, paging: add test for paging warnings We add a test that validates that indexed queries do not throw a warning related to vector search paging Fixes: SCYLLADB-248 Closes scylladb/scylladb#28077 (cherry picked from commit `c8aa49b196`) Closes scylladb/scylladb#28138	2026-01-20 12:57:54 +02:00
Ernest Zaslavsky	177985a69d	aws_error: fix nested exception handling The loop that unwraps nested exception, rethrows nested exception and saves pointer to the temporary std::exception& inner on stack, then continues. This pointer is, thus, pointing to a released temporary Closes scylladb/scylladb#28143 (cherry picked from commit `829bd9b598`) Closes scylladb/scylladb#28243	2026-01-20 11:22:32 +01:00
Patryk Jędrzejczak	8b15975cb8	test: test_group0_schema_versioning: wait for schema sync in system.local `test_schema_versioning_with_recovery` is currently flaky. It performs a write with CL=ALL and then checks if the schema version is the same on all nodes by calling `verify_table_versions_synced`. All nodes are expected to sync their schema before handling the replica write. The node in RECOVERY mode should do it through a schema pull, and other nodes should do it through a group 0 read barrier. The problem is in `verify_local_schema_versions_synced` that compares the schema versions in `system.local`. The node in RECOVERY mode updates the schema version in `system.local` after it acknowledges the replica write as completed. Hence, the check can fail. We fix the problem by making the function wait until the schema versions match. Note that RECOVERY mode is about to be retired together with the whole gossip-based topology in 2026.2. So, this test is about to be deleted. However, we still want to fix it, so that it doesn't bother us in older branches. Fixes #23803 Closes scylladb/scylladb#28114 (cherry picked from commit `6b5923c64e`) Closes scylladb/scylladb#28178	2026-01-19 16:35:49 +01:00
Amnon Heiman	1d8089408a	distributed_loader: system_replicated_keys as system keyspace This patch adds system_replicated_keys to the list of known system keyspaces. Signed-off-by: Amnon Heiman <amnon@scylladb.com> (cherry picked from commit `c6d1c63ddb`)	2026-01-19 14:59:26 +00:00
Amnon Heiman	b473736cd6	replicated_key_provider: make KSNAME public Move KSNAME constant from internal static to public member of replicated_key_provider_factory class. It will be used to identify it as a system keyspace. Signed-off-by: Amnon Heiman <amnon@scylladb.com> (cherry picked from commit `83c1103917`)	2026-01-19 14:59:26 +00:00
Gleb Natapov	4e4bfee41e	topology coordinator: complete pending operation for a replaced node A replaced node may have pending operation on it. The replace operation will move the node into the 'left' state and the request will never be completed. More over the code does not expect left node to have a request. It will try to process the request and will crash because the node for the request will not be found. The patch checks is the replaced node has peening request and completes it with failure. It also changes topology loading code to skip requests for nodes that are in a left state. This is not strictly needed, but makes the code more robust. Fixes #27990 Closes scylladb/scylladb#28009 (cherry picked from commit `bee5f63cb6`) Closes scylladb/scylladb#28180	2026-01-19 09:42:20 +02:00
Asias He	09da47a42c	repair: Fix sstable_list_to_mark_as_repaired with multishard writer It was obseved: ``` test_repair_disjoint_row_2nodes_diff_shard_count was spuriously failing due to segfault. backtrace pointed to a failure when allocating an object from the chain of freed objects, which indicates memory corruption. (gdb) bt at ./seastar/include/seastar/core/shared_ptr.hh:275 at ./seastar/include/seastar/core/shared_ptr.hh:430 Usual suspect is use-after-free, so ran the reproducer in the sanitize mode, which indicated shared ptr was being copied into another cpu through the multi shard writer: seastar - shared_ptr accessed on non-owner cpu, at: ... -------- seastar::smp_message_queue::async_work_item<mutation_writer::multishard_writer::make_shard_writer... ``` The multishard writer itself was fine, the problem was in the streaming consumer for repair copying a shared ptr. It could work fine with same smp setting, since there will be only 1 shard in the consumer path, from rpc handler all the way to the consumer. But with mixed smp setting, the ptr would be copied into the cpus involved, and since the shared ptr is not cpu safe, the refcount change can go wrong, causing double free, use-after-free. To fix, we pass a generic incremental repair handler to the streaming consumer. The handler is safe to be copied to different shards. It will be a no op if incremental repair is not enabled or on a different shard. A reproducer test is added. The test could reproduce the crash consistently before the fix and work well after the fix. Fixes #27666 Closes scylladb/scylladb#27870 (cherry picked from commit `0aabf51380`) Closes scylladb/scylladb#28064	2026-01-19 09:39:49 +02:00
Asias He	c5aa29404d	repair: Add tablet repair progress report support This patch adds tablet repair progress report support so that the user could use the /task_manager/task_status API to query the progress. In order to support this, a new system table is introduced to record the user request related info, i.e, start of the request and end of the request. The progress is accurate when tablet split or merge happens in the middle of the request, since the tokens of the tablet are recorded when the request is started and when repair of each tablet is finished. The original tablet repair is considered as finished when the finished ranges cover the original tablet token ranges. After this patch, the /task_manager/task_status API will report correct progress_total and progress_completed. Fixes #22564 Fixes #26896 Closes scylladb/scylladb#27679 (cherry picked from commit `4f77dd058d`) Closes scylladb/scylladb#28065	2026-01-19 09:39:13 +02:00
Nikos Dragazis	29c534c6e7	test: database_test: Fix serialization of partition key The `make_key` lambda erroneously allocates a fixed 8-byte buffer (`sizeof(s.size())`) for variable-length strings, potentially causing uninitialized bytes to be included. If such bytes exist and they are not valid UTF-8 characters, deserialization fails: ``` ERROR 2026-01-16 08:18:26,062 [shard 0:main] testlog - snapshot_list_contains_dropped_tables: cql env callback failed, error: exceptions::invalid_request_exception (Exception while binding column p1: marshaling error: Validation failed - non-UTF8 character in a UTF8 string, at byte offset 7) ``` Fixes #28195. Signed-off-by: Nikos Dragazis <nikolaos.dragazis@scylladb.com> Closes scylladb/scylladb#28197 (cherry picked from commit `8aca7b0eb9`) Closes scylladb/scylladb#28209	2026-01-19 09:38:45 +02:00
Łukasz Paszkowski	c04007c755	database: Log message after critical_disk_utilization mode is set This is a follow-up of the previous fix: https://github.com/scylladb/scylladb/pull/26030 The test test_user_writes_rejection starts a 3-node cluster and creates a large file on one of the nodes, to trigger the out-of-space prevention mechanism, which should reject writes on that node. It waits for the log message 'Setting critical disk utilization mode: true' and then executes a write expecting the node to reject it. Currently, the message is logged before the `_critical_disk_utilization` variable is actually updated. This causes the test to fail sporadically if it runs quickly enough. The fix splits the logging into two steps: 1. "Asked to set critical disk utilization mode" - logged before any action 2) "Set critical disk utilization mode" - logged after `_critical_disk_utilization` has been updated The tests are updated to wait for the second message. Fixes https://github.com/scylladb/scylladb/issues/26004 Closes scylladb/scylladb#26392 (cherry picked from commit `7ec369b900`) Closes scylladb/scylladb#26626	2026-01-19 06:42:01 +02:00
Botond Dénes	b738be094f	Merge '[Backport 2025.4] Make commitlog replay handle files with corrupt file header (non-zero) as data loss, not startup failure' from Scylladb[bot] Fixes #26744 If a segment to replay is broken such that the main header is not zero, but still broken, we throw header_checksum_error. This was not handled in replayer, which grouped this into the "user error/fundamental problem" category. However, assuming we allow for "real" disk corruption, this should really be treated same as data corruption, i.e. reported data loss, not failure to start up. The `test_one_big_mutation_corrupted_on_startup` test accidentally sometimes provoked this issue, by doing random file wrecking, which on rare occasions provoked this, and thus failed test due to scylla not starting up, instead of losing data as expected. - (cherry picked from commit `9b5f3d12a3`) - (cherry picked from commit `e48170ca8e`) - (cherry picked from commit `8c4ac457af`) Parent PR: #27556 Closes scylladb/scylladb#27682 * github.com:scylladb/scylladb: test::cluster::dtest::tools::files: Remove file commitlog_replay: Handle fully corrupt files same as partial corruption. test::pylib::suite::base: Split options.name test specifier only once	2026-01-19 06:39:29 +02:00
Anna Stuchlik	64039588db	doc: clarify the information about SSTable version support Fixes https://github.com/scylladb/scylladb/issues/27765 Closes scylladb/scylladb#27835 (cherry picked from commit `791ab4ed02`) Closes scylladb/scylladb#28122 scylla-2025.4.2-candidate-20260118022006 scylla-2025.4.2	2026-01-16 16:19:47 +02:00
Calle Wilund	999dfb0e5e	db::commitlog: Fix sanity check error on race between segment flushing and oversized alloc Fixes #27992 When doing a commit log oversized allocation, we lock out all other writers by grabbing the _request_controller semaphore fully (max capacity). We thereafter assert that the semaphore is in fact zero. However, due to how things work with the bookkeep here, the semaphore can in fact become negative (some paths will not actually wait for the semaphore, because this could deadlock). Thus, if, after we grab the semaphore and execution actually returns to us (task schedule), new_buffer via segment::allocate is called (due to a non-fully-full segment), we might in fact grab the segment overhead from zero, resulting in a negative semaphore. The same problem applies later when we try to sanity check the return of our permits. Fix is trivial, just accept less-than-zero values, and take same possible ltz-value into account in exit check (returning units) Added whitebox (special callback interface for sync) unit test that provokes/creates the race condition explicitly (and reliably). Closes scylladb/scylladb#27998 (cherry picked from commit `a7cdb602e1`) Closes scylladb/scylladb#28099	2026-01-16 16:19:18 +02:00
Sergey Zolotukhin	96275adf1c	test: disable test_start_bootstrapped_with_invalid_seed The test intermittently fails when an invalid DNS name is resolved, likely due to ISP DNS error hijacking (see scylladb/scylladb#28153). Disable this test to unblock CI. Fixes scylladb/scylladb#28153 Closes scylladb/scylladb#28162 (cherry picked from commit `799d837295`)	2026-01-15 17:01:30 +02:00
Avi Kivity	8ae3db2aff	Update seastar submodule (accept unbounded recursion) * seastar 60e4b3b921...65f936e384 (1): > net: posix_server_socket_impl: coroutinize accept(), fix unbounded recursion Fixes #28166	2026-01-15 14:23:09 +02:00

1 2 3 4 5 ...

50301 Commits