scylladb

mirror of https://github.com/scylladb/scylladb.git synced 2026-05-24 16:52:12 +00:00

Author	SHA1	Message	Date
Pavel Emelyanov	57d267a97e	test: Do not check tablets mutations on nodes that don't have them The check is performed by selecting from mutation_fragments(table), but it's known that this query crashes Scylla when there's no tablet replica on that node. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-05-30 08:33:26 +03:00
Pavel Emelyanov	5b8523273b	test: Fix the way tablets RF-change test parses mutation_fragments When the test changes RF from 2 to 3, the extra node executes "rebuild" transition which means that it streams tablets replicas from two other peers. When doing it, the node receives two sets of sstables with mutations from the given tablet. The test part that checks if the extra node received the mutations notices two mutation fragments on the new replica and errorneously fails by seeing, that RF=3 is not equal to the number of mutations found, which is 4. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-05-30 08:33:26 +03:00
Pavel Emelyanov	6497ed68ed	test/tablets: Unmark RF-changing test with xfail Now the scailing works and test must check it does Signed-off-by: Pavel Emelyanov <xemul@scylladb.com>	2024-05-30 08:33:26 +03:00
Piotr Smaron	39c1237e25	docs: document ALTER KEYSPACE with tablets	2024-05-30 08:33:26 +03:00
Piotr Smaron	e04964ba17	Return response only when tablets are reallocated Up until now we waited until mutations are in place and then returned directly to the caller of the ALTER statement, but that doesn't imply that tablets were deleted/created, so we must wait until the whole processing is done and return only then.	2024-05-30 08:33:26 +03:00
Dawid Medrek	fb5b9012e6	cql-pytest: Verify RF is changes by at most 1 when tablets on This commit adds a test verifying that we can only change the RF of a keyspace for any DC by at most 1 when using tablets. Fixes #18029	2024-05-30 08:33:26 +03:00
Dawid Medrek	749197f0a4	cql3/alter_keyspace_statement: Do not allow for change of RF by more than 1 We want to ensure that when the replication factor of a keyspace changes, it changes by at most 1 per DC if it uses tablets. The rationale for that is to make sure that the old and new quorums overlap by at least one node. After these changes, attempts to change the RF of a keyspace in any DC by more than 1 will fail.	2024-05-30 08:33:26 +03:00
Piotr Smaron	1f4428153f	Reject ALTER with 'replication_factor' tag This patch removes the support for the "wildcard" replication_factor option for ALTER KEYSPACE when the keyspace supports tablets. It will still be supported for CREATE KEYSPACE so that a user doesn't have to know all datacenter names when creating the keyspace, but ALTER KEYSPACE will require that and the user will have to specify the exact change in replication factors they wish to make by explicitly specifying the datacenter names. Expanding the replication_factor option in the ALTER case is unintuitive and it's a trap many users fell into. See #8881, #15391, #16115	2024-05-30 08:33:26 +03:00
Piotr Smaron	544c424e89	Implement ALTER tablets KEYSPACE statement support This commit adds support for executing ALTER KS for keyspaces with tablets and utilizes all the previous commits. The ALTER KS is handled in alter_keyspace_statement, where a global topology request in generated with data attached to system.topology table. Then, once topology state machine is ready, it starts to handle this global topology event, which results in producing mutations required to change the schema of the keyspace, delete the system.topology's global req, produce tablets mutations and additional mutations for a table tracking the lifetime of the whole req. Tracking the lifetime is necessary to not return the control to the user too early, so the query processor only returns the response while the mutations are sent.	2024-05-30 08:33:25 +03:00
Piotr Smaron	73b59b244d	Parameterize migration_manager::announce by type to allow executing different raft commands Since ALTER KS requires creating topology_change raft command, some functions need to be extended to handle it. RAFT commands are recognized by types, so some functions are just going to be parameterized by type, i.e. made into templates. These templates are instantiated already, so that only 1 instances of each template exists across the whole code base, to avoid compiling it in each translation unit.	2024-05-30 08:33:15 +03:00
Piotr Smaron	5afa3028a3	Introduce TABLET_KEYSPACE event to differentiate processing path of a vnode vs tablets ks	2024-05-30 08:33:15 +03:00
Piotr Smaron	885c7309ee	Extend system.topology with 3 new columns to store data required to process alter ks global topo req Because ALTER KS will result in creating a global topo req, we'll have to pass the req data to topology coordinator's state machine, and the easiest way to do it is through sytem.topology table, which is going to be extended with 3 extra columns carrying all the data required to execute ALTER KS from within topology coordinator.	2024-05-30 08:33:15 +03:00
Piotr Smaron	adfad686b3	Allow query_processor to check if global topo queue is empty With current implementation only 1 global topo req can be executed at a time, so when ALTER KS is executed, we'll have to check if any other global topo req is ongoing and fail the req if that's the case.	2024-05-30 08:33:15 +03:00
Piotr Smaron	1a70db17a6	Introduce new global topo `keyspace_rf_change` req It will be used when processing ALTER KS statement, but also to create a separate processing path for a KS with tablets (as opposed to a vnode KS).	2024-05-30 08:33:15 +03:00
Piotr Smaron	bd4b781dc8	New raft cmd for both schema & topo changes Allows executing combined topology & schema mutations under a single RAFT command	2024-05-30 08:33:15 +03:00
Piotr Smaron	51b8b04d97	Add storage service to query processor Query processor needs to access storage service to check if global topology request is still ongoing and to be able to wait until it completes.	2024-05-30 08:33:15 +03:00
Paweł Zakrzewski	242caa14fe	tablets: tests for adding/removing replicas Note we're suppressing a UBSanitizer overflow error in UTs. That's because our linter complains about a possible overflow, which never happens, but tests are still failing because of it.	2024-05-30 08:33:15 +03:00
Paweł Zakrzewski	cedb47d843	tablet_allocator: make load_balancer_stats_manager configurable by name This is needed, because the same name cannot be used for 2 separate entities, because we're getting double-metrics-registration error, thus the names have to be configurable, not hardcoded.	2024-05-30 08:33:15 +03:00
Pavel Emelyanov	da816bf50c	test/tablets: Check that after RF change data is replicated properly There's a test that checks system.tablets contents to see that after changing ks replication factor via ALTER KEYSPACE the tablet map is updated properly. This patch extends this test that also validates that mutations themselves are replicated according to the desired replication factor. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> Closes scylladb/scylladb#18644	2024-05-30 08:31:48 +03:00
Botond Dénes	8bff078a89	Merge '[Backport 6.0] cdc, raft topology: fix and test cdc in the recovery mode' from ScyllaDB This PR ensures that CDC keeps working correctly in the recovery mode after leaving the raft-based topology. We update `system.cdc_local` in `topology_state_load` to ensure a node restarting in the recovery mode sees the last CDC generation created by the topology coordinator. Additionally, we extend the topology recovery test to verify that the CDC keeps working correctly during the whole recovery process. In particular, we test that after restarting nodes in the recovery mode, they correctly use the active CDC generation created by the topology coordinator. Fixes scylladb/scylladb#17409 Fixes scylladb/scylladb#17819 (cherry picked from commit `4351eee1f6`) (cherry picked from commit `68b6e8e13e`) (cherry picked from commit `388db33dec`) (cherry picked from commit `2111cb01df`) Refs #18820 Closes scylladb/scylladb#18938 * github.com:scylladb/scylladb: test: test_topology_recovery_basic: test CDC during recovery test: util: start_writes_to_cdc_table: add FIXME to increase CL test: util: start_writes_to_cdc_table: allow restarting with new cql storage_service: update system.cdc_local in topology_state_load	2024-05-29 16:14:02 +03:00
Botond Dénes	68d12daa7b	Merge '[Backport 6.0] Fix parsing of initial tablets by ALTER' from ScyllaDB If the user wants to change the default initial tablets value, it uses ALTER KEYSPACE statement. However, specifying `WITH tablets = { initial: $value }` will take no effect, because statement analyzer only applies `tablets` parameters together with the `replication` ones, so the working statement should be `WITH replication = $old_parameters AND tablets = ...` which is not very convenient. This PR changes the analyzer so that altering `tablets` happens independently from `replication`. Test included. fixes: #18801 (cherry picked from commit `8a612da155`) (cherry picked from commit `a172ef1bdf`) (cherry picked from commit `1003391ed6`) Refs #18899 Closes scylladb/scylladb#18918 * github.com:scylladb/scylladb: cql-pytest: Add validation of ALTER KEYSPACE WITH TABLETS cql3: Fix parsing of ALTER KEYSPACE's tablets parameters cql3: Remove unused ks_prop_defs/prepare_options() argument	2024-05-29 16:13:26 +03:00
Patryk Jędrzejczak	e1616a2970	test: test_topology_ops: stop a write worker after the first error `test_topology_ops` is flaky, which has been uncovered by gating in scylladb/scylladb#18707. However, debugging it is harder than it should be because write workers can flood the logs. They may send a lot of failed writes before the test fails. Then, the log file can become huge, even up to 20 GB. Fix this issue by stopping a write worker after the first error. This test is important for 6.0, so we can backport this change. (cherry picked from commit `7c1e6ba8b3`) Closes scylladb/scylladb#18914	2024-05-29 16:11:46 +03:00
Kefu Chai	62f5171a55	docs: fix typos in upgrade document s/Montioring/Monitoring/ Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> (cherry picked from commit `f1f3f009e7`) Closes scylladb/scylladb#18912	2024-05-29 16:11:08 +03:00
Piotr Smaron	fd928601ad	cql: fix a crash lurking in `ks_prop_defs::get_initial_tablets` `tablets_options->erase(it);` invalidates `it`, but it's still referred to later in the code in the last `else`, and when that code is invoked, we get a `heap-use-after-free` crash. Fixes: #18926 (cherry picked from commit `8a77a74d0e`) Closes scylladb/scylladb#18949	2024-05-29 16:10:24 +03:00
Aleksandra Martyniuk	ae474f6897	test: fix test_tombstone_gc.py Tests in test_tombstone_gc.py are parametrized with string instead of bool values. Fix that. Use the value to create a keyspace with or without tablets. Fixes: #18888. (cherry picked from commit `b7ae7e0b0e`) Closes scylladb/scylladb#18948	2024-05-29 16:09:45 +03:00
Anna Stuchlik	099338b766	doc: add the tablet limitation to the manual recovery procedure This commit adds the information that the manual recovery procedure is not supported if tablets are enabled. In addition, the content in the Manual Recovery Procedure is reorganized by adding the Prerequisites and Procedure subsections - in this way, we can limit the number of Note and Warning boxes that made the page hard to follow. Fixes https://github.com/scylladb/scylladb/issues/18895 (cherry picked from commit `cfa3cd4c94`) Closes scylladb/scylladb#18947	2024-05-29 16:08:44 +03:00
Anna Stuchlik	375610ace8	doc: document RF limitation This commit adds the information that the Replication Factor must be the same or higher than the number of nodes. (cherry picked from commit `87f311e1e0`) Closes scylladb/scylladb#18946	2024-05-29 16:08:05 +03:00
Botond Dénes	1b64e80393	Merge '[Backport 6.0] Harden the repair_service shutdown path' from ScyllaDB This series ignores errors in `load_history()` to prevent `abort_requested_exception` coming from `get_repair_module().check_in_shutdown()` from escaping during `repair_service::stop()`, causing ``` repair_service::~repair_service(): Assertion `_stopped' failed. ``` Fixes https://github.com/scylladb/scylladb/issues/18889 Backport to 6.0 required due to `523895145d` (cherry picked from commit `38845754c4`) (cherry picked from commit `c32c418cd5`) Refs #18890 Closes scylladb/scylladb#18944 * github.com:scylladb/scylladb: repair: load_history: warn and ignore all errors repair_service: debug stop	2024-05-29 16:07:40 +03:00
Benny Halevy	fa330a6a4d	repair: load_history: warn and ignore all errors Currently, the call to `get_repair_module().check_in_shutdown()` may throw `abort_requested_exception` that causes `repair_service::stop()` to fail, and trigger assertion failure in `~repair_service`. We alredy ignore failure from `update_repair_time`, so expand the logic to cover the whole function body. Fixes scylladb/scylladb#18889 Signed-off-by: Benny Halevy <bhalevy@scylladb.com> (cherry picked from commit `c32c418cd5`)	2024-05-28 17:07:26 +00:00
Benny Halevy	68544d5bb3	repair_service: debug stop Seen the following unexplained assertion failure with pytest -s -v --scylla-version=local_tarball --tablets repair_additional_test.py::TestRepairAdditional::test_repair_option_pr_multi_dc ``` INFO 2024-05-27 11:18:05,081 [shard 0:main] init - Shutting down repair service INFO 2024-05-27 11:18:05,081 [shard 0:main] task_manager - Stopping module repair INFO 2024-05-27 11:18:05,081 [shard 0:main] task_manager - Unregistered module repair INFO 2024-05-27 11:18:05,081 [shard 1:main] task_manager - Stopping module repair INFO 2024-05-27 11:18:05,081 [shard 1:main] task_manager - Unregistered module repair scylla: repair/row_level.cc:3230: repair_service::~repair_service(): Assertion `_stopped' failed. Aborting on shard 0. Backtrace: /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x3f040c /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x41c7a1 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x3dbaf /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x8e883 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x3dafd /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x2687e /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x2679a /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x36186 0x26f2428 0x10fb373 0x10fc8b8 0x10fc809 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x456c6d /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x456bcf 0x10fc65b 0x10fc5bc 0x10808d0 0x1080800 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x3ff22f /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x4003b7 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x3ff888 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x36dea8 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libseastar.so+0x36d0e2 0x101cefa 0x105a390 0x101bde7 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x27b89 /home/bhalevy/.ccm/scylla-repository/local_tarball/libreloc/libc.so.6+0x27c4a 0x101a764 ``` Decoded: ``` ~repair_service at ./repair/row_level.cc:3230 ~shared_ptr_count_for at ././seastar/include/seastar/core/shared_ptr.hh:491 (inlined by) ~shared_ptr_count_for at ././seastar/include/seastar/core/shared_ptr.hh:491 ~shared_ptr at ././seastar/include/seastar/core/shared_ptr.hh:569 (inlined by) seastar::shared_ptr<repair_service>::operator=(seastar::shared_ptr<repair_service>&&) at ././seastar/include/seastar/core/shared_ptr.hh:582 (inlined by) seastar::shared_ptr<repair_service>::operator=(decltype(nullptr)) at ././seastar/include/seastar/core/shared_ptr.hh:588 (inlined by) operator() at ././seastar/include/seastar/core/sharded.hh:727 (inlined by) seastar::future<void> seastar::futurize<seastar::future<void> >::invoke<seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}&>(seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}&) at ././seastar/include/seastar/core/future.hh:2035 (inlined by) seastar::futurize<std::invoke_result<seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}>::type>::type seastar::smp::submit_to<seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}>(unsigned int, seastar::smp_submit_to_options, seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}&&) at ././seastar/include/seastar/core/smp.hh:367 seastar::futurize<std::invoke_result<seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}>::type>::type seastar::smp::submit_to<seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}>(unsigned int, seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}::operator()(unsigned int) const::{lambda()#1}&&) at ././seastar/include/seastar/core/smp.hh:394 (inlined by) operator() at ././seastar/include/seastar/core/sharded.hh:725 (inlined by) seastar::future<void> std::__invoke_impl<seastar::future<void>, seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}&, unsigned int>(std::__invoke_other, seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}&, unsigned int&&) at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/invoke.h:61 (inlined by) std::enable_if<is_invocable_r_v<seastar::future<void>, seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}&, unsigned int>, seastar::future<void> >::type std::__invoke_r<seastar::future<void>, seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}&, unsigned int>(seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}&, unsigned int&&) at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/invoke.h:114 (inlined by) std::_Function_handler<seastar::future<void> (unsigned int), seastar::sharded<repair_service>::stop()::{lambda(seastar::future<void>)#1}::operator()(seastar::future<void>) const::{lambda(unsigned int)#1}>::_M_invoke(std::_Any_data const&, unsigned int&&) at /usr/bin/../lib/gcc/x86_64-redhat-linux/13/../../../../include/c++/13/bits/std_function.h:290 ``` FWIW, gdb crashed when opening the coredump. This commit will help catch the issue earlier when repair_service::stop() fails (and it must never fail) Signed-off-by: Benny Halevy <bhalevy@scylladb.com> (cherry picked from commit `38845754c4`)	2024-05-28 17:07:26 +00:00
Piotr Dulikowski	bc711a169d	Merge '[Backport 6.0] qos/raft_service_level_distributed_data_accessor: print correct error message when trying to modify a service level in recovery mode' from ScyllaDB Raft service levels are read-only in recovery mode. This patch adds check and proper error message when a user tries to modify service levels in recovery mode. Fixes https://github.com/scylladb/scylladb/issues/18827 (cherry picked from commit `2b56158d13`) (cherry picked from commit `ee08d7fdad`) (cherry picked from commit `af0b6bcc56`) Refs #18841 Closes scylladb/scylladb#18913 * github.com:scylladb/scylladb: test/auth_cluster/test_raft_service_levels: try to create sl in recovery service/qos/raft_sl_dda: reject changes to service levels in recovery mode service/qos/raft_sl_dda: extract raft_sl_dda steps to common function	2024-05-28 16:45:52 +02:00
Patryk Jędrzejczak	0d0c037e1d	test: test_topology_recovery_basic: test CDC during recovery In topology on raft, management of CDC generations is moved to the topology coordinator. We extend the topology recovery test to verify that the CDC keeps working correctly during the whole recovery process. In particular, we test that after restarting nodes in the recovery mode, they correctly use the active CDC generation created by the topology coordinator. A node restarting in the recovery mode should learn about the active generation from `system.cdc_local` (or from gossip, but we don't want to rely on it). Then, it should load its data from `system.cdc_generations_v3`. Fixes scylladb/scylladb#17409 (cherry picked from commit `2111cb01df`)	2024-05-28 14:02:00 +00:00
Patryk Jędrzejczak	4d616ccb8c	test: util: start_writes_to_cdc_table: add FIXME to increase CL (cherry picked from commit `388db33dec`)	2024-05-28 14:02:00 +00:00
Patryk Jędrzejczak	25d3398b93	test: util: start_writes_to_cdc_table: allow restarting with new cql This patch allows us to restart writing (to the same table with CDC enabled) with a new CQL session. It is useful when we want to continue writing after closing the first CQL session, which happens during the `reconnect_driver` call. We must stop writing before calling `reconnect_driver`. If a write started just before the first CQL session was closed, it would time out on the client. We rename `finish_and_verify` - `stop_and_verify` is a better name after introducing `restart`. (cherry picked from commit `68b6e8e13e`)	2024-05-28 14:02:00 +00:00
Patryk Jędrzejczak	ed3ac1eea4	storage_service: update system.cdc_local in topology_state_load When the node with CDC enabled and with the topology on raft disabled bootstraps, it reads system.cdc_local for the last generation. Nodes with both enabled use group0 to get the last generation. In the following scenario with a cluster of one node: 1. the node is created with CDC and the topology on raft enabled 2. the user creates table T 3. the node is restarted in the recovery mode 4. the CDC log of T is extended with new entries 5. the node restarts in normal mode The generation created in the step 3 is seen in system_distributed.cdc_generation_timestamps but not in system.cdc_generations_v3, thus there are used streams that the CDC based on raft doesn't know about. Instead of creating a new generation, the node should use the generation already committed to group0. Save the last CDC generation in the system.cdc_local during loading the topology state so that it is visible for CDC not based on raft. Fixes scylladb/scylladb#17819 (cherry picked from commit `4351eee1f6`)	2024-05-28 14:01:59 +00:00
Anna Stuchlik	7229c820cf	doc: describe Tablets in ScyllaDB This commit adds the main description of tablets and their benefits. The article can be used as a reference in other places across the docs where we mention tablets. (cherry picked from commit `b5c006aadf`) Closes scylladb/scylladb#18916	2024-05-28 11:27:53 +02:00
Pavel Emelyanov	67878af591	cql-pytest: Add validation of ALTER KEYSPACE WITH TABLETS There's a test that checks how ALTER changes the initial tablets value, but it equips the statement with `replication` parameters because of limitations that parser used to impose. Now the `tablets` parameters can come on their own, so add a new test. The old one is kept from compatibility considerations. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> (cherry picked from commit `1003391ed6`)	2024-05-28 02:07:58 +00:00
Pavel Emelyanov	2dbc555933	cql3: Fix parsing of ALTER KEYSPACE's tablets parameters When the `WITH` doesn't include the `replication` parameters, the `tablets` one is ignoded, even if it's present in the statement. That's not great, those two parameter sets are pretty much independent and should be parsed individually. Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> (cherry picked from commit `a172ef1bdf`)	2024-05-28 02:07:58 +00:00
Pavel Emelyanov	3b9c86dcf5	cql3: Remove unused ks_prop_defs/prepare_options() argument Signed-off-by: Pavel Emelyanov <xemul@scylladb.com> (cherry picked from commit `8a612da155`)	2024-05-28 02:07:58 +00:00
Raphael S. Carvalho	b6f3891282	test: Fix flakiness in topology_experimental_raft/test_tablets One source of flakiness is in test_tablet_metadata_propagates_with_schema_changes_in_snapshot_mode due to gossiper being aborted prematurely, and causing reconnection storm. Another is test_tablet_missing_data_repair which is flaky due an issue in python driver that session might not reconnect on rolling restart (tracked by https://github.com/scylladb/python-driver/issues/230) Refs #15356. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> (cherry picked from commit `e7246751b6`)	2024-05-27 18:21:21 +00:00
Raphael S. Carvalho	46220bd839	service: Use tablet read selector to determine which replica to account table stats Since we introduced the ability to revert migrations, we can no longer rely on ordering of transition stages to determine whether to account pending or leaving replica. Let's use read selector instead, which correctly has info which replica type has correct stats info. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> (cherry picked from commit `551bf9dd58`)	2024-05-27 18:21:21 +00:00
Raphael S. Carvalho	55a45e3486	storage_service: Fix race between tablet split and stats retrieval If tablet split is finalized while retrieving stats, the saved erm, used by all shards, will be invalidated. It can either cause incorrect behavior or crash if id is not available. It's worked by feeding local tablet map into the "coordinator" collecting stats from all shards. We will also no longer have a snapshot of erm shared between shards to help intra-node migration. This is simplified by serializing token metadata changes and the retrieval of the stats (latter should complete pretty fast, so it shouldn't block the former for any significant time). Fixes #18085. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> (cherry picked from commit `abcc68dbe7`)	2024-05-27 18:21:21 +00:00
Michał Jadwiszczak	1dd522edc8	test/auth_cluster/test_raft_service_levels: try to create sl in recovery (cherry picked from commit `af0b6bcc56`)	2024-05-27 18:20:36 +00:00
Michał Jadwiszczak	6d655e6766	service/qos/raft_sl_dda: reject changes to service levels in recovery mode When a cluster goes into recovery mode and service levels were migrated to raft, service levels become temporarily read-only. This commit adds a proper error message in case a user tries to do any changes. (cherry picked from commit `ee08d7fdad`)	2024-05-27 18:20:36 +00:00
Michał Jadwiszczak	54b9fdab03	service/qos/raft_sl_dda: extract raft_sl_dda steps to common function When setting/dropping a service level using raft data accessor, the same validation steps are executed (this_shard_id = 0 and guard is present). To not duplicate the calls in both functions, they can be extracted to a helper function. (cherry picked from commit `2b56158d13`)	2024-05-27 18:20:36 +00:00
Raphael S. Carvalho	13f8486cd7	replica: Fix tablet's compaction_groups_for_token_range() with unowned range File-based tablet streaming calls every shard to return data of every group that intersects with a given range. After dynamic group allocation, that breaks as the tablet range will only be present in a single shard, so an exception is thrown causing migration to halt during streaming phase. Ideally, only one shard is invoked, but that's out of the scope of this fix and compaction_groups_for_token_range() should return empty result if none of the local groups intersect with the range. Signed-off-by: Raphael S. Carvalho <raphaelsc@scylladb.com> (cherry picked from commit `eb8ef38543`) Closes scylladb/scylladb#18859	2024-05-27 15:20:04 +03:00
Kefu Chai	747ffd8776	migration_manager: do not reference moved-away smart pointer this change is inspired by clang-tidy. it warns like: ``` [752/852] Building CXX object service/CMakeFiles/service.dir/migration_manager.cc.o Warning: /home/runner/work/scylladb/scylladb/service/migration_manager.cc:891:71: warning: 'view' used after it was moved [bugprone-use-after-move] 891 \| db.get_notifier().before_create_column_family(keyspace, view, mutations, ts); \| ^ /home/runner/work/scylladb/scylladb/service/migration_manager.cc:886:86: note: move occurred here 886 \| auto mutations = db::schema_tables::make_create_view_mutations(keyspace, std::move(view), ts); \| ^ ``` in which, `view` is an instance of view_ptr which is a type with the semantics of shared pointer, it's backed by a member variable of `seastar::lw_shared_ptr<const schema>`, whose move-ctor actually resets the original instance. so we are actually accessing the moved-away pointer in ```c++ db.get_notifier().before_create_column_family(keyspace, view, mutations, ts) ``` so, in this change, instead of moving away from `view`, we create a copy, and pass the copy to `db::schema_tables::make_create_view_mutations()`. this should be fine, as the behavior of `db::schema_tables::make_create_view_mutations()` does not rely on if the `view` passed to it is a moved away from it or not. the change which introduced this use-after-move was `88a5ddabce` Refs `88a5ddabce` Fixes #18837 Signed-off-by: Kefu Chai <kefu.chai@scylladb.com> (cherry picked from commit `125464f2d9`) Closes scylladb/scylladb#18873	2024-05-27 15:18:29 +03:00
Anna Stuchlik	a87683c7be	doc: remove outdated MV error from Troubleshooting This commit removes the MV error message, which only affect older versions of ScyllaDB, from the Troubleshooting section. Fixes https://github.com/scylladb/scylladb/issues/17205 (cherry picked from commit `92bc8053e2`) Closes scylladb/scylladb#18855	2024-05-27 15:12:22 +03:00
Anna Stuchlik	eff7b0d42d	doc: replace Raft-disabled with Raft-enabled procedure This commit fixes the incorrect Raft-related information on the Handling Cluster Membership Change Failures page introduced with https://github.com/scylladb/scylladb/pull/17500. The page describes the procedure for when Raft is disabled. Since 6.0, Raft for consistent schema management is enabled and mandatory (cannot be disabled), this commit adds the procedure for Raft-enabled setups. (cherry picked from commit `6626d72520`) Closes scylladb/scylladb#18858	2024-05-27 15:11:09 +03:00
David Garcia	7dbcfe5a39	docs: docs: autogenerate metrics Autogenerates metrics documentation using the scripts/get_description.py script introduced in #17479 docs: add beta (cherry picked from commit `9eef3d6139`) Closes scylladb/scylladb#18857	2024-05-27 15:10:48 +03:00

... 4 5 6 7 8 ...

43077 Commits